What changed
A new dataset, named MarcoFurrer/swiss-building-law-rag-bench, has been published on Hugging Face, offering a benchmark for Retrieval-Augmented Generation (RAG) systems. This dataset was developed as part of a bachelor's thesis focused on optimizing RAG pipelines for German legal texts. The dataset contains Q&A pairs grounded in Swiss cantonal building law documents. It includes a German dataset with 318 entries, providing German Q&A pairs linked to specific article-level passages. Additionally, a multilingual dataset features 270 entries covering German, French, and Italian, with 9 query-document language combinations.
The dataset's structure includes fields such as id, question, answer, source_answer_passage, source_file, source_article, canton, doc_type, and question_type. The multilingual version adds question_lang and doc_lang. The source documents are official cantonal law publications, with a companion GitHub repository providing a notebook (01_download_dataset.ipynb) to download them. Evaluation metrics include standard IR measures like Recall@k, Precision@k, and MRR@k, with bootstrap 95% confidence intervals reported for all metrics. The dataset itself is licensed under CC BY 4.0, while the source PDFs remain under their respective cantonal authorities' ownership.
Dataset Contents:
- data/german/golden_dataset.jsonl: 318 entries, German Q&A pairs grounded to article-level passages.
- data/multilingual/golden_dataset.jsonl: 270 entries, cross-lingual Q&A (DE/FR/IT).
- results/*.csv: Ablation result tables from 14 experiments.
Evaluation Metrics:
- Recall@k
- Precision@k
- MRR@k
Standard IR metrics are evaluated with match="article", considering a chunk correct if it contains the source article identifier and originates from the correct source file. Bootstrap 95% confidence intervals (1,000 iterations) are provided for all metrics.
Why it matters for builders
This benchmark dataset is crucial for developers working on RAG systems, especially those dealing with complex, domain-specific information like legal documents. It offers a standardized way to evaluate and compare the performance of different RAG pipelines. By providing a structured set of questions and grounded answers derived from actual legal texts, it enables builders to fine-tune their systems for accuracy and relevance in legal contexts.
Practical impact
Developers can leverage the MarcoFurrer/swiss-building-law-rag-bench dataset to test and improve their RAG implementations. The dataset's multilingual capabilities are particularly useful for applications requiring cross-lingual understanding of legal documents. Builders can use the provided evaluation metrics (Recall@k, Precision@k, MRR@k) to quantitatively assess their models' performance and identify areas for optimization. The companion notebook for downloading source documents also facilitates a more comprehensive integration into existing workflows.
Caveats and source limits
The dataset contains 9 rows in its default view and has a total file size of 421 kB. The provided metrics are based on 1,000 bootstrap iterations for confidence intervals. While the dataset is grounded in official legal documents, its scope is specific to Swiss cantonal building law. The benchmark's performance figures are derived from the thesis's experiments and have not yet been independently verified by external benchmarks. The dataset was published on Hugging Face with 0 likes and 43 downloads at the time of reporting, indicating it is a recent addition to the community resources.
Sources
Claim check: 6/6 supported claims - 6 evidence links - 100% avg confidence
- A dataset named MarcoFurrer/swiss-building-law-rag-bench has been released on Hugging Face.supported - huggingface.co
- The dataset is an evaluation benchmark for Retrieval-Augmented Generation (RAG) systems on Swiss cantonal building law documents.supported - huggingface.co
- The dataset includes a German dataset with 318 entries and a multilingual dataset with 270 entries (DE/FR/IT).supported - huggingface.co
- Evaluation metrics include Recall@k, Precision@k, and MRR@k with 95% confidence intervals.supported - huggingface.co
- The dataset is licensed under CC BY 4.0.supported - huggingface.co
- The dataset had 43 downloads and 0 likes on Hugging Face.supported - huggingface.co
Caveats
- Single-source caution: verify critical details at the linked source.
Radar score 68/100 - how it was calculated
- Reliability 79: Source trust and evidence quality
- Freshness 78: Fresh model metadata
- Novelty 60: Novelty blends source metadata and enrichment
- Technical 61: Structured technical source signals
- Developer 67: Model developer utility
- Ecosystem 48: Single-source caution
- Confidence 98: Claims have reliable evidence