Description Simulated hallucination rate across retrieved-passage-count (k) and RAG method for 600 configurations spanning five tasks and five methods (naive concatenation, FiD, REPLUG, Self-RAG, robust-trained RAG), grounded in the foundational and 2023 retrieval-augmented-generation literature. Forward-looking task predicts faithfulness at deployment-scale retrieval (k=20) from a single cheap top-passage probe (k=1).
Simulates hallucination risk in RAG as a function of retrieval depth and method robustness to noisy passages. 600 configs across five tasks (two India-relevant: vernacular-language QA, agricultural advisory QA), k swept 1–20. Ten sources (2020–2024) in rag_theory.py, including RAG, FiD, REPLUG, Self-RAG, and the RGB benchmark.
Lets Teams Prototyping Rag Estimate Deployment-depth Faithfulness From A Single Cheap Probe Before Running A Full Expensive Evaluation.
Attribution 4.0 International (CC BY- 4.0)
To preview this file, you need to be a registered user. Please complete the registration process to gain access and continue viewing the content.
© 2026 - Copyright AIKosh. All rights reserved.