Simulated chain-of-thought faithfulness (perturbation-test-style) for 1200 configurations across five India-relevant reasoning tasks, five CoT methods, bias presence, and reasoning-step count, grounded in Turpin et al's unfaithful-explanation findings and Anthropic's perturbation-based faithfulness metrics. Forward-looking task predicts unfaithfulness from just a few cheap perturbation probes.
Simulates whether a model's stated reasoning causally determines its answer versus being post-hoc rationalization, via Lanham et al's perturbation-test methodology. 1200 configs across five tasks (three India-relevant: vernacular commonsense QA, legal clause interpretation, loan eligibility explanation), five methods (free CoT, zero-shot CoT, self-consistency, selection-inference, question decomposition). Ten sources (2022–2023) — two are Anthropic arXiv papers with peer-review status flagged honestly.
Lets Teams Auditing A Reasoning System Estimate Unfaithfulness From A Few Cheap Perturbation Probes Before Running The Full Expensive Audit.
Attribution 4.0 International (CC BY- 4.0)
To preview this file, you need to be a registered user. Please complete the registration process to gain access and continue viewing the content.
© 2026 - Copyright AIKosh. All rights reserved.