Simulated within-rollout accumulated policy error and failure risk for 1920 imitation-learning configurations across four methods (behavior cloning, DAgger, DART, GAIL), five India-relevant tasks, demonstration count, rollout horizon, and causal-confusion conditions, grounded in Ross and Bagnell's quadratic-vs-linear compounding-error bounds. Forward-looking task predicts full-rollout failure from only the first 25 percent of a single rollout.
This dataset simulates covariate shift, the central failure mode of imitation learning: a behavior-cloned policy accumulates small per-step errors that drift it off the expert's demonstrated states, compounding quadratically over the rollout horizon, a bound Ross and Bagnell (2010) prove and Rajaraman et al (2020) show is an unavoidable information-theoretic lower bound. DAgger fixes this by querying the expert on states the current policy visits, reducing the bound to linear. The simulation spans 1920 configurations across five India-relevant tasks including autonomous e-rickshaw driving and agri-robot weeding. Ten grounding sources are documented in imitation_theory.py and the README Research basis.
Lets Safety Monitors Watching A Live Imitation-learned Policy, Such As An Autonomous E-rickshaw Or Agricultural Robot, Estimate From Only The First Quarter Of A Rollout Whether Accumulating Error Is Headed For A Failure By The End.
Attribution 4.0 International (CC BY- 4.0)
To preview this file, you need to be a registered user. Please complete the registration process to gain access and continue viewing the content.
© 2026 - Copyright AIKosh. All rights reserved.