A TF-IDF + Logistic Regression classifier that identifies which government scheme (PM-KISAN or Ayushman Bharat) a citizen's Hinglish question is about, enabling automatic routing to the correct guideline document for grounded answering.
This model predicts which government scheme a citizen's question is about — currently PM-KISAN or Ayushman Bharat PM-JAY — given the question text in Hinglish. It's the routing component a real scheme-assistance system would need before it can even look up the correct guideline document to ground its answer in.
This Model Predicts Which Government Scheme A Citizen's Question Is About — Currently Pm-kisan Or Ayushman Bharat Pm-jay — Given The Question Text In Hinglish. It's The Routing Component A Real Scheme-assistance System Would Need Before It Can Even Look Up The Correct Guideline Document To Ground Its Answer In. The Model Is A Tf-idf Vectorizer (Unigrams And Bigrams) Feeding A Logistic Regression Classifier, Trained On The Companion Scheme-anchor-check Qa Dataset (14 Questions Across 2 Schemes). This Model Is Intended To Be Used Together With The Scheme-anchor-check Evaluation Pipeline: Once A Question Is Routed To The Correct Scheme, A Separate Grounded-qa Process (Using The Actual Verified Scheme Guideline As Context) Answers The Question And Gets Scored On Faithfulness To The Source Document — This Model Handles The Routing Step Of That Pipeline, Not The Answering Or Faithfulness-scoring Steps. Limitation: Trained On Only 2 Schemes And 14 Examples — Accuracy And Usefulness Will Improve Significantly As More Real, Verified Government Schemes Are Added To The Training Set.
Attribution 4.0 International (CC BY- 4.0)
© 2026 - Copyright AIKosh. All rights reserved.