Image PII redaction tuned for indian medical documents
Most PII redactors stop at plain text. These models power document-pii-redactor, which also redacts document images and is light enough to deploy on CPU. They are trained to understand Indian names, documents, and contexts, and the text model works across Indian languages. The main contribution is the PII token classifier — OCR is just the pluggable input stage in front of it. It defaults to lightweight Tesseract, which keeps memory low and works well for PDFs and good-quality images; for more difficult or blurred images, Bring-your-own OCR lets a model like Nemotron OCR (or Textract, Google Vision, etc) plug straight in for better results (example notebook).
Other
ekacare
Transformers
PyTorch
Open
Healthcare, Wellness and Family Welfare
10/08/26 14:17:37
0
Other
© 2026 - Copyright AIKosh. All rights reserved.