Indian Flag
Government Of India
A-
A
A+
ORGANISATION
DocVQA

DocVQA

Document image dataset paired with question-answer pairs, used to train and evaluate document visual question-answering models.

About Dataset

DocVQA is a dataset of scanned document images paired with natural-language questions and corresponding answers, designed to support research in document visual question answering. The documents span diverse layouts and content types such as forms, letters, and reports, requiring models to jointly reason over visual layout and embedded text to answer questions accurately. The dataset was created to push document understanding beyond simple text extraction toward layout-aware comprehension.

Purpose of Dataset

Docvqa Is Used To Train And Benchmark Models That Combine Computer Vision And Natural Language Understanding To Answer Questions About Document Images. It Supports Research In Document Ai, Information Extraction, And Multimodal Reasoning Over Structured And Semi-structured Documents. The Dataset Is Widely Used To Evaluate Vision-language Models On Real-world Document Understanding Tasks Such As Form Processing And Invoice Analysis.

Activity Overview Activity Overview

  • Downloads0
  • Redirect 1
  • File Size 0
  • Views 8

Tags Tags

  • Multimodal
  • Visual Question Answering
  • OCR
  • Document Understanding

License Control License Control

Other