FERMAT is a benchmark designed to test the multimodal reasoning and auto-evaluation capabilities of VLMs using real-world handwritten math problems.
FERMAT has 244 handwritten math solutions, carefully annotated across core mathematical domains - 🔢 Arithmetic | 📏 Algebra | 📐 Geometry | 📏 Mensuration | 🎲 Probability | 📊 Statistics | 📐 Trigonometry | 📈 Calculus.
Each solution features realistic student mistakes categorized along four key axes:
Additionally, some solutions contain superficial variations that don't actually affect correctness (e.g., "16 cm" vs. "16.0 cm")—perfect for testing the subtlety of your models!
Attribution 4.0 International (CC BY- 4.0)
© 2026 - Copyright AIKosh. All rights reserved.