Nearly 1,000 crowd-sourced Python programming problems, used to benchmark program synthesis and code-generation models.
MBPP (Mostly Basic Python Problems) is a benchmark dataset containing approximately 1,000 crowd-sourced Python programming problems intended to be solvable by entry-level programmers. Each problem includes a natural-language task description, a reference solution, and a small set of test cases for automatic verification. The dataset was developed by Google Research to provide a simpler, broader complement to more difficult code-generation benchmarks.
Mbpp Is Used To Evaluate The Basic Program Synthesis And Code-generation Capabilities Of Language Models, Particularly Their Ability To Translate Straightforward Natural-language Instructions Into Correct, Executable Python Code. It Is Commonly Used Alongside Humaneval To Provide A Fuller Picture Of A Model's Coding Proficiency Across Different Difficulty Levels, And To Study Few-shot And Fine-tuned Code Generation Performance.
Attribution 4.0 International (CC BY- 4.0)
© 2026 - Copyright AIKosh. All rights reserved.