Nearly 1,000 crowd-sourced Python programming problems, used to benchmark program synthesis and code-generation models.
MBPP (Mostly Basic Python Problems) is a benchmark dataset containing approximately 1,000 crowd-sourced Python programming problems intended to be solvable by entry-level programmers. Each problem includes a natural-language task description, a reference solution, and a small set of test cases for automatic verification. The dataset was developed by Google Research to provide a simpler, broader complement to more difficult code-generation benchmarks.
MBPP is used to evaluate the basic program synthesis and code-generation capabilities of language models, particularly their ability to translate straightforward natural-language instructions into correct, executable Python code. It is commonly used alongside HumanEval to provide a fuller picture of a model's coding proficiency across different difficulty levels, and to study few-shot and fine-tuned code generation performance.
Attribution 4.0 International (CC BY- 4.0)
© 2026 - Copyright AIKosh. All rights reserved.