15,000 human-generated instruction-following records across multiple task categories, released under a permissive license.
Databricks Dolly 15k is a dataset of approximately 15,000 instruction-following records written entirely by Databricks employees, covering task categories such as brainstorming, classification, question answering, summarization, and creative writing. Unlike many instruction datasets derived from model outputs, Dolly 15k was created by human annotators from scratch, making it free of restrictions tied to other models' usage terms. Each record includes an instruction, optional context, and a human-written response.
Dolly 15k Is Used To Instruction-tune Large Language Models So They Can Follow Natural-language Directions Across A Diverse Set Of Task Types. Its Permissive Licensing Makes It Particularly Valuable For Building And Commercially Deploying Open-source Instruction-tuned Models Without Dependency On Closed-model-generated Data. Researchers Also Use It To Study Human-authored Instruction Data Quality Compared To Synthetically Generated Alternatives.
Attribution-ShareAlike 3.0 (CC BY-SA 3.0)
© 2026 - Copyright AIKosh. All rights reserved.