Code Debugging & Review Reasoning Dataset
Realistic bugs in Python, TypeScript, Java and SQL with the step-by-step reasoning a senior engineer used to find, explain and fix them, plus a test that proves the fix.
On request: this dataset is produced to your specification after you order. Typically 3–4 weeks for 200 records and 6–8 weeks for 1,000.
{
"id": "cdr-000412",
"test": "assert add_tag('a') == ['a']\nassert add_tag('b') == ['b']",
"symptom": "Tags from earlier calls show up in new results.",
"bug_type": "mutable-default-argument",
"language": "python",
"buggy_code": "def add_tag(tag, tags=[]):\n tags.append(tag)\n return tags",
"difficulty": "easy",
"fixed_code": "def add_tag(tag, tags=None):\n if tags is None:\n tags = []\n tags.append(tag)\n return tags",
"root_cause": "Mutable default argument shared across calls.",
"author_role": "senior backend engineer, 9 years",
"reasoning_steps": [
"The list only grows across calls, so state is shared between calls.",
"The only shared object is the default value of `tags`.",
"Python evaluates default arguments once, when the function is defined, so every call reuses the same list."
]
}Pricing
Fixed prices for standard packages, produced to your spec. Pick a package and a license to see the exact total, or ask for a custom quote.
| Feature | Starter | Pro | Custom |
|---|---|---|---|
| examples | 200 | 1,000 | Any quantity |
| What's included |
|
| Your quantity, annotations and deadline |
| Datasheet, license, checksums | Included | Included | Included |
| Research license | $690$3.45 / example | $2,490$2.49 / example | Quote |
| Commercial license | $1,990$9.95 / example | $7,900$7.90 / example | Quote |
| Exclusive license | $6,900$34.50 / example | $23,900$23.90 / example | Quote |
How ordering works
- 1Choose a package and license and place your order. No payment is taken online.
- 2Within 48 hours our team contacts you to confirm the spec and invoices a 50% deposit.
- 3We produce and deliver in parts so you can review early. Typically 3–4 weeks for 200 records and 6–8 weeks for 1,000.
- 4Pay the balance and download the final dataset.
Custom quote
Need another quantity, an Enterprise license or annotations for this dataset? Available here:
- Bug type and difficulty labels
- Code review comments
- Multi-file repository bugs
- Your languages and frameworks
Overview
Coding assistants are good at writing new code and weaker at finding why existing code fails. Public bug datasets are mined from open-source commits: they show the diff, but not the reasoning that led to it, and many are already in model training data.
Each record in this dataset is written by a senior engineer: a realistic piece of buggy code, the symptom a user would report, the step-by-step reasoning to locate the cause, the fix and a test that fails before and passes after. Every record is executed to confirm it.
Example records
Illustrative records in the delivered format. Request the sample pack for real records written by our experts.
Record 1 of 1 · illustrative
{ "id": "cdr-000412", "test": "assert add_tag('a') == ['a']\nassert add_tag('b') == ['b']", "symptom": "Tags from earlier calls show up in new results.", "bug_type": "mutable-default-argument", "language": "python", "buggy_code": "def add_tag(tag, tags=[]):\n tags.append(tag)\n return tags", "difficulty": "easy", "fixed_code": "def add_tag(tag, tags=None):\n if tags is None:\n tags = []\n tags.append(tag)\n return tags", "root_cause": "Mutable default argument shared across calls.", "author_role": "senior backend engineer, 9 years", "reasoning_steps": [ "The list only grows across calls, so state is shared between calls.", "The only shared object is the default value of `tags`.", "Python evaluates default arguments once, when the function is defined, so every call reuses the same list." ] }
Get a free sample pack
We email you real examples from this dataset.
Technical specifications
- Authors
- Senior software engineers (5+ years), reviewed by a second engineer
- Languages
- Python, TypeScript/JavaScript, Java, Go, SQL; others on request
- Bug types
- Logic, off-by-one, concurrency, null handling, SQL joins, API misuse, performance, security
- Verification
- Every test fails on the buggy code and passes on the fix (executed in CI)
- Fields
- Buggy code, symptom, reasoning steps, root cause, fix, test, bug type, difficulty
- Format
- JSONL (one record per line), UTF-8; CSV on request
- AI use
- Written by people; screened for LLM-generated text. Any AI assistance is disclosed in the datasheet
- License
- Research, Commercial, Enterprise or Exclusive
Use cases
Coding model fine-tuning
Teach models to diagnose before they patch, with explicit reasoning about the root cause.
Code review assistants
Train reviewers that explain why a change is risky, not just flag it.
Uncontaminated evaluation
Exclusive, never-published bugs give an honest measure of debugging ability.
Developer education
Explanations written for humans make good material for teaching junior engineers.
How this data is made
- 1Problems sourced or written by vetted domain experts
- 2Each solution written step by step by a person, never generated
- 3A second expert checks every step and the final answer
- 4Disagreements resolved by a lead reviewer
- 5Screened for copied text and LLM-generated writing
- 6Difficulty, topic and skill labels added
- 7Exported as JSONL with a schema you can rely on
- 8Datasheet and license packaged with the delivery
Provenance & legal
- 100% made by people: no scraping, no generative AI
- Datasheet documenting how the data was made, checked and its limitations
- Commercial license that lets you keep models trained on the data
- Supports training-data documentation under the EU AI Act
Dataset-specific notes
- Code is written for the dataset, not copied from employers or open-source projects, so licensing is clean.
- Every record is executed: the test must fail before the fix and pass after.
- Security bugs are limited to defensive examples (input validation, injection fixes), never working exploits.
Delivery format
One JSON record per line, with the fields shown in the example records above and a SCHEMA.md describing each one.
Each delivery contains:
code-debugging-reasoning-v1/ ├── data/train.jsonl ├── data/sample.jsonl ├── SCHEMA.md ├── DATASHEET.md ├── LICENSE.pdf └── checksums.sha256
Frequently asked questions
Is the code copied from GitHub?
No. Engineers write the code for the dataset, so it is not in public training data and carries no open-source license obligations.
How do you check the fixes?
Each record ships with a test. We run it in CI: it must fail on the buggy version and pass on the fixed one.
Can you cover our stack?
Yes. Tell us the languages, frameworks and bug categories you care about and we will quote the mix.
How does ordering and payment work?
Create a free account and place your order on this page; nothing is charged online. Our team contacts you to confirm the spec and invoices a 50% deposit, with the balance due on final delivery. The deposit is refunded in full if we cannot deliver the agreed spec.
What is the difference between non-exclusive and Exclusive?
With Research or Commercial we may license the examples produced for your order to other buyers later. With Exclusive they are never licensed to anyone else or added to our catalog.
Can I order a different quantity or extra annotations?
Yes. Request a custom quote with the quantity, annotations and deadline you need. Prices per record fall as quantity grows.
Need a variation?
Different topics, languages, difficulty, quantity or labels? Describe it and we reply within 48 hours with matching samples and a quote.
Request a quoteRelated datasets
Browse all datasets- Software QA Test CasesEvaluation & testsManual test cases and bug reports written by professional QA engineers for real app flows: sign-up, search, checkout, payments, forms and settings.On requestFrom $490
- AI Agent Task SuiteEvaluation & testsRealistic multi-step tasks for browser, tool-using and coding agents, each with a start state, allowed tools, a human reference solution and an objective success check.On requestFrom $790
- Professional Reasoning TracesReasoningHow experienced accountants, supply-chain planners, marketers and engineers think through real work decisions, written out step by step by the professionals themselves.On requestFrom $790
- Vietnamese Math ReasoningReasoningSecondary-school and university-entrance math problems in Vietnamese, each solved step by step by a qualified teacher and checked by a second one.On requestFrom $490