CursorBench 4.0
We evaluate agents on ambiguous, multi-file tasks from real Cursor sessions. Higher scores are better.
More about CursorBench| Model | |||||
|---|---|---|---|---|---|
| 1 | Fable 5.1 Max | 51.8% | $17.28 | 117,236 | 128 |
| 2 | Fable 5.1 Extra High | 51.6% | $13.01 | 87,294 | 101 |
| 3 | Fable 5.1 High | 49.2% | $9.08 | 58,438 | 77 |
| 4 | Fable 5.1 Medium | 46.8% | $7.05 | 45,411 | 63 |
| 5 | Opus 5 Max | 46.6% | $11.95 | 85,384 | 106 |
| 6 | Opus 5 Extra High | 46.1% | $11.43 | 80,094 | 103 |
| 7 | Fable 5.1 Low | 45.1% | $5.44 | 34,795 | 51 |
| 8 | Opus 5 High | 44.7% | $9.00 | 61,405 | 86 |
| 9 | Opus 5 Medium | 43.3% | $6.94 | 45,272 | 72 |
| 10 | GPT-5.6 Sol Max | 41.7% | $8.23 | 42,944 | 99 |
| 11 | Muse Spark 1.3 Max | 41.6% | $2.64 | 52,005 | 98 |
| 12 | Grok 4.6 Extra High | 41.4% | $6.10 | 49,814 | 56 |
| 13 | GPT-5.6 Terra Max | 41.3% | $5.14 | 60,814 | 107 |
| 14 | Opus 5 Low | 40.7% | $4.87 | 31,995 | 57 |
| 15 | Grok 4.6 High | 40.4% | $5.20 | 41,387 | 48 |
| 16 | Gemini 3.8 Flash High | 39.6% | $4.70 | 162,565 | 324 |
| 17 | GPT-5.6 Sol Extra High | 37.7% | $4.40 | 24,729 | 55 |
| 18 | Muse Spark 1.3 Extra High | 37.5% | $2.10 | 40,891 | 83 |
| 19 | Gemini 3.8 Flash Medium | 37.3% | $4.06 | 128,364 | 290 |
| 20 | Grok 4.6 Medium | 36.1% | $3.48 | 24,893 | 40 |
| 21 | GPT-5.6 Luna Max | 35.9% | $1.03 | 87,284 | 208 |
| 22 | GPT-5.6 Sol High | 35.7% | $2.85 | 16,174 | 41 |
| 23 | Sonnet 5 Max | 34.1% | $7.17 | 149,257 | 140 |
| 24 | GPT-5.6 Terra Extra High | 33.6% | $1.81 | 23,436 | 43 |
| 25 | Grok 4.6 Low | 33.4% | $2.25 | 16,307 | 32 |
| 26 | Muse Spark 1.3 High | 33.4% | $1.66 | 30,654 | 69 |
| 27 | GPT-5.6 Luna Extra High | 33.0% | $0.44 | 40,598 | 98 |
| 28 | Muse Spark 1.3 Medium | 32.6% | $1.49 | 27,255 | 64 |
| 29 | Sonnet 5 Extra High | 32.0% | $4.55 | 83,373 | 102 |
| 30 | GPT-5.6 Sol Medium | 31.1% | $1.77 | 10,111 | 32 |
| 31 | Sonnet 5 High | 30.8% | $3.48 | 61,146 | 85 |
| 32 | GPT-5.6 Terra High | 30.7% | $1.11 | 13,162 | 33 |
| 33 | GPT-5.6 Luna High | 29.4% | $0.25 | 23,368 | 64 |
| 34 | Muse Spark 1.3 Low | 29.3% | $0.93 | 17,483 | 47 |
| 35 | Sonnet 5 Medium | 28.0% | $2.31 | 39,114 | 65 |
| 36 | Composer 2.5 | 27.7% | $0.68 | 17,347 | 41 |
| 37 | GPT-5.6 Terra Medium | 27.6% | $0.64 | 7,307 | 25 |
| 38 | GPT-5.6 Terra Low | 25.2% | $0.52 | 5,914 | 23 |
| 39 | GPT-5.6 Sol Low | 24.6% | $0.87 | 4,885 | 21 |
| 40 | Muse Spark 1.3 Minimal | 24.3% | $0.56 | 10,620 | 34 |
| 41 | Sonnet 5 Low | 24.1% | $1.39 | 23,772 | 46 |
| 42 | GPT-5.6 Luna Medium | 22.2% | $0.08 | 7,642 | 32 |
| 43 | GPT-5.6 Luna Low | 16.0% | $0.03 | 3,288 | 18 |
Changelog
Tasks
- CursorBench 4.0
- Introduced new long-horizon problems focused on edit, refactor, investigation, intent understanding, managing jobs, and design adherence.
Reporting
- Updated Sonnet 5 results to account for adjusted pricing.
Reporting
- Updated GPT-5.6 Terra and Luna results to account for adjusted pricing.
Reporting
- Updated GPT-5.6 Sol, Terra, and Luna results to account for cache write costs.
Tasks
- CursorBench 3.2
- Introduced instruction following and advanced tool use problems.
Tasks
- CursorBench 3.1
- Introduced problems focused on codebase understanding, bugfinding, planning, and code review.
- Improved grading criteria for some edit tasks.
Tasks
- CursorBench 3.0
- Initial set of tasks focused on edit, refactor, and bugfix problems.
Avg cost / task is computed by applying each model's published per-million-token pricing (input, cache read, cache write, and output) to the tokens it used on each task. Results are subject to variance; small differences in scores may not be statistically meaningful.