CursorBench 4.0

We evaluate agents on ambiguous, multi-file tasks from real Cursor sessions. Higher scores are better.

More about CursorBench
A scatter and line chart comparing Fable 5.1, Opus 5, Grok 4.6, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, Sonnet 5, Gemini 3.8 Flash, Muse Spark 1.3, and Composer 2.5 scores against average cost per task.55%CursorBench 4.0 score50%45%40%35%30%25%20%$18$15$12$9$6$3$0Average cost per taskFable 5.1Opus 5Gemini 3.8 FlashMuse Spark 1.3GPT-5.6 SolSonnet 5Grok 4.6
Model
1Fable 5.1 Max51.8%$17.28117,236128
2Fable 5.1 Extra High51.6%$13.0187,294101
3Fable 5.1 High49.2%$9.0858,43877
4Fable 5.1 Medium46.8%$7.0545,41163
5Opus 5 Max46.6%$11.9585,384106
6Opus 5 Extra High46.1%$11.4380,094103
7Fable 5.1 Low45.1%$5.4434,79551
8Opus 5 High44.7%$9.0061,40586
9Opus 5 Medium43.3%$6.9445,27272
10GPT-5.6 Sol Max41.7%$8.2342,94499
11Muse Spark 1.3 Max41.6%$2.6452,00598
12Grok 4.6 Extra High41.4%$6.1049,81456
13GPT-5.6 Terra Max41.3%$5.1460,814107
14Opus 5 Low40.7%$4.8731,99557
15Grok 4.6 High40.4%$5.2041,38748
16Gemini 3.8 Flash High39.6%$4.70162,565324
17GPT-5.6 Sol Extra High37.7%$4.4024,72955
18Muse Spark 1.3 Extra High37.5%$2.1040,89183
19Gemini 3.8 Flash Medium37.3%$4.06128,364290
20Grok 4.6 Medium36.1%$3.4824,89340
21GPT-5.6 Luna Max35.9%$1.0387,284208
22GPT-5.6 Sol High35.7%$2.8516,17441
23Sonnet 5 Max34.1%$7.17149,257140
24GPT-5.6 Terra Extra High33.6%$1.8123,43643
25Grok 4.6 Low33.4%$2.2516,30732
26Muse Spark 1.3 High33.4%$1.6630,65469
27GPT-5.6 Luna Extra High33.0%$0.4440,59898
28Muse Spark 1.3 Medium32.6%$1.4927,25564
29Sonnet 5 Extra High32.0%$4.5583,373102
30GPT-5.6 Sol Medium31.1%$1.7710,11132
31Sonnet 5 High30.8%$3.4861,14685
32GPT-5.6 Terra High30.7%$1.1113,16233
33GPT-5.6 Luna High29.4%$0.2523,36864
34Muse Spark 1.3 Low29.3%$0.9317,48347
35Sonnet 5 Medium28.0%$2.3139,11465
36Composer 2.527.7%$0.6817,34741
37GPT-5.6 Terra Medium27.6%$0.647,30725
38GPT-5.6 Terra Low25.2%$0.525,91423
39GPT-5.6 Sol Low24.6%$0.874,88521
40Muse Spark 1.3 Minimal24.3%$0.5610,62034
41Sonnet 5 Low24.1%$1.3923,77246
42GPT-5.6 Luna Medium22.2%$0.087,64232
43GPT-5.6 Luna Low16.0%$0.033,28818

Changelog

Tasks

  • CursorBench 4.0
    • Introduced new long-horizon problems focused on edit, refactor, investigation, intent understanding, managing jobs, and design adherence.

Reporting

  • Updated Sonnet 5 results to account for adjusted pricing.

Reporting

  • Updated GPT-5.6 Terra and Luna results to account for adjusted pricing.

Reporting

  • Updated GPT-5.6 Sol, Terra, and Luna results to account for cache write costs.

Tasks

  • CursorBench 3.2
    • Introduced instruction following and advanced tool use problems.

Tasks

  • CursorBench 3.1
    • Introduced problems focused on codebase understanding, bugfinding, planning, and code review.
    • Improved grading criteria for some edit tasks.

Tasks

  • CursorBench 3.0
    • Initial set of tasks focused on edit, refactor, and bugfix problems.

Avg cost / task is computed by applying each model's published per-million-token pricing (input, cache read, cache write, and output) to the tokens it used on each task. Results are subject to variance; small differences in scores may not be statistically meaningful.