CursorBench 4.0

We evaluate agents on ambiguous, multi-file tasks from real Cursor sessions. Higher scores are better.

More about CursorBench
A scatter and line chart comparing Fable 5.1, Opus 5, Grok 4.7, Grok 4.6, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, Sonnet 5, Gemini 3.8 Flash, Muse Spark 1.3, and Composer 2.5 scores against average cost per task.55%CursorBench 4.0 score50%45%40%35%30%25%20%$18$15$12$9$6$3$0Average cost per taskFable 5.1Opus 5Gemini 3.8 FlashGPT-5.6 SolSonnet 5Grok 4.7
Model
1Fable 5.1 Max51.8%$17.28117,236128
2Fable 5.1 Extra High51.6%$13.0187,294101
3Fable 5.1 High49.2%$9.0858,43877
4Fable 5.1 Medium46.8%$7.0545,41163
5Opus 5 Max46.6%$11.9585,384106
6Grok 4.7 Extra High46.3%$6.0170,14188
7Opus 5 Extra High46.1%$11.4380,094103
8Fable 5.1 Low45.1%$5.4434,79551
9Opus 5 High44.7%$9.0061,40586
10Grok 4.7 High43.9%$4.6956,38271
11Opus 5 Medium43.3%$6.9445,27272
12GPT-5.6 Sol Max41.7%$8.2342,94499
13Grok 4.7 Medium41.6%$3.4936,68360
14Muse Spark 1.3 Max41.6%$2.6452,00598
15Grok 4.6 Extra High41.4%$6.1049,81456
16GPT-5.6 Terra Max41.3%$5.1460,814107
17Opus 5 Low40.7%$4.8731,99557
18Grok 4.6 High40.4%$5.2041,38748
19Gemini 3.8 Flash High39.6%$4.70162,565324
20GPT-5.6 Sol Extra High37.7%$4.4024,72955
21Muse Spark 1.3 Extra High37.5%$2.1040,89183
22Gemini 3.8 Flash Medium37.3%$4.06128,364290
23Grok 4.6 Medium36.1%$3.4824,89340
24GPT-5.6 Luna Max35.9%$1.0387,284208
25GPT-5.6 Sol High35.7%$2.8516,17441
26Sonnet 5 Max34.1%$7.17149,257140
27GPT-5.6 Terra Extra High33.6%$1.8123,43643
28Grok 4.6 Low33.4%$2.2516,30732
29Muse Spark 1.3 High33.4%$1.6630,65469
30Grok 4.7 Low33.1%$1.5815,67740
31GPT-5.6 Luna Extra High33.0%$0.4440,59898
32Muse Spark 1.3 Medium32.6%$1.4927,25564
33Sonnet 5 Extra High32.0%$4.5583,373102
34GPT-5.6 Sol Medium31.1%$1.7710,11132
35Sonnet 5 High30.8%$3.4861,14685
36GPT-5.6 Terra High30.7%$1.1113,16233
37GPT-5.6 Luna High29.4%$0.2523,36864
38Muse Spark 1.3 Low29.3%$0.9317,48347
39Sonnet 5 Medium28.0%$2.3139,11465
40Composer 2.527.7%$0.6817,34741
41GPT-5.6 Terra Medium27.6%$0.647,30725
42GPT-5.6 Terra Low25.2%$0.525,91423
43GPT-5.6 Sol Low24.6%$0.874,88521
44Muse Spark 1.3 Minimal24.3%$0.5610,62034
45Sonnet 5 Low24.1%$1.3923,77246
46GPT-5.6 Luna Medium22.2%$0.087,64232
47GPT-5.6 Luna Low16.0%$0.033,28818

Changelog

Tasks

  • CursorBench 4.0
    • Introduced new long-horizon problems focused on edit, refactor, investigation, intent understanding, managing jobs, and design adherence.

Reporting

  • Updated Sonnet 5 results to account for adjusted pricing.

Reporting

  • Updated GPT-5.6 Terra and Luna results to account for adjusted pricing.

Reporting

  • Updated GPT-5.6 Sol, Terra, and Luna results to account for cache write costs.

Tasks

  • CursorBench 3.2
    • Introduced instruction following and advanced tool use problems.

Tasks

  • CursorBench 3.1
    • Introduced problems focused on codebase understanding, bugfinding, planning, and code review.
    • Improved grading criteria for some edit tasks.

Tasks

  • CursorBench 3.0
    • Initial set of tasks focused on edit, refactor, and bugfix problems.

Avg cost / task is computed by applying each model's published per-million-token pricing (input, cache read, cache write, and output) to the tokens it used on each task. Results are subject to variance; small differences in scores may not be statistically meaningful.