CursorBench 3.2

We evaluate agents on ambiguous, multi-file tasks from real Cursor sessions. Higher scores are better.

More about CursorBench
A scatter and line chart comparing Fable 5.1, Fable 5, Opus 5, Opus 4.8, Grok 4.6, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, GPT-5.5, Sonnet 5, GLM 5.2, Composer 2.5, Gemini 3.7 Flash, Gemini 3.6 Flash, Kimi K3, and Kimi K2.7 Code scores against average cost per task.75%CursorBench 3.2 score70%65%60%55%50%45%$10$5$0Average cost per taskGrok 4.6Fable 5.1Opus 5Gemini 3.7 FlashKimi K3GPT-5.6 SolSonnet 5Composer 2.5GPT-5.6 TerraGPT-5.6 Luna
Model
1Fable 5.1 Max73.4%$9.6472,06070
2Fable 5.1 Extra High72.8%$6.9651,34955
3Grok 4.6 Extra High70.8%$2.8141,13646
4Fable 5 Max70.5%$17.32103,52572
5Opus 5 Max70.0%$8.2361,83878
6Grok 4.6 High69.9%$2.3432,44939
7Fable 5.1 High69.4%$4.8033,15344
8Opus 5 Extra High69.3%$7.3554,23972
9Fable 5 Extra High68.4%$11.7364,97156
10Fable 5.1 Medium68.0%$3.5323,80136
11GPT-5.6 Sol Max67.2%$5.6928,32048
12Grok 4.6 Medium67.1%$1.2817,94229
13Opus 5 High66.7%$3.9127,93248
14Fable 5 High66.5%$8.7743,74748
15Fable 5.1 Low66.2%$2.9019,52231
16Fable 5 Medium65.2%$6.8030,36641
17GPT-5.6 Terra Max64.9%$2.3132,96947
18GPT-5.6 Sol Extra High64.5%$3.8819,69938
19Opus 5 Medium64.3%$3.2923,61244
20GPT-5.6 Sol High63.5%$2.7913,86732
21Opus 5 Low62.8%$2.5518,52937
22Opus 4.8 Max62.3%$5.7771,41144
23Fable 5 Low62.1%$4.4618,18231
24Gemini 3.7 Flash High61.6%$1.2038,44899
25Sonnet 5 Max61.5%$4.3092,88286
26GPT-5.6 Luna Max61.1%$0.3987,97361
27Grok 4.6 Low61.0%$0.7010,65823
28Kimi K3 Max60.8%$2.7038,42857
29GPT-5.6 Sol Medium60.0%$1.959,74727
30Kimi K3 High59.7%$1.8926,84647
31Opus 4.8 Extra High59.4%$4.5051,12140
32GPT-5.6 Terra Extra High59.2%$1.1516,08929
33Gemini 3.7 Flash Medium59.0%$0.9530,95382
34Sonnet 5 Extra High58.7%$2.7752,87167
35GPT-5.5 High58.4%$2.0512,18328
36GPT-5.5 Extra High58.4%$2.8517,53432
37Opus 4.8 High58.0%$3.1533,54833
38GPT-5.6 Luna Extra High57.7%$0.2322,48048
39Sonnet 5 High56.9%$2.1339,48357
40GPT-5.6 Luna High56.8%$0.1615,14140
41Opus 4.8 Medium56.1%$2.8128,38432
42Composer 2.556.1%$0.4414,28633
43GLM 5.2 Max55.0%$1.7635,94658
44GPT-5.6 Terra High54.2%$0.719,46823
45GPT-5.5 Medium53.8%$1.518,52225
46Gemini 3.7 Flash Low53.8%$0.7420,59468
47Gemini 3.6 Flash High53.5%$1.5630,43664
48Opus 4.8 Low53.1%$2.0219,62427
49GPT-5.6 Sol Low52.6%$1.015,10419
50Sonnet 5 Medium52.4%$1.4426,20046
51GLM 5.2 High51.5%$1.1921,82949
52Gemini 3.6 Flash Medium51.2%$1.4828,51162
53Kimi K3 Low50.5%$0.9913,00733
54GPT-5.6 Terra Medium50.3%$0.496,22220
55Kimi K2.7 Code49.7%$1.4331,24758
56GPT-5.6 Luna Medium47.7%$0.087,09528
57Sonnet 5 Low47.7%$0.8716,26933
58Gemini 3.6 Flash Low47.4%$1.1320,52950
59GPT-5.6 Terra Low46.9%$0.425,31219
60GPT-5.5 Low46.6%$0.985,16820
61GPT-5.6 Luna Low37.6%$0.033,20917

Changelog

Reporting

  • Updated Sonnet 5 results to account for adjusted pricing.

Reporting

  • Updated GPT-5.6 Terra and Luna results to account for adjusted pricing.

Reporting

  • Updated GPT-5.6 Sol, Terra, and Luna results to account for cache write costs.

Tasks

  • CursorBench 3.2
    • Introduced instruction following and advanced tool use problems.

Tasks

  • CursorBench 3.1
    • Introduced problems focused on codebase understanding, bugfinding, planning, and code review.
    • Improved grading criteria for some edit tasks.

Tasks

  • CursorBench 3.0
    • Initial set of tasks focused on edit, refactor, and bugfix problems.

Avg cost / task is computed by applying each model's published per-million-token pricing (input, cache read, cache write, and output) to the tokens it used on each task. Results are subject to variance; small differences in scores may not be statistically meaningful.