Cursor Insights Lab

Developers use Cursor across the full range of software work, giving us a unique view into how AI is changing everyday development and what the most advanced users are doing first. Cursor Insights is where we share what we’re learning.

The productivity index

  • AI edits / developer
  • Median PR size

The Productivity Index shows how much software developers are building with Cursor. We track two measures of development activity: AI edits per developer and median PR size.

Each measure is indexed to 100 in the first week of January. AI edits per developer nearly tripled, and median PR size more than doubled — evidence that developers are shipping more, and larger, changes with AI in the loop.

Developer Habits Report: Spring 2026

Our inaugural Developer Habits Report described the transformation happening across software development according to 5 themes:

  1. Developer acceleration.
  2. The economics of intelligence.
  3. The power user gap.
  4. The rise of context.
  5. The shift to automation.
Read more

A few more turns go a long way

Longer conversations are more likely to open a PR

  • Conversations that open a PR
  • Opened PRs that merge

As cloud-agent conversations get longer, they are much more likely to produce a PR. About a third of one-turn conversations open one, compared with more than four in five conversations with 12 or more turns. The share of those PRs that merge falls slightly, from 67% to 59%, but the increase in PRs more than makes up for that decline.

Human cloud agents, June 2026. PR rate is the share of conversations that open a pull request. Merge rate is the share of those PRs that merge.

Longer conversations are more likely to open a PR

  • Conversations that open a PR
  • Opened PRs that merge

Where power users pull away

Subagents show the widest adoption gap

Power users adopt several Cursor features at much higher rates than users in the median band. The largest gap is in subagents, which are used by 81% of power users on June 29 compared with 15% of the median band.

Difference in the share of each cohort that used the feature on June 29, 2026. Cohorts are based on agent-panel token usage.

Subagents show the widest adoption gap

The cost of intelligence

Updated: September, 2026

CursorBench 4.0 score vs cost per task

A scatter and line chart comparing Fable 5.1, Opus 5, Grok 4.6, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, Sonnet 5, Gemini 3.8 Flash, Muse Spark 1.3, and Composer 2.5 scores against average cost per task.55%CursorBench 4.0 score50%45%40%35%30%25%20%$18$15$12$9$6$3$0Average cost per taskFable 5.1Opus 5Gemini 3.8 FlashMuse Spark 1.3GPT-5.6 SolSonnet 5Grok 4.6

We maintain a ranking of model eval scores plotted against average cost per task, showing which models provide the most intelligence per dollar.

#ModelScoreAvg. Cost
  1. 1Fable 5.1 High49.2%$9.08
  2. 2Opus 5 High44.7%$9.00
  3. 3Grok 4.6 Extra High41.4%$6.10
  4. 4Gemini 3.8 Flash High39.6%$4.70
  5. 5Muse Spark 1.3 High33.4%$1.66
  6. 6GPT-5.6 Sol Medium31.1%$1.77
  7. 7Sonnet 5 High30.8%$3.48
  8. 8Composer 2.527.7%$0.68
  9. 9GPT-5.6 Terra Medium27.6%$0.64
  10. 10GPT-5.6 Luna Medium22.2%$0.08

Avg cost / task is computed by applying each model's published per-million-token pricing (input, cache read, cache write, and output) to the tokens it used on each task. Results are subject to variance; small differences in scores may not be statistically meaningful.

CursorBench 4.0 score vs cost per task

A scatter and line chart comparing Fable 5.1, Opus 5, Grok 4.6, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, Sonnet 5, Gemini 3.8 Flash, Muse Spark 1.3, and Composer 2.5 scores against average cost per task.55%CursorBench 4.0 score50%45%40%35%30%25%20%$18$15$12$9$6$3$0Average cost per taskFable 5.1Opus 5Gemini 3.8 FlashMuse Spark 1.3GPT-5.6 SolSonnet 5Grok 4.6