Cursor Insights Lab

Developers use Cursor across the full range of software work, giving us a unique view into how AI is changing everyday development and what the most advanced users are doing first. Cursor Insights is where we share what we’re learning.

The productivity index

  • AI edits / developer
  • Median PR size

The Productivity Index shows how much software developers are building with Cursor. We track two measures of development activity: AI edits per developer and median PR size.

Each measure is indexed to 100 in the first week of January. AI edits per developer nearly tripled, and median PR size more than doubled — evidence that developers are shipping more, and larger, changes with AI in the loop.

Developer Habits Report: Spring 2026

Our inaugural Developer Habits Report described the transformation happening across software development according to 5 themes:

  1. Developer acceleration.
  2. The economics of intelligence.
  3. The power user gap.
  4. The rise of context.
  5. The shift to automation.
Read more

The asymmetry of token consumption

Dan Hollick

The top 10% of users consume nearly two-thirds of all tokensShare of users and token usage over the last four weeks

  • Top 1%
  • p90–p99
  • p50–p90
  • Bottom 50%

Over the last four weeks, the most engaged 10% of Cursor users accounted for nearly two-thirds of all token use. We compare engagement cohorts across their language mix and model choices, then look at how much agent work they run at once.

Parallelism produces the clearest gap. In the most engaged cohort, 86.6% started a chat while another agent was running, compared with 25.2% in the p50–p90 band. We explore whether this way of working offers an early look at how more developers will use agents as they become better at completing and checking work on their own.

Read more

The top 10% of users consume nearly two-thirds of all tokensShare of users and token usage over the last four weeks

  • Top 1%
  • p90–p99
  • p50–p90
  • Bottom 50%

A few more turns go a long way

Longer conversations are more likely to open a PR

  • Conversations that open a PR
  • Opened PRs that merge

As cloud-agent conversations get longer, they are much more likely to produce a PR. About a third of one-turn conversations open one, compared with more than four in five conversations with 12 or more turns. The share of those PRs that merge falls slightly, from 67% to 59%, but the increase in PRs more than makes up for that decline.

Human cloud agents, June 2026. PR rate is the share of conversations that open a pull request. Merge rate is the share of those PRs that merge.

Longer conversations are more likely to open a PR

  • Conversations that open a PR
  • Opened PRs that merge

Where power users pull away

Subagents show the widest adoption gap

Power users adopt several Cursor features at much higher rates than users in the median band. The largest gap is in subagents, which are used by 81% of power users on June 29 compared with 15% of the median band.

Difference in the share of each cohort that used the feature on June 29, 2026. Cohorts are based on agent-panel token usage.

Subagents show the widest adoption gap

The cost of intelligence

Updated: September, 2026

CursorBench 3.2 score vs cost per task

A scatter and line chart comparing Fable 5.1, Fable 5, Opus 5, Opus 4.8, Grok 4.6, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, GPT-5.5, Sonnet 5, GLM 5.2, Composer 2.5, Gemini 3.8 Flash, Gemini 3.7 Flash, Kimi K3, and Kimi K2.7 Code scores against average cost per task.75%CursorBench 3.2 score70%65%60%55%50%45%$10$5$0Average cost per taskGrok 4.6Fable 5.1Gemini 3.8 FlashOpus 5Kimi K3GPT-5.6 SolSonnet 5Composer 2.5GPT-5.6 TerraGPT-5.6 Luna

We maintain a ranking of model eval scores plotted against average cost per task, showing which models provide the most intelligence per dollar.

#ModelScoreAvg. Cost
  1. 1Grok 4.6 High69.9%$2.34
  2. 2Fable 5.1 High69.4%$4.80
  3. 3Gemini 3.8 Flash High69.2%$2.38
  4. 4Opus 5 High66.7%$3.91
  5. 5Fable 5 High66.5%$8.77
  6. 6Gemini 3.7 Flash High61.6%$1.20
  7. 7Kimi K3 Max60.8%$2.70
  8. 8GPT-5.6 Sol Medium60.0%$1.95
  9. 9Opus 4.8 High58.0%$3.15
  10. 10Sonnet 5 High56.9%$2.13

Avg cost / task is computed by applying each model's published per-million-token pricing (input, cache read, cache write, and output) to the tokens it used on each task. Results are subject to variance; small differences in scores may not be statistically meaningful.

CursorBench 3.2 score vs cost per task

A scatter and line chart comparing Fable 5.1, Fable 5, Opus 5, Opus 4.8, Grok 4.6, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, GPT-5.5, Sonnet 5, GLM 5.2, Composer 2.5, Gemini 3.8 Flash, Gemini 3.7 Flash, Kimi K3, and Kimi K2.7 Code scores against average cost per task.75%CursorBench 3.2 score70%65%60%55%50%45%$10$5$0Average cost per taskGrok 4.6Fable 5.1Gemini 3.8 FlashOpus 5Kimi K3GPT-5.6 SolSonnet 5Composer 2.5GPT-5.6 TerraGPT-5.6 Luna