Grok 4.6

A frontier model for long-running agents and ambitious interactive and visual work.

Announcements

推出 Grok 4.6

推出 Grok 4.6

專為長時間運行的 Agent,以及更具企圖心的互動與視覺工作而打造。

Grok 4.6 模型卡

Grok 4.6 模型卡

Grok 4.6 的官方模型卡,介紹模型能力基準測試與安全防護評估。

Built for coding and knowledge work alike

Long-horizon problems

Grok 4.6 stays with complex tasks across many steps. It uses tools, checks its work, adjusts its approach, and keeps moving toward a finished result.

Ambitious engineering

Grok 4.6 works across large codebases and extended engineering projects. It researches unfamiliar systems, edits across files, runs tests, and verifies the result.

In-depth knowledge work

Beyond code, Grok 4.6 works across documents, spreadsheets, PDFs, and other professional artifacts. Give it a multi-step project that requires research, analysis, and synthesis.

Interactive and visual work

Grok 4.6 turns broad product ideas into working first versions. It can establish an application's structure and visual language, build the core interactions, and refine the result through feedback.

Built for more than software engineering

Grok 4.6 builds on Grok 4.5 with a focus on long-running agents and ambitious interactive and visual work. It handles projects that span research, analysis, implementation, and several rounds of refinement.

It matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, a composite of nine benchmarks. Cursor subscription plans for individuals and teams include significant usage of the model.

A strong foundation

We trained Grok 4.6 jointly with SpaceXAI. A longer supplemental training run used curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and training recipe.

We then used Grok 4.5 to regenerate supervised fine-tuning trajectories across reasoning efforts, agent harnesses, STEM, software engineering, and knowledge work. Reinforcement learning covered general coding, knowledge work, kernel optimization, web development, computer-aided design, and other agentic environments.

Benchmarks

Grok 4.6 HighGrok 4.5 HighGPT-5.6 Sol MaxFable 5 Max
AA Intelligence Index61566162
GDPVal-AA v21753152617281741
CursorBench v3.269.9%66.7%67.2%70.5%
DeepSWE v1.165.9%54%73%70%
FrontierCode v1.1Extended61.3%56.6%60.6%64.9%
APEX-Agents57.5%47.1%56.7%59.2%
Terminal-Bench v3.026%15.7%34.6%34.1%
APEX-SWE56.4%53.6%58.8%
AA-Briefcase1577131315021574
Harvey LABVals15.8%12.9%2.5%11.3%

Third-party model scores are the best self-reported or publicly available results.

(Above) Grok 4.6 results across agentic coding and knowledge work benchmarks.

Available everywhere you work

Grok 4.6 is available today in Cursor across desktop, web, iOS, CLI, and our SDK.

Desktop

Manual to agentic coding, in one familiar editor.

CLI

Run agents in any terminal, script, or editor.

Interactive demo with multiple windows showing Cursor's AI-powered features. The interface is displayed over a subtle, solid brand background.

Web & Mobile

Spawn cloud agents on the go from your browser or phone.

Set up Datadog APM for all API routes
Explored 12 files, 3 searches
I'll add the Datadog SDK, instrument all API routes with tracing, and set up error tracking.
Worked for 7m 42s
Processed screen recording
Done — here's the Datadog dashboard showing all instrumented routes.
Acme Labs
Summary
Integrated Datadog APM across all 24 API routes with error tracking and custom metrics.
Add a follow up...

Other Surfaces

Start agents from Slack, GitHub, Linear, JetBrains IDEs, and more.

This element contains an interactive demo for sighted users. It's a demonstration of Cursor integrated within Slack, showing AI-powered assistance inside team communication. The interface is displayed over a subtle, solid brand background.

FAQ

Grok 4.6 is a frontier model from Cursor and SpaceXAI for complex coding and knowledge work. It builds on Grok 4.5 with improved instruction following, long-horizon agentic work, and stronger first passes on interactive and visual projects.

Grok 4.6 is available today in Cursor across desktop, web, iOS, CLI, and our SDK. You can also build your own agents on top of Grok 4.6 with our SDK docs.

Grok 4.6 is part of the Cursor Models pool on individual and team plans, alongside Grok 4.5 and Composer 2.5. Standard on-demand usage is $2/M input, $0.50/M cached input, and $6/M output tokens. The Fast variant is $4/M input, $1/M cached input, and $12/M output tokens. We're offering 2x included usage for the first week. See the model docs for full details.

Try Grok 4.6 now.