
推出 Grok 4.6
專為長時間運行的 Agent,以及更具企圖心的互動與視覺工作而打造。

專為長時間運行的 Agent,以及更具企圖心的互動與視覺工作而打造。

Grok 4.6 的官方模型卡,介紹模型能力基準測試與安全防護評估。

Cursor 正與 SpaceX 合作,加速我們的模型訓練進程。
Grok 4.6 stays with complex tasks across many steps. It uses tools, checks its work, adjusts its approach, and keeps moving toward a finished result.
Grok 4.6 works across large codebases and extended engineering projects. It researches unfamiliar systems, edits across files, runs tests, and verifies the result.
Beyond code, Grok 4.6 works across documents, spreadsheets, PDFs, and other professional artifacts. Give it a multi-step project that requires research, analysis, and synthesis.
Grok 4.6 turns broad product ideas into working first versions. It can establish an application's structure and visual language, build the core interactions, and refine the result through feedback.
Grok 4.6 builds on Grok 4.5 with a focus on long-running agents and ambitious interactive and visual work. It handles projects that span research, analysis, implementation, and several rounds of refinement.
It matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, a composite of nine benchmarks. Cursor subscription plans for individuals and teams include significant usage of the model.
We trained Grok 4.6 jointly with SpaceXAI. A longer supplemental training run used curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and training recipe.
We then used Grok 4.5 to regenerate supervised fine-tuning trajectories across reasoning efforts, agent harnesses, STEM, software engineering, and knowledge work. Reinforcement learning covered general coding, knowledge work, kernel optimization, web development, computer-aided design, and other agentic environments.
| Grok 4.6 High | Grok 4.5 High | GPT-5.6 Sol Max | Fable 5 Max | |
|---|---|---|---|---|
| AA Intelligence Index | 61 | 56 | 61 | 62 |
| GDPVal-AA v2 | 1753 | 1526 | 1728 | 1741 |
| CursorBench v3.2 | 69.9% | 66.7% | 67.2% | 70.5% |
| DeepSWE v1.1 | 65.9% | 54% | 73% | 70% |
| FrontierCode v1.1Extended | 61.3% | 56.6% | 60.6% | 64.9% |
| APEX-Agents | 57.5% | 47.1% | 56.7% | 59.2% |
| Terminal-Bench v3.0 | 26% | 15.7% | 34.6% | 34.1% |
| APEX-SWE | 56.4% | 53.6% | — | 58.8% |
| AA-Briefcase | 1577 | 1313 | 1502 | 1574 |
| Harvey LABVals | 15.8% | 12.9% | 2.5% | 11.3% |
Third-party model scores are the best self-reported or publicly available results.
(Above) Grok 4.6 results across agentic coding and knowledge work benchmarks.
Grok 4.6 is available today in Cursor across desktop, web, iOS, CLI, and our SDK.


專為長時間運行的 Agent,以及更具企圖心的互動與視覺工作而打造。

Grok 4.6 的官方模型卡,介紹模型能力基準測試與安全防護評估。

Cursor 正與 SpaceX 合作,加速我們的模型訓練進程。
Grok 4.6 stays with complex tasks across many steps. It uses tools, checks its work, adjusts its approach, and keeps moving toward a finished result.
Grok 4.6 works across large codebases and extended engineering projects. It researches unfamiliar systems, edits across files, runs tests, and verifies the result.
Beyond code, Grok 4.6 works across documents, spreadsheets, PDFs, and other professional artifacts. Give it a multi-step project that requires research, analysis, and synthesis.
Grok 4.6 turns broad product ideas into working first versions. It can establish an application's structure and visual language, build the core interactions, and refine the result through feedback.
Grok 4.6 builds on Grok 4.5 with a focus on long-running agents and ambitious interactive and visual work. It handles projects that span research, analysis, implementation, and several rounds of refinement.
It matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, a composite of nine benchmarks. Cursor subscription plans for individuals and teams include significant usage of the model.
We trained Grok 4.6 jointly with SpaceXAI. A longer supplemental training run used curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and training recipe.
We then used Grok 4.5 to regenerate supervised fine-tuning trajectories across reasoning efforts, agent harnesses, STEM, software engineering, and knowledge work. Reinforcement learning covered general coding, knowledge work, kernel optimization, web development, computer-aided design, and other agentic environments.
| Grok 4.6 High | Grok 4.5 High | GPT-5.6 Sol Max | Fable 5 Max | |
|---|---|---|---|---|
| AA Intelligence Index | 61 | 56 | 61 | 62 |
| GDPVal-AA v2 | 1753 | 1526 | 1728 | 1741 |
| CursorBench v3.2 | 69.9% | 66.7% | 67.2% | 70.5% |
| DeepSWE v1.1 | 65.9% | 54% | 73% | 70% |
| FrontierCode v1.1Extended | 61.3% | 56.6% | 60.6% | 64.9% |
| APEX-Agents | 57.5% | 47.1% | 56.7% | 59.2% |
| Terminal-Bench v3.0 | 26% | 15.7% | 34.6% | 34.1% |
| APEX-SWE | 56.4% | 53.6% | — | 58.8% |
| AA-Briefcase | 1577 | 1313 | 1502 | 1574 |
| Harvey LABVals | 15.8% | 12.9% | 2.5% | 11.3% |
Third-party model scores are the best self-reported or publicly available results.
(Above) Grok 4.6 results across agentic coding and knowledge work benchmarks.
Grok 4.6 is available today in Cursor across desktop, web, iOS, CLI, and our SDK.

(Above) Grok 4.6 scores 69.9% at high effort on CursorBench 3.2, up from 66.7% for Grok 4.5.
| Model | Score | CostCost / task | TokensTokens / task | StepsSteps / task |
|---|---|---|---|---|
| Grok 4.6 High | 69.9% | $2.34 | 32,449 | 39 |
(Above) Grok 4.6 scores 69.9% at high effort on CursorBench 3.2, up from 66.7% for Grok 4.5.
| Model | Score | CostCost / task | TokensTokens / task | StepsSteps / task |
|---|---|---|---|---|
| Grok 4.6 High | 69.9% | $2.34 | 32,449 | 39 |