
Grok 4.5 模型卡
Grok 4.5 的官方模型卡,記錄了該模型的能力基準測試與安全評估。
Our most intelligent model for long-running agentic work across coding and knowledge work.

Grok 4.5 的官方模型卡,記錄了該模型的能力基準測試與安全評估。

這是我們最聰明的模型,也是首度打造、適用於軟體工程以外領域的模型。

Cursor 正與 SpaceX 合作,加速我們的模型訓練進程。
Grok 4.5 runs hard problems for hours: investigating, using tools creatively, recovering from dead ends, and verifying results. Built for work that needs sustained reasoning and multiple tool calls before a conclusion is reached.
Grok 4.5 is strong at large migrations, multi-service systems, and ultra-long-horizon engineering projects. It reads the codebase, edits across files, runs tests, and keeps going until the patch holds up.
Beyond code, Grok 4.5 excels at deep research, analysis, and professional deliverables across data science, finance, legal, and other forms of knowledge work. Give it a multi-step project that requires research, analysis, and synthesis and review a polished, final deliverable in turn.
Grok 4.5 runs multi-step research with tools: searching, synthesizing sources, and verifying its claims for absolute honesty and accuracy. It's designed for a measured, honest approach with the lowest single-turn hallucination rate on the work it performs.
Grok 4.5 handles difficult, long-running tasks that require creatively using tools to solve problems, whether in software engineering, data science, finance, legal work, or anything else you do on a computer.
It solves multistep tasks in under half the steps of comparable frontier models. Cursor subscription plans for individuals and teams include significant usage of the model, with double usage for the first week.
Grok 4.5 is a mixture-of-experts model that we trained jointly with SpaceXAI. Training included trillions of tokens of Cursor data that capture real developer-agent interactions, so the model learns both from existing software and from how agents work inside a codebase.
The training mix is broader than Composer 2.5. It draws on high-quality STEM tasks, research papers, and other knowledge work, then adds reinforcement learning on difficult, realistic problems that teach the model to investigate, recover from mistakes, and verify results.
| Grok 4.5 | Opus 4.8 | GPT-5.5 | Composer 2.5 | Fable 5 | |
|---|---|---|---|---|---|
| Terminal-Bench 2.1 | 83.3% | 78.9% | 83.4% | 73.0% | 84.3% |
| SWE-Bench Multilingual | 78.0% | 84.4% | 77.8% | 71.6% | — |
| DeepSWE 1.0Artificial Analysis | 62.0% high | 55.8% max | 64.3% xhigh | 18.0% | 66.1% max |
| SWE-Bench Pro | 64.7% high | 69.2% max | 58.6% xhigh | 54.0% | 80.3% |
SWE-Bench Pro and Terminal-Bench show self-reported scores for third-party models. For SWE-Bench multilingual, the GPT-5.5 score comes from our internal run.
(Above) Grok 4.5 results across SWE-Bench Pro, Terminal-Bench, SWE-Bench multilingual, and CursorBench.
Grok 4.5 is available today in Cursor across desktop, web, iOS, CLI, and our SDK.

Our most intelligent model for long-running agentic work across coding and knowledge work.

Grok 4.5 的官方模型卡,記錄了該模型的能力基準測試與安全評估。

這是我們最聰明的模型,也是首度打造、適用於軟體工程以外領域的模型。

Cursor 正與 SpaceX 合作,加速我們的模型訓練進程。
Grok 4.5 runs hard problems for hours: investigating, using tools creatively, recovering from dead ends, and verifying results. Built for work that needs sustained reasoning and multiple tool calls before a conclusion is reached.
Grok 4.5 is strong at large migrations, multi-service systems, and ultra-long-horizon engineering projects. It reads the codebase, edits across files, runs tests, and keeps going until the patch holds up.
Beyond code, Grok 4.5 excels at deep research, analysis, and professional deliverables across data science, finance, legal, and other forms of knowledge work. Give it a multi-step project that requires research, analysis, and synthesis and review a polished, final deliverable in turn.
Grok 4.5 runs multi-step research with tools: searching, synthesizing sources, and verifying its claims for absolute honesty and accuracy. It's designed for a measured, honest approach with the lowest single-turn hallucination rate on the work it performs.
Grok 4.5 handles difficult, long-running tasks that require creatively using tools to solve problems, whether in software engineering, data science, finance, legal work, or anything else you do on a computer.
It solves multistep tasks in under half the steps of comparable frontier models. Cursor subscription plans for individuals and teams include significant usage of the model, with double usage for the first week.
Grok 4.5 is a mixture-of-experts model that we trained jointly with SpaceXAI. Training included trillions of tokens of Cursor data that capture real developer-agent interactions, so the model learns both from existing software and from how agents work inside a codebase.
The training mix is broader than Composer 2.5. It draws on high-quality STEM tasks, research papers, and other knowledge work, then adds reinforcement learning on difficult, realistic problems that teach the model to investigate, recover from mistakes, and verify results.
| Grok 4.5 | Opus 4.8 | GPT-5.5 | Composer 2.5 | Fable 5 | |
|---|---|---|---|---|---|
| Terminal-Bench 2.1 | 83.3% | 78.9% | 83.4% | 73.0% | 84.3% |
| SWE-Bench Multilingual | 78.0% | 84.4% | 77.8% | 71.6% | — |
| DeepSWE 1.0Artificial Analysis | 62.0% high | 55.8% max | 64.3% xhigh | 18.0% | 66.1% max |
| SWE-Bench Pro | 64.7% high | 69.2% max | 58.6% xhigh | 54.0% | 80.3% |
SWE-Bench Pro and Terminal-Bench show self-reported scores for third-party models. For SWE-Bench multilingual, the GPT-5.5 score comes from our internal run.
(Above) Grok 4.5 results across SWE-Bench Pro, Terminal-Bench, SWE-Bench multilingual, and CursorBench.
Grok 4.5 is available today in Cursor across desktop, web, iOS, CLI, and our SDK.

(Above) Grok 4.5 reaches strong CursorBench 3.2 scores with fewer steps per task than other frontier models.
| Model | Score | CostCost / task | TokensTokens / task | StepsSteps / task |
|---|---|---|---|---|
| Grok 4.5 High* | 66.7% | $1.51 | 19,521 | 33 |
(Above) Grok 4.5 reaches strong CursorBench 3.2 scores with fewer steps per task than other frontier models.
| Model | Score | CostCost / task | TokensTokens / task | StepsSteps / task |
|---|---|---|---|---|
| Grok 4.5 High* | 66.7% | $1.51 | 19,521 | 33 |