
Ficha técnica de Grok 4.5
La ficha técnica oficial de Grok 4.5. Documenta los benchmarks de capacidades del modelo y sus evaluaciones de seguridad.
Our most intelligent model for long-running agentic work across coding and knowledge work.

La ficha técnica oficial de Grok 4.5. Documenta los benchmarks de capacidades del modelo y sus evaluaciones de seguridad.

Nuestro modelo más inteligente y el primero que hemos creado para mucho más que la ingeniería de software.

Cursor se asocia con SpaceX para acelerar el entrenamiento de nuestros modelos.
Grok 4.5 runs hard problems for hours: investigating, using tools creatively, recovering from dead ends, and verifying results. Built for work that needs sustained reasoning and multiple tool calls before a conclusion is reached.
Grok 4.5 is strong at large migrations, multi-service systems, and ultra-long-horizon engineering projects. It reads the codebase, edits across files, runs tests, and keeps going until the patch holds up.
Beyond code, Grok 4.5 excels at deep research, analysis, and professional deliverables across data science, finance, legal, and other forms of knowledge work. Give it a multi-step project that requires research, analysis, and synthesis and review a polished, final deliverable in turn.
Grok 4.5 runs multi-step research with tools: searching, synthesizing sources, and verifying its claims for absolute honesty and accuracy. It's designed for a measured, honest approach with the lowest single-turn hallucination rate on the work it performs.
Grok 4.5 handles difficult, long-running tasks that require creatively using tools to solve problems, whether in software engineering, data science, finance, legal work, or anything else you do on a computer.
It solves multistep tasks in under half the steps of comparable frontier models. Cursor subscription plans for individuals and teams include significant usage of the model, with double usage for the first week.
Grok 4.5 is a mixture-of-experts model that we trained jointly with SpaceXAI. Training included trillions of tokens of Cursor data that capture real developer-agent interactions, so the model learns both from existing software and from how agents work inside a codebase.
The training mix is broader than Composer 2.5. It draws on high-quality STEM tasks, research papers, and other knowledge work, then adds reinforcement learning on difficult, realistic problems that teach the model to investigate, recover from mistakes, and verify results.
| Grok 4.5 | Opus 4.8 | GPT-5.5 | Composer 2.5 | Fable 5 | |
|---|---|---|---|---|---|
| Terminal-Bench 2.1 | 83.3% | 78.9% | 83.4% | 73.0% | 84.3% |
| SWE-Bench Multilingual | 78.0% | 84.4% | 77.8% | 71.6% | — |
| DeepSWE 1.0Artificial Analysis | 62.0% high | 55.8% max | 64.3% xhigh | 18.0% | 66.1% max |
| SWE-Bench Pro | 64.7% high | 69.2% max | 58.6% xhigh | 54.0% | 80.3% |
SWE-Bench Pro and Terminal-Bench show self-reported scores for third-party models. For SWE-Bench multilingual, the GPT-5.5 score comes from our internal run.
(Above) Grok 4.5 results across SWE-Bench Pro, Terminal-Bench, SWE-Bench multilingual, and CursorBench.
Grok 4.5 is available today in Cursor across desktop, web, iOS, CLI, and our SDK.

Our most intelligent model for long-running agentic work across coding and knowledge work.

La ficha técnica oficial de Grok 4.5. Documenta los benchmarks de capacidades del modelo y sus evaluaciones de seguridad.

Nuestro modelo más inteligente y el primero que hemos creado para mucho más que la ingeniería de software.

Cursor se asocia con SpaceX para acelerar el entrenamiento de nuestros modelos.
Grok 4.5 runs hard problems for hours: investigating, using tools creatively, recovering from dead ends, and verifying results. Built for work that needs sustained reasoning and multiple tool calls before a conclusion is reached.
Grok 4.5 is strong at large migrations, multi-service systems, and ultra-long-horizon engineering projects. It reads the codebase, edits across files, runs tests, and keeps going until the patch holds up.
Beyond code, Grok 4.5 excels at deep research, analysis, and professional deliverables across data science, finance, legal, and other forms of knowledge work. Give it a multi-step project that requires research, analysis, and synthesis and review a polished, final deliverable in turn.
Grok 4.5 runs multi-step research with tools: searching, synthesizing sources, and verifying its claims for absolute honesty and accuracy. It's designed for a measured, honest approach with the lowest single-turn hallucination rate on the work it performs.
Grok 4.5 handles difficult, long-running tasks that require creatively using tools to solve problems, whether in software engineering, data science, finance, legal work, or anything else you do on a computer.
It solves multistep tasks in under half the steps of comparable frontier models. Cursor subscription plans for individuals and teams include significant usage of the model, with double usage for the first week.
Grok 4.5 is a mixture-of-experts model that we trained jointly with SpaceXAI. Training included trillions of tokens of Cursor data that capture real developer-agent interactions, so the model learns both from existing software and from how agents work inside a codebase.
The training mix is broader than Composer 2.5. It draws on high-quality STEM tasks, research papers, and other knowledge work, then adds reinforcement learning on difficult, realistic problems that teach the model to investigate, recover from mistakes, and verify results.
| Grok 4.5 | Opus 4.8 | GPT-5.5 | Composer 2.5 | Fable 5 | |
|---|---|---|---|---|---|
| Terminal-Bench 2.1 | 83.3% | 78.9% | 83.4% | 73.0% | 84.3% |
| SWE-Bench Multilingual | 78.0% | 84.4% | 77.8% | 71.6% | — |
| DeepSWE 1.0Artificial Analysis | 62.0% high | 55.8% max | 64.3% xhigh | 18.0% | 66.1% max |
| SWE-Bench Pro | 64.7% high | 69.2% max | 58.6% xhigh | 54.0% | 80.3% |
SWE-Bench Pro and Terminal-Bench show self-reported scores for third-party models. For SWE-Bench multilingual, the GPT-5.5 score comes from our internal run.
(Above) Grok 4.5 results across SWE-Bench Pro, Terminal-Bench, SWE-Bench multilingual, and CursorBench.
Grok 4.5 is available today in Cursor across desktop, web, iOS, CLI, and our SDK.

(Above) Grok 4.5 reaches strong CursorBench 3.2 scores with fewer steps per task than other frontier models.
| Model | Score | CostCost / task | TokensTokens / task | StepsSteps / task |
|---|---|---|---|---|
| Grok 4.5 High* | 66.7% | $1.51 | 19,521 | 33 |
(Above) Grok 4.5 reaches strong CursorBench 3.2 scores with fewer steps per task than other frontier models.
| Model | Score | CostCost / task | TokensTokens / task | StepsSteps / task |
|---|---|---|---|---|
| Grok 4.5 High* | 66.7% | $1.51 | 19,521 | 33 |