saved
Introducing Grok 4.7
Hraness cites a source capture. The source author remains the source.
gist
xAI launched Grok 4.7 as its strongest coding and knowledge-work model, served at Grok 4.6 price and speed with a larger base, longer RL on multi-hour tasks, stronger self-checks, and native Grok Bot harness training. On CursorBench 4.0 it claims the price-performance frontier versus Fable 5.1, Opus 5, GPT-5.6 Sol, and Sonnet 5. API rates start at $2/$6 per million tokens; a fast variant doubles speed and price. Availability spans Cursor, Grok Build, the Grok API, third-party harnesses, and cloud routers, with a new refusal and jailbreak stack.
ideas
- Same bill of materials as 4.6, bigger training. Same $2/$6 rates and speed as Grok 4.6, but a larger base, longer RL on hard multi-hour tasks, better self-verification and long context, plus native Grok Bot harness training for conversational and knowledge work.
- CursorBench puts it on the price-performance frontier. Versus Fable 5.1, Opus 5, GPT-5.6 Sol, and Sonnet 5 on CursorBench 4.0 average cost per task. Table highs include CursorBench 46.3%, DeepSWE 71.0% (high effort), EEBench 64.0%, AA Briefcase 1,657, Terminal-Bench 4.0 38.0%, Harvey Legal 19.6%, HealthBench Professional 56.7%.
- Documents and presentations move with GDPval. Grok 4.7 xhigh scores 1695 Elo on GDPval versus Fable 5.1 max 1735, Grok 4.6 high 1605, and GPT-6 Astra max 1542; AA Briefcase and GDPval improve on 4.6 and land near the compared frontier set.
- New safeguard stack leads dual-use refusals. Strongest xAI refusals and jailbreak resistance yet; LatchBio biosafety 62.4%; HackerBench v0.3 lets only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work; select cyber partners get invite-only red-team access.
- Ship today; fast costs 2x. Live in Cursor, Grok Build (free trial), the Grok API, third-party coding harnesses, and model routers/cloud platforms. Fast variant: twice output speed at twice the price.
quotes
“Grok 4.7 is our most capable model for coding and knowledge work.”
“On CursorBench 4.0, which stresses longer-running coding tasks, Grok 4.7 is at the frontier in price-performance.”
“The model is priced starting at $2 per million input tokens and $6 per million output tokens.”
“It is the strongest model we’ve tested on refusals and jailbreak resistance.”