xAI released Grok 4.7 with improved coding capabilities and 2.1 trillion parameters at unchanged pricing of $2/$6 per million tokens. But independent testing reveals the AI model burns 81,000 output tokens per task versus 27,000 for GPT-6 Astra, threatening cost efficiency despite lower sticker prices.

News article

xAI Launches Grok 4.7 With Improved Coding Capabilities

SpaceXAI released Grok 4.7, its latest AI model focused on coding and professional knowledge work, maintaining the same pricing as its predecessor at $2 per million input tokens and $6 per million output tokens

1

. The Grok 4.7 release comes after multiple delays, with Elon Musk previously suggesting timelines ranging from "four weeks out" to "needs a few more days to cook" before the Monday afternoon launch

2

. The model is immediately available through Grok Build, the Grok API, Cursor, and GitHub Copilot across multiple development environments including VS Code, Visual Studio, JetBrains, Xcode and Eclipse

1

.

The new model packs 2.1 trillion parameters, representing a 40% increase from Grok 4.6's 1.5 trillion parameters

2

. SpaceXAI incorporated supplemental training data from SpaceX engineering operations, including Starlink satellite telemetry, manufacturing records, and engineering failure logs, positioning the model to reason better about hardware and physical systems

2

. The company describes Grok 4.7 as capable of working longer on difficult tasks and verifying its work more carefully, having undergone a longer reinforcement-learning run using a harder distribution of tasks weighted toward problems requiring many hours to complete

1

.

Benchmark Scores Trail GPT-6 Astra and Claude Opus 5

Despite improvements over Grok 4.6, the AI model continues to lag behind frontier competitors from OpenAI, Anthropic, and Google on key benchmark scores

1

. On Terminal-Bench, where Grok 4.7 saw its largest improvement, Artificial Analysis measured the model at roughly 26% at its xHigh reasoning effort setting, compared to 59.6% for OpenAI's GPT-6 Astra xHigh and approximately 49% for Anthropic's Claude Opus 5 at max effort

1

. Developer Dan McAteer characterized the Terminal-Bench performance as "horrendous" despite the gains over its predecessor

1

.

On CursorBench 4.0, Grok 4.7 at xHigh effort achieved 46.3%, up from 40.4% for Grok 4.6 High, but still trailing Fable 5.1 Max at 51.8%

1

. The model reached 71.0% on DeepSWE v1.1, compared with 65.2% for Grok 4.6

1

. Grok 4.7 scored 1695 on GDPval, which measures performance on economically valuable knowledge work, while Claude Fable 5.1 topped the chart at 1735

2

. On the Artificial Analysis Intelligence Index composite benchmark measuring reasoning, coding, knowledge, and multi-step agentic task completion, Grok 4.7 achieved a score of 46, trailing Claude Fable 5.1 Max, Claude Opus 5, GPT-6 Astra, GPT-5.6 Sol Max, and Meta's Muse Spark 1.3 Max

3

.

High Token Consumption Threatens Real-World ROI

While Grok 4.7 maintains attractive pricing, independent testing reveals significant token consumption that undermines its cost-effective alternative positioning. Artificial Analysis found Grok 4.7 xHigh used approximately 81,000 output tokens per Intelligence Index task, versus 36,000 for Grok 4.6 High and 27,000 for GPT-6 Astra Max

1

. This represents 125% more output token consumption than Grok 4.6 and 196% more than GPT-6 Astra

1

. The high token consumption means that despite lower per-token rates, the model might cost more on actual tasks, threatening real-world ROI for engineering teams routing agentic workloads

3

.

For teams evaluating cost efficiency, the dollar-denominated cost-per-task metric becomes more relevant than sticker pricing

1

. SpaceXAI also offers a Fast variant at 2X the price ($4/$12 per million tokens), though throughput measurements were not available at publication time

1

. Cursor charges $4/$12 for standard Grok 4.7 and $6/$18 for Fast tier on inputs above 256K tokens, up to the model's 500K maximum context window

1

.

xAI's Strategy and the Widening AI Frontier Gap

xAI has consistently positioned itself as offering affordability over frontier performance, betting that "good enough, cheap, and everywhere" beats "best, but pricier" for everyday use

2

. Musk set expectations lower before launch, writing that Grok 4.7 should land "roughly on par with" Anthropic's Claude Opus 5.0, not the newer Opus 5.1

2

. He outlined a roadmap with Grok 4.8 as a meaningful step up, Grok 4.9 in "Astra/Fable class," and Grok 5 as a possible frontier leader, though these upgrades lack release dates

2

. If SpaceXAI needs at least two iterations before matching Anthropic's Fable and OpenAI's Astra-class models, Grok might never fully catch up to true frontier-level capabilities

3

.

Meanwhile, DeepSeek is training a model with 2 trillion parameters and plans to develop an 8-trillion-parameter model soon after, according to The Information

3

. DeepSeek placed a massive order of 160,000 Ascend 950DT chips with Huawei for its 1GW data center, potentially totaling approximately $2.56 billion

3

. Remarkably, DeepSeek's latest models remain competitive with Grok 4.7 despite lacking SpaceXAI's gigawatts of compute infrastructure

3

. Watch for how xAI addresses the growing performance gap while maintaining its pricing advantage, and whether upcoming iterations can deliver the promised Astra-class capabilities before competitors advance further.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved