GPT-6 Astra: OpenAI's AI Model Excels at Spatial Tasks But Struggles With Writing

Reviewed byNidhi Govil

4 Sources

Share

OpenAI released GPT-6 Astra in September 2026, marking a shift from answering questions to executing complex tasks autonomously. The AI model achieves a 72.6% success rate on desktop automation benchmarks and rebuilds Manhattan in Unreal Engine. But independent testing reveals it writes worse than its predecessor while carrying OpenAI's first critical cybersecurity risk rating.

OpenAI Shifts Focus From Intelligence to Execution

OpenAI released GPT-6 Astra in September 2026, introducing an AI model designed to perform tasks autonomously rather than simply provide answers

4

. Within 48 hours of early access, developers turned the launch into a public stress test that revealed a clear pattern: GPT-6 Astra dominates spatial reasoning, 3D modeling, and agentic tasks while delivering weaker writing performance than GPT-5.6 Sol, its predecessor

1

. The advanced AI model costs $10 per million input tokens and $50 per million output tokens, representing a 2.5x increase over Sol's pricing

1

. OpenAI president Greg Brockman used the launch briefing to announce the arrival of AGI

1

.

Independent testing from Artificial Analysis found GPT-6 Astra's overall intelligence score nearly identical to GPT-5.6 Sol and trailing some rival models on general reasoning

4

. The same testers rated its writing below its predecessor and measured a drop of roughly 80 Elo points on a benchmark of economically valuable professional work

1

. This performance split highlights how OpenAI built GPT-6 Astra to execute multi-step actions rather than think harder.

Desktop Automation Benchmark Shows Major Gains

The headline feature is computer use, which allows the AI model to drive a mouse and keyboard on a real desktop instead of returning instruction lists

1

. On OSWorld 2.0, a desktop automation benchmark that scores what percentage of ordinary desktop chores an agent finishes independently, OpenAI reported 72.6% completion at roughly 40 minutes per task, compared to 65.7% at 75 minutes for Sol

1

. Independent write-ups from DataCamp and MindStudio confirmed similar results

4

.

On ScreenSpot-Pro, a test measuring how well models locate and click screen elements without external help, both OpenAI and outside reviewers placed GPT-6 Astra in the low nineties, well past Sol's high seventies

4

. These scores demonstrate the model's ability to perform tasks autonomously across extended sessions without losing track of goals.

Visual Understanding Drives 3D Modeling Breakthroughs

GPT-6 Astra demonstrates exceptional visual understanding and spatial awareness across real-world applications

1

. Matt Shumer, former CEO of HyperWrite, gave Astra a week inside Unreal Engine and watched it generate a replica of Manhattan, working street by street to perfect each one

1

. Max Weinbach fed the model photographs of Apple Park and received a reconstruction in Blender that he described as doing "an absurd job"

1

.

Tom Krcha handed GPT-6 Astra a single image of a house and received the full interior as editable geometry running at 60 frames per second, complete with appliances and toys

1

. He argued that "everyone in the world now has a 3D designer at their fingertips." A developer posting as SuSu took the concept to city scale, watching Astra rebuild the Chinese city of Hangzhou and surrounding towns in Three.js in 24 minutes, with West Lake, Leifeng Pagoda, tea terraces, and wetlands all in place

1

.

Game Development Shows Fastest Creative Output

Game development emerged as the most popular use case where GPT-6 Astra excels

1

.

Source: Geeky Gadgets

Source: Geeky Gadgets

The AI model created a playable first-person shooter with detailed environments and mechanics in just 30 minutes, a task typically demanding extensive time and resources

3

. Anshu Chimala, former UX/UI designer at Apple, generated a 3D game in one shot in 45 minutes for barely a couple percent of his usage quota, calling Astra "some kind of turbo-AGI machine god for 3D games"

1

.

Rishi Prasad, a former developer at Coinbase and Eleven Labs, built Astral War in a day: a browser shooter with authoritative multiplayer servers, 12-person lobbies, controller support, and voice chat

1

. He described "a huge, step-function leap in visual fidelity" over what he built a month earlier with Claude Opus 5. By automating labor-intensive processes such as coding, level design, and asset creation, GPT-6 Astra allows developers to focus on refining gameplay and enhancing user experiences

3

.

Robotics Applications Achieve 95% Success Rate

In robotics, GPT-6 Astra achieved a 95% success rate in managing robotic arms for complex tasks, outperforming previous AI models

3

. This precision proves critical for applications ranging from industrial automation to autonomous systems. For engineers and researchers, the advanced algorithms reduce error rates and minimize setbacks in high-stakes environments such as manufacturing and healthcare, where even minor errors carry significant consequences

3

.

Head-to-Head Testing Against Fable 5.1

After 100 hours of testing across 15 distinct scenarios, GPT-6 Astra performed well in 10 out of 15 use cases when compared against Fable 5.1

2

.

Source: Geeky Gadgets

Source: Geeky Gadgets

Astra demonstrated particular strength in tasks requiring structured outputs, browser automation, and cost efficiency, with total costs of $326.98 compared to Fable 5.1's $513.36

2

. However, Fable 5.1 completed tasks faster with a total runtime of 9 hours and 35 minutes versus GPT-6 Astra's 11 hours and 19 minutes

2

.

Fable 5.1 excelled in creative domains such as web design and game development, producing visually refined results

2

. Its website cloning capabilities consistently produced visually accurate replicas with fewer errors. GPT-6 Astra's slower pace stemmed from its tendency to ask clarifying questions, which ensured tailored and accurate outputs but added to overall runtime

2

.

Token Cost Economics Favor Long Sessions

On common coding benchmarks like DeepSWE and Frontier Code, MindStudio found GPT-6 Astra running close to even with Sol and rival models

4

. Where Astra pulls ahead is Terminal-Bench 4.0, a test built around long terminal sessions requiring chained commands and error recovery. Artificial Analysis found GPT-6 Astra matching a top rival coding model on its Coding Agent Index while running under half the token cost per finished task

4

.

The upgrade delivers lower token cost and faster completion across sessions stretching for hours rather than sharper code from single prompts

4

. For short, simple requests, the 2.5x premium rarely pays off. For long agentic tasks involving hundreds of tool calls, lower token use can offset or beat the higher price on a per-task basis

4

.

First Critical Cybersecurity Risk Rating

OpenAI classified GPT-6 Astra at a new critical risk level for cybersecurity, marking the first time OpenAI rated a model at this threshold

1

. The AI model can find unknown software flaws and build working attacks without a human pointing at the hole first

1

. New safeguards followed, including tighter review of high-risk actions inside ChatGPT and Codex

4

. This cybersecurity risk rating adds fresh weight to how companies manage access and deploy the model in production environments.

Professional Work Moves Toward Full Automation

On Agents' Last Exam, which spans financial modeling, engineering, and media production inside real software, GPT-6 Astra edged past both Sol and a leading rival reasoning model while using notably fewer tokens

4

. The shift means GPT-6 Astra can move between research, calculation, software interaction, and final output with less manual handoff between stages. This capability turns AI from a drafting assistant into something closer to a task owner, automating labor-intensive processes that previously required constant human oversight

4

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved