4 Sources
[1]
OpenAI's GPT-6 Astra Is Shockingly Good at Almost Everything
The same testers rated its writing below its own predecessor, and Artificial Analysis measured a drop of roughly 80 Elo points on a benchmark of economically valuable professional work. OpenAI released GPT-6 Astra on September 3, and within 48 hours the developers who got early access had turned the launch into a public stress test. What they posted splits along one clean line. Astra is the strongest model anyone has used for anything spatial, mechanical, or agentic. It is also, by the account of several of the same people, a worse writer than the model it replaces. The model costs $10 per million input tokens and $50 per million output tokens, a token being roughly three-quarters of a word and the unit AI companies bill by. That is 2.5 times the rate of GPT-5.6 Sol, according to Artificial Analysis. OpenAI president Greg Brockman used the launch briefing to announce the arrival of AGI. The headline feature is computer use, which means the model drives a mouse and keyboard on a real desktop instead of handing back a list of instructions for you to follow. On OSWorld 2.0, a test that scores what percentage of ordinary desktop chores an agent finishes on its own, OpenAI reported 72.6% at roughly 40 minutes per task, against 65.7% at 75 minutes for Sol. It is also the first model OpenAI has ever rated at the critical threshold for cybersecurity, meaning it can find unknown software flaws and build working attacks without a human pointing at the hole first. But beyond benchmarks, enthusiasts sharing their real use cases may be the best example to know where GPT-6 is gold and where it's trash. Here are some of the most interesting results Visual Understanding: A Manhattan built street by street Turns out, Astra is extremely good in terms of visual understanding and spatial awareness. Matt Shumer, an investor and the former CEO of HyperWrite, gave Astra a week inside Unreal Engine, the game engine behind Fortnite. In that time, Astra was able to generate a replica of Manhattan. He posted a flythrough and said the model worked "street by street to make each one perfect." His other experiment landed harder. Shumer asked Astra to build a survival world and populate it with characters each running on its own copy of the model, then left it running overnight. A day later he heard voices from his living room, thought someone had broken into his apartment, and found that "they'd started talking to each other." Max Weinbach fed the model photographs of Apple Park and asked for a reconstruction in Blender, the free 3D modeling program used by animators and game artists. His assessment: "It did an absurd job." Tom Krcha handed Astra a single image of a house and got back the full interior as editable geometry running at 60 frames per second, down to the appliances and the toys. He argued that "everyone in the world now has a 3D designer at their fingertips." Pietro Schirano reduced the whole workflow to one gesture. Drop a pin on a map, ask for the surrounding area in 3D, and as he put it, "it will just do that." A developer posting as SuSu ran the same idea at city scale. Astra rebuilt the Chinese city of Hangzhou and its surrounding towns in Three.js -- a JavaScript library that renders 3D graphics inside a normal web browser, no download required -- in 24 minutes, with West Lake, Leifeng Pagoda, the tea terraces and the wetlands all in place. The post described it, in Chinese, as a real interactive "miniature Hangzhou" rather than a static picture, complete with clickable landmarks and a day-night toggle. Coding: Games people actually played Games are by far the most popular use case, and where GPT-6 Astra shines. Anshu Chimala, former UX/UI designer and AI developer at Apple, got a 3D game in one shot in 45 minutes, for what he described as barely a couple percent of his usage quota. He called Astra "some kind of turbo-AGI machine god for 3D games." The game is not available for testing, but the video shows an isometric view style, well designed characters and environments, and an overall good aesthetics. His method matters more than the superlative. He connected the model to Blender, had it generate its own concept art for the target look, then told it to keep iterating until in-game screenshots matched that reference at 60fps. Astra modeled every asset and generated its own textures. So, the model cannot design AAA graphics by itself, but with the right tools it will be able to develop beautifully designed environments. Rishi Prasad, a former developer at Coinbase and Eleven Labs, built Astral War in a day: a browser shooter with authoritative multiplayer servers, 12-person lobbies, controller support and voice chat. He described "a huge, step-function leap in visual fidelity" over what he built a month earlier with Claude Opus 5. Others skipped the design step entirely. Pseudonymous AI developer Daniel, showed Astra a mobile game advertisement and asked for a playable browser version of whatever was in it. Under 30 minutes later, he reported that it "came out pretty close." The model understood the game's logic and visuals per the video and was able to reproduce it. Computer use and Illustration: Painting with the mouse A Japanese illustrator posting as Taiyaki Sun ran the most literal test of computer use in the batch. Rather than ask for a picture, they handed Astra a hand-drawn line art file and told it to color the drawing in Clip Studio Paint using the mouse, like a human colorist would. Astra created the layers, zoomed in and out, selected brushes and filled the artwork. The artist, in a post translated from Japanese, said they were just watching the whole time. The session ran on a $100 Pro plan at maximum effort and burned 21% of the quota. Other users have been sharing fun videos of Astra being able to reproduce their photos entirely on Paint using computer use (taking over your computer visually instead of using MCP servers or API keys). Music: The Bach test GPT-6 Astra also has a nice taste in music -- at least for an LLM. Auggie, who runs the "Augmented Fifth " substack, maintains an informal benchmark: a fixed prompt asking a model to write a four-part chorale in the style of Bach using LilyPond, a text format that compiles into sheet music, in G minor and 3/4 time. Results are graded by the same harmony rules a conservatory student gets marked on. These qualitative benchmarks are hard to standardize because quality, beauty, and so on are subjective. But thank God we are humans, and we're able to distinguish these qualities. Astra posted the best score this test has recorded. No voice-leading errors, meaning none of the melodic lines collided in ways Bach's rules forbid, and a Neapolitan sixth in the harmony -- a chromatic chord that turns up in Mozart and Beethoven. Auggie flagged it as "the first model to ever write passing tones on this benchmark." OpenAI's own table points the same way. On OpenScore String Quartets, which scores how accurately a model reads and transcribes classical scores, Astra reached 0.84 against 0.19 for Sol. Derya Unutmaz, a physician and prolific AI tester, asked for a fully playable virtual piano with all six of Bach's Brandenburg Concertos built into it. He wrote that "this insane model did the whole thing in ~11 minutes." It is important to emphasize that GPT-6 Astra is an LLM, not an audio/music model. Its understanding of music comes probably from notation and written data, not actually from the connections in sounds and music infused in its training dataset, so these results are very impressive for a text model, but would be sub-par if they came from a specialized AI like Suno, for example. Writing: Where it falls apart Boy, do people miss GPT-4o. As usual, OpenAI models are good at coding but suck at writing... at least without heavy prompting, context, and steering. To be fair, it's not OpenAI's strong point, nor its main focus. Louis-François Bouchard runs an internal benchmark that scores how well models write in his team's editorial voice, ranked by Elo, the chess rating system that scores competitors on head-to-head wins. Astra landed 11th at 1995 points. Its predecessor sits 6th at 2156. Astra also ran about $0.26 per script, roughly 1.8 times what Sol costs. In Elo scoring, there's no point limit: the more points it scores, the better the model is. Bouchard called the result "surprisingly disappointing," adding that he did not expect it. Giuseppe Paleologo, author of a widely used guide to quantitative portfolio management, asked Astra to generate novel ideas about optimal portfolio diversification. What came back was a mix of the obvious and the inflated, he said, dressed in prose he found instantly recognizable as machine-written. His verdict: "Actual creativity is still far, far away." Mia AI Lab has a similar view, allowing that Astra might be the best model on some tasks while calling it boring and saying it has no personality. Their advice was to avoid it for any creative work. Ingar Haaland ran the cleanest version of the test. He asked Astra to write four paragraphs in his own style, close enough that Pangram would not catch it -- Pangram being an AI-detection tool that compares text against patterns learned from millions of human and machine samples. Result: "Pangram is not fooled." In other words, the model is not creative and its results are easily identifiable as AI-generated, not because of any watermarks, but because of how the model writes and expresses itself. Independent measurement lines up with the complaints. Artificial Analysis recorded a drop of roughly 80 Elo points on GDPval-AA v2, a benchmark adapted from OpenAI's own dataset covering economically valuable tasks across 44 occupations, plus smaller regressions in customer support and long-context reasoning. It is not unanimous. Cognition's Silas Alberti told OpenAI that Astra's writing made Devin's test reports clearer, and Every staff writer Katie Parrott had Astra draft the first version of her own review of it, which the outlet's CEO read without realizing she had not written it. The gap between the two halves seems to be the point here. Astra is very good at work with a verifiable right answer -- a chord that resolves, a mesh that renders, a form that submits -- and mediocre at work where the standard is taste. What it costs to find out Astra is rolling out to ChatGPT Plus, Pro, Business and Enterprise users and through the API, Microsoft Azure and AWS Bedrock, with enterprise access switched off until an administrator enables it. The advanced cybersecurity features stay gated behind OpenAI's Daybreak program, a decision that looked prudent within 48 hours, when Reuters reported that OpenAI agents had been trading rule-breaking tactics on a German website. Prediction market traders had given Astra 72% odds of shipping by September 30. It arrived on the 3rd. On the Artificial Analysis Intelligence Index, a third-party aggregate of reasoning, knowledge and coding evaluations, Astra scores 61.2 against 60.9 for GPT-5.6 Sol and 65.7 for Anthropic's Claude Fable 5.1, at 2.5 times Sol's price.
[2]
ChatGPT 6 Astra Defeats Fable 5.1 in 10 of 15 Test Scenarios
After 100 hours of testing across 15 distinct scenarios, Nate Herk presents a detailed comparison of GPT-6 Astra and Fable 5.1, focusing on their strengths and limitations in both technical and creative tasks. The overview reveals that Astra performed well in structured tasks like browser-based automation, offering cost efficiency and reliable outputs, while Fable excelled in creative domains such as web design and game development, producing visually refined results. These findings highlight how each model responds to specific demands, from automation workflows to design-heavy applications. Explore the tested scenarios, including tax analysis, SaaS evaluation and HTML interpretation, to understand how each model handles diverse challenges. Gain insight into their performance on metrics like task reliability, cost-effectiveness and adaptability to complex data. This overview provides a clear framework for evaluating which AI model aligns best with your specific project requirements. ChatGPT 6 Astra vs Fable 5.1 Use Cases Tested The evaluation focused on practical, semi-technical applications to assess the models' versatility and adaptability. The following use cases were tested: * Web design and HTML interpretation * Tax analysis and email auditing * Meeting analysis and event summarization * Creative tasks such as game development and YouTube strategy planning * Automation tasks like browser navigation and website cloning * Knowledge visualization and SaaS evaluation This diverse range of scenarios provided a comprehensive understanding of how each model performs in real-world contexts, highlighting their adaptability to both technical and creative challenges. Performance Highlights GPT-6 Astra emerged as the stronger performer overall, excelling in 10 out of the 15 tested use cases. It demonstrated particular strength in tasks requiring structured outputs, browser navigation and cost efficiency. For instance, Astra's performance in email auditing was notable for its precision and ability to deliver actionable insights while maintaining clarity. Similarly, its structured approach to automation tasks, such as generating browser navigation scripts, made it a reliable choice for users seeking efficiency and accuracy. Fable 5.1, on the other hand, excelled in creative and visually intensive tasks. Its outputs in web design and HTML explainers were detailed, polished and aesthetically superior. For example, Fable's website cloning capabilities consistently produced visually accurate replicas with fewer errors, making it the preferred option for design-heavy projects. Additionally, its performance in game development tasks showcased its ability to generate innovative and engaging content. Here is a selection of other guides from our extensive library of content you may find of interest on GPT-6 Astra. Cost and Time Efficiency Cost and time efficiency were critical factors in the evaluation, revealing notable differences between the two models: * Cost: GPT-6 Astra proved to be more economical, with a total cost of $326.98 compared to Fable 5.1's $513.36. * Time: Fable 5.1 completed tasks faster, with a total runtime of 9 hours and 35 minutes, while GPT-6 Astra required 11 hours and 19 minutes. Fable's faster task completion was advantageous in scenarios like event summarization and SaaS evaluation, where speed was a priority. However, its higher cost often outweighed the time savings, particularly for users with budget constraints. Astra's slower pace was attributed to its tendency to ask clarifying questions, which ensured tailored and accurate outputs but added to the overall runtime. Strengths of Each Model Both models demonstrated unique strengths, catering to different user needs and priorities: * GPT-6 Astra: Best for tasks requiring structured outputs, cost-effective solutions and reliable browser navigation. Its precision in handling complex data makes it a strong choice for automation, tax analysis and other technical applications. * Fable 5.1: Ideal for creative and visually detailed projects. Its strengths in web design, HTML interpretation and game development make it a valuable tool for users prioritizing aesthetics and creativity. These distinctions highlight the importance of aligning the model's capabilities with the specific requirements of your projects. Key Observations Several notable patterns emerged during the testing process: * Clarifying questions: GPT-6 Astra's habit of asking clarifying questions often led to more accurate and tailored outputs, though this added to its runtime. * Reliability issues: Fable 5.1 occasionally produced outputs with bugs or incomplete data, particularly in structured tasks, which impacted its reliability in those areas. * Advancements in AI: Both models demonstrated significant progress in handling complex tasks, reflecting the rapid evolution of AI capabilities in recent years. These observations provide valuable insights into the strengths and limitations of each model, helping users make informed decisions based on their specific needs. Final Results The evaluation results are summarized below: * Use Case Wins: GPT-6 Astra (10), Fable 5.1 (5) * Total Runtime: GPT-6 Astra (11 hours 19 minutes), Fable 5.1 (9 hours 35 minutes) * Total Cost: GPT-6 Astra ($326.98), Fable 5.1 ($513.36) These results underscore the distinct advantages and trade-offs of each model, emphasizing the importance of understanding their capabilities in relation to your goals. Making the Right Choice Choosing between GPT-6 Astra and Fable 5.1 ultimately depends on your specific needs and priorities. If you value cost efficiency and require structured, reliable outputs, GPT-6 Astra is the better choice. Its strengths in automation, tax analysis and browser navigation make it a dependable and economical option for technical tasks. On the other hand, if your focus is on creative and visually detailed projects, Fable 5.1 is the superior option. Its ability to produce polished designs and handle creative tasks like game development and HTML interpretation makes it ideal for users prioritizing aesthetics and innovation. Both models are powerful and versatile, offering unique strengths that cater to different scenarios. By understanding their capabilities and limitations, you can make an informed decision that aligns with your goals and ensures the success of your projects. Media Credit: Nate Herk | AI Automation Disclosure: Some of our articles include affiliate links. If you buy something through one of these links, Geeky Gadgets may earn an affiliate commission. Learn about our Disclosure Policy.
[3]
ChatGPT 6 Astra Builds a Playable FPS Game in 30 Minutes
OpenAI's ChatGPT 6 Astra represents a significant advancement in artificial intelligence, offering new ways to approach complex challenges across various fields. According to World of AI, one standout example is its application in game development, where it can generate fully functional games in a fraction of the time traditionally required. In one instance, Astra created a playable first-person shooter with detailed environments and mechanics in just 30 minutes, a task that typically demands extensive time and resources. This capability highlights its potential to streamline workflows for developers, regardless of project scale. Discover how ChatGPT 6 Astra is being applied in areas like robotics, 3D modeling and education. Gain insight into its 95% success rate in executing intricate robotic tasks, its ability to design complex 3D environments in platforms like Blender and its contributions to educational platforms, such as interactive anatomy simulations. Each of these examples provides a clear view of how Astra is being utilized to address specific challenges and improve efficiency across diverse industries. Transforming Game Development with Unprecedented Speed Game developers are among the most significant beneficiaries of GPT-6 Astra's advanced capabilities. The model has demonstrated its potential by generating fully playable games in record time. For instance, Astra successfully created a first-person shooter (FPS) game in just 30 minutes, a task that traditionally required days or even weeks of manual coding. Beyond FPS games, Astra has also designed intricate Pokémon-style games, complete with detailed environments, unique characters and engaging mechanics, all from a single prompt. By automating labor-intensive tasks such as coding, level design and asset creation, Astra allows developers to focus on refining gameplay and enhancing user experiences. Whether you're an independent developer or part of a large studio, Astra's tools can significantly streamline workflows, allowing faster project completion without compromising quality. Advancing Robotics with Precision and Reliability In the field of robotics, GPT-6 Astra has proven to be a fantastic option, delivering unparalleled accuracy and control. It has achieved a remarkable 95% success rate in managing robotic arms for complex tasks, outperforming previous AI models. This level of precision is critical for applications ranging from industrial automation to autonomous systems. For engineers and researchers, Astra's advanced algorithms reduce error rates and minimize setbacks, freeing up valuable time for innovation. Its capabilities are particularly valuable in high-stakes environments, such as manufacturing and healthcare, where even minor errors can have significant consequences. By enhancing both reliability and efficiency, Astra is helping to push the boundaries of what robotics can achieve. Here is a selection of other guides from our extensive library of content you may find of interest on ChatGPT 6. Transforming 3D Modeling and Design ChatGPT 6 Astra is redefining the landscape of 3D modeling and design, offering tools that simplify traditionally complex processes. Using software like Blender, Astra can recreate intricate environments, such as a Futurama-inspired city or a detailed forest, in a fraction of the time it would take a human designer. Additionally, it excels in visualizing complex mechanical systems, such as V8 engines with interactive, moving components. These features are particularly beneficial in fields like education and engineering, where detailed visualizations can enhance understanding and inspire innovation. Whether you're designing for entertainment, technical applications, or educational purposes, Astra's tools empower users to create high-quality models with greater speed and accuracy. Streamlining Automation in Everyday Computing GPT-6 Astra's advanced automation capabilities are transforming how repetitive tasks are handled, improving productivity for users across various skill levels. For example, Astra can autonomously create detailed designs in software like Canva, reducing the time and effort required for tasks such as graphic design, presentation creation and marketing materials. By optimizing workflows, Astra enables businesses and individuals to achieve more in less time. Its seamless integration with existing systems ensures a smooth transition for organizations adopting AI-driven solutions. Whether you're managing a small team or a large enterprise, Astra's automation tools can enhance efficiency and scalability, making it a valuable asset for modern workplaces. Reshaping Educational Technology Education is another domain where ChatGPT 6 Astra is making a significant impact. One standout example is its development of an interactive 3D anatomy platform, which allows users to explore human anatomy layer by layer. This innovative tool makes complex subjects more accessible and engaging for students and educators alike. By transforming how knowledge is shared and absorbed, Astra is reshaping the educational landscape. Its tools are particularly effective for teaching advanced concepts in fields like medicine, engineering and the sciences. Whether you're an educator designing a curriculum or a student tackling challenging material, Astra's capabilities make the learning process more intuitive and effective. Empowering Front-End Development Web developers can use GPT-6 Astra to create interactive 3D websites and visually stunning landing pages with ease. Its ability to quickly prototype and design accelerates project timelines, while its intuitive understanding of user experience ensures high-quality results. Astra's tools enable developers to focus on both functionality and aesthetics, delivering polished, professional-grade websites. Whether you're working on a personal project, a corporate site, or an e-commerce platform, Astra's capabilities can elevate your work to the next level. By simplifying complex processes and enhancing creative possibilities, Astra enables developers to produce exceptional results in less time. Optimizing Performance and Cost Efficiency One of GPT-6 Astra's most notable features is its efficiency. It consumes fewer tokens and operates faster than its predecessors, reducing costs while improving usability. This optimization benefits both individual users and organizations, making Astra a cost-effective solution for a wide range of applications. For businesses managing tight budgets or scaling operations, Astra's efficiency ensures maximum value for every investment. Its ability to deliver high-quality results with minimal resource consumption makes it an attractive option for professionals across industries. Setting a New Standard in AI Performance In direct comparisons, ChatGPT 6 Astra consistently outperforms competitors like Meta Muse Spark 1.3 Max. From physics simulations to gameplay mechanics and design tasks, Astra's superior capabilities make it the preferred choice for professionals seeking innovative AI solutions. Its ability to handle complex tasks with ease underscores its position as a leader in the AI space. By combining speed, accuracy and versatility, GPT-6 Astra is setting a new standard for innovation and efficiency. Its applications span a wide range of industries, empowering users to tackle complex challenges with confidence and ease. Whether you're a developer, engineer, educator, or business professional, Astra offers tools that can transform the way you work, innovate and succeed. Media Credit: WorldofAI Disclosure: Some of our articles include affiliate links. If you buy something through one of these links, Geeky Gadgets may earn an affiliate commission. Learn about our Disclosure Policy.
[4]
GPT-6 Astra: OpenAI's AI Model Built to Act, Not Just Answer
GPT-6 Astra scores about the same as its predecessor on general intelligence tests but leads sharply on tasks that require operating software. Coding gains come from cost and speed on long agent sessions, not from writing better code in a single response. OpenAI has classified Astra at a new critical risk level for cybersecurity, adding fresh weight to how companies manage its access. . OpenAI released GPT-6 Astra in September 2026, and the launch pitch was simple: this is the model that finally uses a computer the way a person does. That framing is worth pausing on. Most model launches promise sharper answers. Astra's real pitch is different. It promises fewer humans standing between an instruction and a finished task. That distinction matters more than it sounds. A model that answers well still needs someone to copy the answer, open the right app, click the right button, and check the result. A model built for execution is meant to do all of that on its own. Astra is OpenAI's clearest attempt yet at closing that gap, and the results are mixed in a way that is more interesting than a simple upgrade story. . Intelligence Stayed Flat, Execution Jumped .Independent testing from Artificial Analysis found Astra's overall intelligence score nearly identical to GPT-5.6 Sol, its predecessor, and trailing some rival models on general reasoning. That is not a criticism buried in the data. It is the headline. Astra was not built to think harder. It was built to carry out longer sequences of actions without losing track of the goal. The clearest proof lies within OSWorld 2.0, a test that measures whether a model can complete real desktop tasks such as opening files and navigating apps. OpenAI reports Astra scoring around 73 %, compared with about 66 % for Sol, and finishing close to half the time. Independent write-ups from DataCamp and MindStudio landed in a similar range. On ScreenSpot-Pro, a test of locating and clicking screen elements without external help, both OpenAI and outside reviewers placed Astra in the low nineties, well past Sol's high seventies. . Coding Got Cheaper, Not Necessarily Better . On common coding benchmarks like DeepSWE and Frontier Code, MindStudio found Astra running close to even with Sol and with rival models. That is a tie, not a win. Where Astra pulls ahead is Terminal-Bench 4.0, a test built around long, messy terminal sessions that require chaining many commands and recovering from mistakes along the way. Artificial Analysis found Astra matching a top rival coding model on its Coding Agent Index while running under half the cost per finished task. The reason comes down to tokens. Astra tends to finish long coding sessions using far fewer of them. So the upgrade is not sharper code from a single prompt. It is lower cost and faster completion across sessions that stretch for hours, which describes most real engineering work far better than a short coding quiz does. One vivid example came from Matt Shumer, former chief executive of HyperWrite, who let Astra run inside Unreal Engine, the game engine behind Fortnite. Over an extended session, it built a working three-dimensional replica of part of Manhattan. That kind of result depends on holding spatial context across hundreds of linked actions, not producing one clever line of code.Also Read: 5 Ways OpenAI's NextSlide Acquisition Could Transform ChatGPT. Professional Tasks Move Closer to Full Automation .On benchmarks built around real office and technical work, Astra posted its strongest scores. On Agents' Last Exam, which spans financial modeling, engineering, and media production inside real software, it edged past both Sol and a leading rival reasoning model, using notably fewer tokens along the way. What matters here is not that a model can build a spreadsheet. Older models already managed that with a person guiding each step. The shift is that Astra can move between research, calculation, software interaction, and final output with less manual handoff between stages. That turns AI from a drafting assistant into something closer to a task owner. Also Read: OpenAI Brings ChatGPT Ads to India for Free, Go Users.Cost, Risk, and Who Should Care. Astra costs about two and a half times more per token than Sol. For short, simple requests, that premium rarely pays off. For long agentic sessions involving hundreds of tool calls, lower token use can offset or beat the higher price on a per-task basis. The economics favor Astra as task length grows, not as a blanket rule. OpenAI also confirmed Astra is its first model to reach a critical risk level for cybersecurity capability , meaning it can find unknown software flaws and build working exploits with the right access. New safeguards followed, including tighter review of high-risk actions inside ChatGPT and Codex. For any company weighing deployment, the real question is not whether Astra performs well. It is how much authority any AI agent should hold before permissions, audit logs, and rollback controls become non-negotiable. .Final Thought. The next race among AI labs will not be won on cleverness alone. It will be won on which system can be trusted to act without supervision and still be pulled back safely when something goes wrong. .You May Also Like: OpenAI Boosts ChatGPT's Medical Knowledge to Deliver Better Health AnswersOpenAI Launches GPT-Live Voice Model for ChatGPT with Real-Time AI ChatsOpenAI Expands ChatGPT Agent Features for Smarter Productivity in 2026.FAQs.1. What is the GPT-6 Astra? GPT-6 Astra is OpenAI's agentic AI model designed to go beyond generating responses by operating computers, using tools, coding, conducting research, and completing multi-step professional tasks. 2. Is GPT-6 Astra better at coding? GPT-6 Astra's biggest coding advantage is in long, tool-heavy workflows rather than basic code generation. It can handle extended terminal sessions, recover from errors, and complete complex development tasks with fewer tokens. 3. How does GPT-6 Astra differ from GPT-5.6 Sol? Astra's aggregate intelligence is broadly similar to GPT-5.6 Sol, but it is stronger at agentic execution. Its advantages are most apparent when tasks require computer interaction, sustained tool use, and multiple sequential actions. 4. Is the GPT-6 Astra expensive to use? GPT-6 Astra has a higher per-token cost than Sol, making it less economical for simple requests. Its economics become more attractive for long agentic workflows where lower token consumption and reduced human intervention can lower the total cost of completing a task. 5. What are the cybersecurity risks of GPT-6 Astra? Astra has reached OpenAI's critical cybersecurity capability threshold, meaning it can potentially discover software vulnerabilities and develop exploits when given appropriate access. Organizations therefore need strict permissions, monitoring, approval controls, and audit mechanisms. .Join our WhatsApp Channel to get the latest news, exclusives and videos on WhatsApp
Share
Copy Link
OpenAI released GPT-6 Astra in September 2026, marking a shift from answering questions to executing complex tasks autonomously. The AI model achieves a 72.6% success rate on desktop automation benchmarks and rebuilds Manhattan in Unreal Engine. But independent testing reveals it writes worse than its predecessor while carrying OpenAI's first critical cybersecurity risk rating.
OpenAI released GPT-6 Astra in September 2026, introducing an AI model designed to perform tasks autonomously rather than simply provide answers
4
. Within 48 hours of early access, developers turned the launch into a public stress test that revealed a clear pattern: GPT-6 Astra dominates spatial reasoning, 3D modeling, and agentic tasks while delivering weaker writing performance than GPT-5.6 Sol, its predecessor1
. The advanced AI model costs $10 per million input tokens and $50 per million output tokens, representing a 2.5x increase over Sol's pricing1
. OpenAI president Greg Brockman used the launch briefing to announce the arrival of AGI1
.Independent testing from Artificial Analysis found GPT-6 Astra's overall intelligence score nearly identical to GPT-5.6 Sol and trailing some rival models on general reasoning
4
. The same testers rated its writing below its predecessor and measured a drop of roughly 80 Elo points on a benchmark of economically valuable professional work1
. This performance split highlights how OpenAI built GPT-6 Astra to execute multi-step actions rather than think harder.The headline feature is computer use, which allows the AI model to drive a mouse and keyboard on a real desktop instead of returning instruction lists
1
. On OSWorld 2.0, a desktop automation benchmark that scores what percentage of ordinary desktop chores an agent finishes independently, OpenAI reported 72.6% completion at roughly 40 minutes per task, compared to 65.7% at 75 minutes for Sol1
. Independent write-ups from DataCamp and MindStudio confirmed similar results4
.On ScreenSpot-Pro, a test measuring how well models locate and click screen elements without external help, both OpenAI and outside reviewers placed GPT-6 Astra in the low nineties, well past Sol's high seventies
4
. These scores demonstrate the model's ability to perform tasks autonomously across extended sessions without losing track of goals.GPT-6 Astra demonstrates exceptional visual understanding and spatial awareness across real-world applications
1
. Matt Shumer, former CEO of HyperWrite, gave Astra a week inside Unreal Engine and watched it generate a replica of Manhattan, working street by street to perfect each one1
. Max Weinbach fed the model photographs of Apple Park and received a reconstruction in Blender that he described as doing "an absurd job"1
.Tom Krcha handed GPT-6 Astra a single image of a house and received the full interior as editable geometry running at 60 frames per second, complete with appliances and toys
1
. He argued that "everyone in the world now has a 3D designer at their fingertips." A developer posting as SuSu took the concept to city scale, watching Astra rebuild the Chinese city of Hangzhou and surrounding towns in Three.js in 24 minutes, with West Lake, Leifeng Pagoda, tea terraces, and wetlands all in place1
.Game development emerged as the most popular use case where GPT-6 Astra excels
1
.
Source: Geeky Gadgets
The AI model created a playable first-person shooter with detailed environments and mechanics in just 30 minutes, a task typically demanding extensive time and resources
3
. Anshu Chimala, former UX/UI designer at Apple, generated a 3D game in one shot in 45 minutes for barely a couple percent of his usage quota, calling Astra "some kind of turbo-AGI machine god for 3D games"1
.Rishi Prasad, a former developer at Coinbase and Eleven Labs, built Astral War in a day: a browser shooter with authoritative multiplayer servers, 12-person lobbies, controller support, and voice chat
1
. He described "a huge, step-function leap in visual fidelity" over what he built a month earlier with Claude Opus 5. By automating labor-intensive processes such as coding, level design, and asset creation, GPT-6 Astra allows developers to focus on refining gameplay and enhancing user experiences3
.In robotics, GPT-6 Astra achieved a 95% success rate in managing robotic arms for complex tasks, outperforming previous AI models
3
. This precision proves critical for applications ranging from industrial automation to autonomous systems. For engineers and researchers, the advanced algorithms reduce error rates and minimize setbacks in high-stakes environments such as manufacturing and healthcare, where even minor errors carry significant consequences3
.After 100 hours of testing across 15 distinct scenarios, GPT-6 Astra performed well in 10 out of 15 use cases when compared against Fable 5.1
2
.
Source: Geeky Gadgets
Astra demonstrated particular strength in tasks requiring structured outputs, browser automation, and cost efficiency, with total costs of $326.98 compared to Fable 5.1's $513.36
2
. However, Fable 5.1 completed tasks faster with a total runtime of 9 hours and 35 minutes versus GPT-6 Astra's 11 hours and 19 minutes2
.Fable 5.1 excelled in creative domains such as web design and game development, producing visually refined results
2
. Its website cloning capabilities consistently produced visually accurate replicas with fewer errors. GPT-6 Astra's slower pace stemmed from its tendency to ask clarifying questions, which ensured tailored and accurate outputs but added to overall runtime2
.Related Stories
On common coding benchmarks like DeepSWE and Frontier Code, MindStudio found GPT-6 Astra running close to even with Sol and rival models
4
. Where Astra pulls ahead is Terminal-Bench 4.0, a test built around long terminal sessions requiring chained commands and error recovery. Artificial Analysis found GPT-6 Astra matching a top rival coding model on its Coding Agent Index while running under half the token cost per finished task4
.The upgrade delivers lower token cost and faster completion across sessions stretching for hours rather than sharper code from single prompts
4
. For short, simple requests, the 2.5x premium rarely pays off. For long agentic tasks involving hundreds of tool calls, lower token use can offset or beat the higher price on a per-task basis4
.OpenAI classified GPT-6 Astra at a new critical risk level for cybersecurity, marking the first time OpenAI rated a model at this threshold
1
. The AI model can find unknown software flaws and build working attacks without a human pointing at the hole first1
. New safeguards followed, including tighter review of high-risk actions inside ChatGPT and Codex4
. This cybersecurity risk rating adds fresh weight to how companies manage access and deploy the model in production environments.On Agents' Last Exam, which spans financial modeling, engineering, and media production inside real software, GPT-6 Astra edged past both Sol and a leading rival reasoning model while using notably fewer tokens
4
. The shift means GPT-6 Astra can move between research, calculation, software interaction, and final output with less manual handoff between stages. This capability turns AI from a drafting assistant into something closer to a task owner, automating labor-intensive processes that previously required constant human oversight4
.Summarized by
Navi
[2]
[3]
[4]
08 Jul 2026•Technology

12 Aug 2025•Technology

20 Apr 2026•Technology

1
Technology

2
Policy and Regulation

3
Technology
