"Welcome to the AGI Era": OpenAI releases GPT-6 Astra
Article | Xiaojing, Tencent Technology
"Welcome to the AGI era."
Greg Brockman, co-founder and president of OpenAI, concluded the release of GPT-6 Astra with this statement.
On September 3rd, local time in the United States, OpenAI officially released GPT-6 Astra. Compared to previous models primarily responsible for answering, generating, and invoking tools, Astra takes a further step by continuously executing complete tasks, from receiving instructions and operating software to adjusting subsequent actions based on results. This is also one of the reasons OpenAI has reintroduced the term "AGI era."
OpenAI positions Astra as "the world's most powerful computer usage model." This capability has extended from simple tasks like browsing, emailing, and spreadsheets to professional software such as Power BI, KiCad, and FreeCAD, beginning to enter more complex workflows in data analysis, engineering design, software testing, and troubleshooting.
In the video showcased by OpenAI, Astra can start from a simple graphic and continue to complete tasks such as 3D games and product pages. OpenAI also provided examples including circuit board design, Blender modeling, legal documents, Excel, and scientific analysis.

The numerous benchmark results released by OpenAI show that the improvement in capabilities is concentrated in tasks requiring continuous actions, such as computer operations, long-range coding, engineering design, and scientific research. Astra achieved 100% in ExploitBench and significantly surpassed the previous generation model in tests related to computer operations and CAD.
Currently, Astra has been opened to some organizations and will gradually be available to ChatGPT Plus, Pro, Business, and Enterprise users in the coming days. It can also be accessed via OpenAI API and Amazon Bedrock. The standard API pricing is $10 per million input tokens and $50 per million output tokens, which is 2.5 times that of GPT-5.6 Sol.
01
First, let AI learn to "use a computer"
The biggest change emphasized by Astra is "using a computer." In the past, users had to connect AI to specific software to complete tasks. Companies needed to develop APIs, plugins, retrieval systems, and various connectors for the model to invoke internal tools.
Astra attempts to bypass some of this work. It can directly see the computer interface and complete operations using the mouse, keyboard, and browser. Examples provided by OpenAI include filling out online forms, updating CRM customer records, scheduling calendars, searching the web, and organizing search results into emails or documents.
It can also open Python notebooks to analyze scientific data, process data in Power BI, complete engineering designs using KiCad and FreeCAD, create websites and conduct front-end testing, as well as install software, check for errors, and handle issues that arise on the screen.
OpenAI demonstrated Astra completing PCB layout in KiCad. The model started from an electronic schematic, placed components on the circuit board, and then completed copper wiring. The entire process was compressed into 15 seconds. For engineers, this type of work, which used to require manual completion, can now be executed by the model.

Another demonstration involved completing 3D modeling in Blender and Unreal Engine 5. Astra first built a house in Blender and then converted the model into a walkable scene in Unreal Engine 5, allowing designers and clients to enter the scene in advance to view the space.

OpenAI also showcased scenarios involving game development, Excel, Power BI, automotive transmissions, legal documents, and the 1040 form.
In the OSWorld 2.0 offline subset test, Astra scored 72.6%, while GPT-5.6 Sol scored 65.7%. After simulating actual delays, OpenAI found that Astra completed each task in an average of about 40 minutes, compared to about 75 minutes for GPT-5.6 Sol, reducing the time by approximately 47%.
In the Agents' Last Exam, Astra scored 59.3%, while GPT-5.6 Sol scored 53.6% and Claude Opus 5 scored 55.5%. Additionally, under the highest score settings, Astra used approximately 65% fewer output tokens than Claude Opus 5.
The score for ScreenSpot-Pro reached 92.7%, while GPT-5.6 Sol scored 76.9%.
Speed has also improved. OpenAI updated the Codex framework simultaneously, and combined with Astra, the task completion speed in the Mind2Web test reached 1.9 times the current experience of GPT-5.6 Sol.
The official demonstration also included several everyday scenarios: searching for pediatricians, finding apartments, scheduling DMV appointments, looking for low-carb snacks, and analyzing kindergartens. In the demonstration, Astra completed a task in 2 minutes and 54 seconds.
Investor and AI practitioner Matt Shumer shared an experience on X. He had Astra create a world in Unreal Engine, then added human agents driven by Astra to the world, allowing these agents to coexist.

A day later, he heard sounds coming from the living room while in his bedroom, initially thinking someone had entered the apartment. He later discovered that the sounds were coming from those Astra agents, which had begun communicating with each other.
This experience is not yet fully stable, but Shumer believes that the effect of having multiple AI agents enter the same virtual world and interact autonomously was surprising enough for him.
Silas Alberti, Vice President of Advanced Research at Cognition, stated that the company plans to integrate Astra into the Devin framework on the day of release. In internal testing, Astra's capabilities in computer usage, writing, and codebase understanding showed significant improvements, with clearer video comprehension and more concise reports.
02 From writing code to completing entire tasks
After the enhancement in computer usage capabilities, the scope of work covered by Astra has further expanded. OpenAI positions it as a model for software engineering, professional work, and scientific research. It can handle longer tasks and adjust its work direction based on task changes.
This capability is particularly evident in coding tests.
In the Terminal-Bench 4.0 test, Astra scored 57.9%, while GPT-5.6 Sol scored 37.3%, and Claude Fable 5.1 scored 55.8%.
In DeepSWE v1.1, Astra scored 74.1%, while GPT-5.6 Sol scored 72.7%. In internal database migration tasks, Astra achieved 63.9%, while GPT-5.6 Sol scored 42.7%.
Astra has also introduced an improvement for long Codex tasks.
In the past, when long tasks encountered context window limitations, the model typically needed to compress previous work into a summary, which often resulted in the loss of details, such as why a particular fix failed or what issues a component had previously encountered.
Astra can retain important information during the work process in Codex and retrieve it in subsequent contexts. Early messages and tool outputs remain searchable, allowing the model to rediscover previous requirements, test results, and tool operation records.
Similar changes have occurred in professional work. In the BenchCAD test, Astra achieved a geometric overlap score of 95.9%, while GPT-5.6 Sol scored 83.3% and Claude Fable 5.1 scored 84.3%. In BrowseComp, Astra scored 91.5%, while GPT-5.6 Sol scored 90.4%. In the OpenScore String Quartets test, Astra achieved 0.84, while GPT-5.6 Sol scored 0.19.

OpenAI also demonstrated Astra's ability to create presentations, spreadsheets, and documents.
When given a few slides from OpenAI's presentation templates, it can create a new presentation following the original tone, layout, and structure. It also handles information trade-offs in tasks. OpenAI stated that Astra has been specially trained to bring important context into the final results, reducing the repetition of already completed work.

When task instructions are incomplete, Astra will also determine when to ask the user. If missing information will affect the final result, it will pose targeted questions. If the missing information does not affect the main direction, it can continue working and make reasonable assumptions. In Codex, even if the user does not respond temporarily, the model can continue processing other parts; when encountering significant decisions, it will wait for user confirmation.
Niko Grupen, Director of Application Research at Harvey, stated that in early legal task tests, Astra showed clearer distinctions between documents and existing records, making it easier to identify unsupported assumptions and translate information gaps into specific drafting proposals.
John Crepezzi from the Jane Street AI assistant team noted that Astra performed outstandingly in internal coding tests. When used for agent-based coding, its communication style is easier for developers to understand, and the generated code requires fewer modifications to reach production quality.
The scientific field is another key area. In FrontierMath Tier 4 v2, Astra scored 80.5% with an accuracy of 97.6%, while GPT-5.6 Sol scored 83.0%; GPQA Diamond reached 96.0%, while GPT-5.6 Sol scored 94.6%.

The Terminal-Bench Science 0.1 test evaluates scientific research processes, including data analysis, simulations, and model fitting. Astra scored 64.6%, while Claude Fable 5.1 scored 52.6% and GPT-5.6 Sol scored 22.4%.
In the scientific applications demonstrated by OpenAI, Astra can enter specialized scientific software, check sequencing quality, visualize genetic variations, and decide the next analysis direction based on data.
With the combination of scientific reasoning and computer operations, the model can directly work within the software environment originally used by researchers.
Greg Burnham from Epoch AI, which specializes in capability assessments, summarized this change at the release event as the end of one era and the beginning of another. OpenAI emphasized that Astra's value has extended from answering scientific questions to directly participating in scientific workflows.

03 The stronger the capability, the higher the security threshold
Astra also has a very special aspect this time: cybersecurity.
In ExploitBench, Astra achieved 100%, while GPT-5.6 Sol scored 78.5%; in ExploitGym, Astra scored 42.4%, while GPT-5.6 Sol scored 30.3%.

OpenAI has also established an internal test covering recent vulnerabilities from June to August 2026. In this test, Astra's arbitrary code execution rate was significantly higher than that of GPT-5.6 Sol, and it used fewer output tokens. During the testing process, Astra also discovered two previously unknown zero-day vulnerabilities, which OpenAI stated have been disclosed to the relevant maintainers.
The SRE-Bench tests the ability to reverse engineer software binaries without source code and understand their core logic. Astra solved 88.0% of tasks in a single attempt and reached 99.2% after four attempts, while GPT-5.6 Sol scored 55.9% and 68.7%, respectively.
The improvement in capabilities has also brought new security requirements. OpenAI stated that Astra has reached a critical threshold in its cybersecurity readiness framework. The model can discover previously unknown software vulnerabilities and may further form exploit chains.
Therefore, the current version of Astra will refuse to execute more advanced cybersecurity tasks, such as creating proof of concept for vulnerabilities. OpenAI plans to gradually expand access for trusted defenders through Daybreak for vulnerability validation, malware analysis, and detection engineering.
In terms of alignment, OpenAI also provided notable data. In the internal computer usage security benchmark, Astra's misconduct rate was 2.4%, while GPT-5.6 Sol was 22.0%; after adding AutoReview, Astra's rate was 1.8%, while GPT-5.6 Sol was 4.5%. In the internal evasion benchmark, Astra scored 0.00%, while GPT-5.6 Sol scored 0.29%.
In the ExploitGym honeypot test, Astra scored 0%, while GPT-5.6 Sol scored 48.2%. This test primarily observes whether the model attempts to bypass restrictions due to task difficulty. In the internal hallucination benchmark, Astra scored 4.2%, while GPT-5.6 Sol scored 12.2%.

OpenAI stated that Astra has improved in understanding task boundaries, handling ambiguous instructions, and explaining its capabilities to users. Its probability of capability hallucination is about one-third that of GPT-5.6 Sol.

However, OpenAI also acknowledged a problem: the stronger the model's capabilities, the more difficult it may be to monitor the reasoning process. Chief Scientist Jakub Pachocki believes that the model's ability to solve more complex problems with fewer natural language reasoning tokens also means that it becomes increasingly difficult for humans to judge what the model is actually doing through its textual reasoning process.
Therefore, OpenAI has incorporated misalignment monitoring into Astra's deployment. The system will check the model's reasoning and actions, and if unauthorized behavior is detected, it can pause the task.
This mechanism can also have practical implications. Legitimate tasks may sometimes be slowed down, paused, or even stopped. In ChatGPT or Codex, users may need to confirm before continuing; in API workflows, flagged tasks may be stopped directly.
As a result, a very realistic change has emerged with the release of Astra: as AI begins to gain more computer permissions, the concerns for enterprises have expanded from "Will the model answer incorrectly?" to "What can the model operate, what can it access, and when must it stop?"
OpenAI also mentioned that Astra is its first model to use over 100,000 DBUs for pre-training on the Stargate infrastructure. Aidan Clark, Vice President of Research at OpenAI, stated that based on the evaluation results from the pre-training phase, Astra's capability leap exceeds that of GPT-5.6 Sol compared to its predecessor.
OpenAI attributes this change to the combination of large-scale pre-training and reinforcement learning, with a further shift in training focus towards connecting information, executing longer tasks, and working continuously in complex environments.
04 More expensive but "more capable"
Astra's pricing has also changed significantly.
The standard price for OpenAI API is $10 per million input tokens and $50 per million output tokens. Previously, GPT-5.6 Sol was $4 per million input tokens and $20 per million output tokens. Both input and output prices have increased by 2.5 times.
Cache read and write are billed separately. OpenAI also offers a fast mode, with speeds reaching up to 2.5 times that of standard processing, priced at twice that of the standard mode.

Looking solely at token prices, Astra is clearly more expensive. However, OpenAI hopes enterprises will calculate costs differently. Brockman believes that as models increasingly resemble agents, the cost per token is becoming harder to accurately reflect the true cost. A model may have cheap tokens, but if it requires repeated attempts and manual corrections, the final cost of completing a task may be higher.
OpenAI's data also supports this assessment. In the Terminal-Bench Science 0.1 test, Astra scored 64.6%, compared to Claude Fable 5.1's 52.6%, estimating a 31% reduction in API costs. In low-cost settings, Astra scored 61.1%, while GPT-5.6 Sol's best result was 22.4%, estimating a 27% reduction in API costs.
In BenchCAD, Astra scored 95.9%, with estimated API costs about 43% lower than GPT-5.6 Sol and about 86% lower than Claude Fable 5.1. In Terminal-Bench 4.0, Astra scored 57.9%, while GPT-5.6 Sol scored 37.3% and Claude Fable 5.1 scored 55.8%. OpenAI estimates that the cost per task is approximately 9% and 63% lower, respectively.
Data from third-party evaluation agency Artificial Analysis presents another perspective.

GPT-6 Astra scored 61.2 in the Artificial Analysis Intelligence Index, close to GPT-5.6 Sol's 60.9, but with approximately 10% fewer output tokens. Due to the price increase of 2.5 times, the cost per task is actually about 75% higher than that of GPT-5.6 Sol.
In the Artificial Analysis Coding Agent Index, Astra scored 67.0, close to Claude Fable 5's 67.2 and Claude Opus 5's 68.1. Artificial Analysis believes that in the Codex environment, Astra significantly reduces the tokens required to complete tasks, and at the highest effort level, the cost per task is close to that of GPT-5.6 Sol; compared to Claude Fable 5, the cost of completing similar tasks is less than half.
However, Astra does not lead in all tests. Data from Artificial Analysis shows that it experienced a decline of about 80 Elo in GDPval-AA v2, and also showed some regression in tests such as τ³-Banking, SciCode, and AA-LCR. Humanity's Last Exam improved by about 6 points.
This is also a noteworthy aspect of Astra's release. OpenAI did not disclose Astra's scores in GDPval this time. GDPval itself is a test used by OpenAI to measure real-world economic performance, covering 44 occupations and 1,320 tasks, including legal briefs, engineering design, spreadsheets, presentations, customer support, and care plans.
From the capabilities demonstrated by Astra, GDPval has a strong correlation with the professional work scenarios it emphasizes, but OpenAI did not include this score in the main release materials.
Therefore, Astra's real-world work capabilities are currently more reflected through results from Agents' Last Exam, BenchCAD, AutomationBench, internal design tasks, and data science tasks.

Ultimately, whether Astra can become an AI employee that enterprises are truly willing to use long-term will depend on its cost, speed, accuracy, and the number of human interventions when completing actual tasks.
As for whether Astra is AGI, Brockman's answer does not shy away from the question. He believes that there is no universally accepted standard for AGI, and whether Astra qualifies as AGI can continue to be discussed. However, if a system can already undertake a large number of tasks in browsing, computer operations, programming, mathematics, science, law, and other professional work, then it is not unreasonable to say it has entered the AGI era.
OpenAI CEO Sam Altman also described Astra as a model that opens a new generation of entrepreneurship, scientific discovery, and creation.

Special contributor Jin Lu also contributed to this article.












