sajawal.dev
BOOKING · SEP '26
·9 MIN

GPT-6 Astra: What it means for engineers and founders

OpenAI just released GPT-6 Astra. I break down its new computer use, coding, science, and cybersecurity capabilities. Learn about the API, pricing, and safety.

SHARE
gpt-6-astra-cover

OpenAI just announced GPT-6 Astra. This is their "most intelligent and aligned" model so far. For me, as an engineer and founder, these releases are not just news, they shape how I build products and mentor others. I have studied the details from OpenAI's pages. Here is my take on what Astra offers.

What Astra can do

Astra is a big step for computer use, browsing, software engineering, cybersecurity, science, and professional work. It started rolling out on September 3rd to a few organizations. Soon, it will be in ChatGPT Plus, Pro, Business, Enterprise, the OpenAI API, Microsoft Azure, and AWS Bedrock. For enterprise users, admins need to turn it on, as it is off by default.

Computer use and automation

OpenAI calls Astra the best computer-use model in the world. It can run websites and desktop apps on its own. It handles many steps in knowledge work.

Think about these tasks: filling online forms, updating CRM records, or organizing calendars. It can do web research and write summaries for emails or documents. For science, it analyzes data, makes plots, and even creates websites. It can also install and test software, then troubleshoot issues it sees on screen.

This model is also faster. On OSWorld 2.0, Astra scored 72.6 percent. It took about 40 minutes per task. The previous model, Sol, scored 65.7 percent and took about 75 minutes. That is almost half the time per task. With updates to the Codex harness, Astra finishes tasks 1.9 times faster than GPT-5.6 Sol on Mind2Web.

Professional work

Astra creates professional documents, slides, spreadsheets, and analyses. It matches your style and templates. It is good at picking out only the important context. It also has better visual judgment for websites, games, apps, and renderings.

In ChatGPT, Astra has a new feature called Sites. You can create, host, and share websites, web apps, and games directly from a prompt. If Astra needs more information, it asks focused questions. In Codex, it can ask questions in the background while still working on other things.

Coding and software engineering

OpenAI says Astra is the best model for software engineering to date. This is a big claim. There is a new Codex feature that lets Astra keep and find context across different windows. This means it remembers notes and earlier windows. It helps reduce losing information during long debugging or refactoring sessions. This feature is experimental now, but it will be the default soon.

Astra also stays focused as tasks change. It takes in new rules without forgetting the main goal.

Science, math, and health

Astra is a major advance for scientific discovery and math. It has already helped solve some long-standing open problems. OpenAI shared two new results on prime gaps with the launch.

It scores 97.6 percent on FrontierMath Tier 4 (v2). On GPQA Diamond, it gets 96.0 percent. Astra combines scientific reasoning with computer use. It works directly in specialized tools to inspect data and explore results.

Cybersecurity

Astra is the first OpenAI model to reach the "Critical" cybersecurity threshold in their Preparedness Framework. This is a big deal. Without safeguards, it scored 100 percent on ExploitBench, compared to Sol's 78.5 percent. On ExploitGym, it got 42.4 percent, while Sol got 30.3 percent. Astra even found and used two zero-days during internal evaluations.

Because of these strong capabilities, advanced cyber tasks, like creating exploit proofs of concept, are blocked for general use. Access for defensive tasks, such as vulnerability validation or malware analysis, will come later through Daybreak and Daybreak Blue programs.

Alignment and safety

Astra is OpenAI's most aligned model. It understands intent better and respects boundaries more. Sol went beyond its authorized scope 48 percent of the time in one test. Astra did not do this at all.

Astra is also three times less likely to misrepresent its own capabilities than Sol. OpenAI is working on how to monitor its reasoning, as Astra solves simpler tasks with fewer steps, making it harder to track. This is a research priority for them.

To keep things safe, they have stronger safeguards against jailbreaks. They also use misalignment monitoring in production. Classifiers check Astra's reasoning and actions. They can pause or stop tasks if needed. This might sometimes slow down legitimate work, even for defensive security tasks. In ChatGPT and Codex, you might be asked to review before continuing. In the API, the task stops.

Benchmarks

Here are some of the benchmark results from OpenAI. "Lower is better" where noted.

Computer Use

BenchmarkAstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5
Agents’ Last Exam59.3%53.6%to48.7%55.5%
OSWorld 2.0 (offline, partial)72.6%65.7%toto70.2%
ScreenSpot‑Pro (no tools)92.7%76.9%to87.3%to

Professional Work

BenchmarkAstraSolFable 5.1Fable 5Opus 5
AutomationBench41.4%18.1%31.4%17.4%26.9%
BenchCAD95.9%83.3%84.3%67.5%82.1%
BrowseComp91.5%90.4%to87.4%90.8%
Internal Design Tasks50.0%47.4%to35.8%to
Internal Data Science Tasks40.9%30.5%to34.7%to

Coding

BenchmarkAstraSolFable 5.1Fable 5Opus 5
Terminal‑Bench 4.057.9%37.3%55.8%42.0%52.3%
DeepSWE v1.174.1%72.7%67.4%69.9%73.7%
FrontierCode 1.1 Extended64.5%60.6%63.6%64.9%63.6%
FrontierCode 1.1 Main53.3%47.5%50.9%53.5%53.4%
Internal Database Migration Tasks63.9%42.7%57.8%50.3%to

Academic (math/science)

BenchmarkAstraSolFable 5.1Fable 5Opus 5
Terminal‑Bench Science 0.164.6%22.4%52.6%21.4%30.0%
FrontierMath Tier 4 (v2)97.6%83.0%87.8%87.8%73.2%
GPQA Diamond96.0%94.6%93.7%92.6%93.7%
Humanity’s Last Exam (w/ tools)57.2%to65.0%63.8%63.6%

Science & Health

BenchmarkAstraSol
GeneBench Pro37.8%28.7%
MedChemBench (internal)49.3%47.4%
LifeSciBench60.3%59.9%
HealthBench Professional (length‑adjusted)63.4%60.5%

Cybersecurity

BenchmarkAstraSolOpus 5
ExploitBench100.0%78.5%70%
ExploitGym42.4%30.3%22.0%
ExploitBench (Jun to Aug 2026)39.0%11.5%to
SRE‑Bench88.0%55.9%12.5%
SEC‑Bench Pro85.4%79.1%to

Alignment (lower is better)

BenchmarkAstraSol
Internal computer‑use safety2.4%22.0%
+ AutoReview1.8%4.3%
Internal circumvention0.00%0.29%
ExploitGym honeypot0.0%48.2%
Impossible ExploitGym100.0%to
Internal hallucination4.2%12.2%

Long Context

BenchmarkAstraSol
OpenAI MRCR v2 8‑needle 256K to 512K100.0%91.5%
OpenAI MRCR v2 8‑needle 512K to 1M96.3%73.8%

Abstract Reasoning

BenchmarkAstraSolOpus 5
ARC‑AGI‑399.9%7.8%30.2%
ARC‑AGI‑295.0%92.5%90.4%
ARC‑AGI‑198.5%97.5%97.5%

Note: For ARC-AGI-3, the 99.9 percent score used OpenAI's responses-API harness. This had two settings tuned for real-world performance.

API specs, pricing, and deployment

In the API, the model name is gpt-6-astra. It supports up to about 1.05 million tokens for context and 128k max output. There is a Fast mode, plus batch and flex options.

Pricing is set at $10 for 1 million input tokens and $50 for 1 million output tokens. Fast mode costs twice as much but offers up to twice the speed. There are also separate rates for cache reads and writes. Batch and flex processing can lower costs for certain tasks. Astra is available through the OpenAI API, Microsoft Azure, and Amazon Bedrock.

For API customers who qualify, Astra supports Zero Data Retention. OpenAI is also testing Private Safety Processing. This aims to improve safety monitoring while keeping data private.

ChatGPT and Codex product changes

Astra usage is part of existing ChatGPT allowances. If you use a lot, you can buy extra credits. Pro, Business, and Enterprise plans also get Astra Pro.

For Codex, the new context preservation feature is key. It remembers notes and searchable history across windows. This helps avoid losing details during debugging or refactoring. Updates to the harness make computer use 1.9 times faster than the current Sol experience on Mind2Web. In ChatGPT, the Sites feature lets Astra create, host, and share websites, web apps, and games directly from prompts.

My takeaway: What this means for you

GPT-6 Astra is not just another model update. It is a significant leap. Sol was already strong in coding, science, and cyber. But Astra changes things with its speed and accuracy in computer use, the quality of professional outputs, abstract reasoning, and its offensive cyber capability. That last point is why some features are gated.

As an engineer, the context preservation in Codex is huge. Losing context during a long debug session is a real pain point. This feature alone could save a lot of development time. For founders, the ability for Astra to autonomously handle multi-step knowledge work means new possibilities for agent-based automation in businesses. Imagine an agent that can truly manage a complex workflow from start to finish.

The pricing for Astra is higher than Sol's lower-cost tiers. This makes sense for a frontier capability. You will need to weigh the cost against the gains in speed and accuracy. For my work with Placewise AI and ReachWise, these advancements mean I can build more sophisticated and reliable AI agents. The focus on alignment and safety is also good. It helps build trust in these powerful new tools.

If you are building new products or optimizing existing workflows, understanding Astra's specific strengths and limitations is critical. This model changes what is possible.

SHARE