Notes & thoughts

GPT-6 Astra Is Here: What We Actually Know

OpenAI started rolling out GPT-6 "Astra" this week — initially to select customers, with the usual fanfare. Sam Altman teased the launch a day ahead and called it "the most intelligent and aligned model in the world," which is what he always says, so I waited a couple of days for the smoke to clear before writing this. Here's the rundown, with a clear line between what's confirmed and what's still leaked gossip.

Performance

The headline number: reports say Astra scored nearly perfect on an "AGI test," which I take with a big grain of salt because nobody agrees on what an AGI test is. The more grounded claim is competitive — leaked benchmarks suggest it comfortably beats Anthropic's Claude Fable 5.1, which means OpenAI has arguably taken the lead back after a year of playing catch-up. Independent boards haven't caught up yet, though: Astra isn't listed on OpenRouter or Artificial Analysis as of today, so there's no third-party number to check the hype against. The other real upgrade is agentic: Astra is noticeably better at multi-step tool use, which matters more to me than any leaderboard.

Artificial Analysis Intelligence Index · Sep 2026 Claude Fable 5.1 (Max Effort) 66 Claude Fable 5.1 (Xhigh Effort) 65 Claude Opus 5 (Max Effort) 63 Kimi K3 (max) · GLM-5.3 (max) 60 GPT-6 Astra — not yet listed ?
Scores from artificialanalysis.ai, the most-watched independent benchmark board. GPT-6 Astra isn't on OpenRouter or Artificial Analysis yet (the newest OpenAI models listed are the GPT-5.6 series from July), so independent numbers are still pending — the dashed bar marks where it would need to land to back the "nearly perfect AGI test" claims.

What it can actually do

The genuinely new trick is that Astra can run desktop apps and automate tasks by voice — less "chatbot," more "agent that lives on your machine." The interesting caveat: its riskiest cybersecurity capabilities are locked down and restricted to a small group of approved testers. So what most of us get is the careful version.

Cost

Early reporting puts it at roughly 2.5× the previous generation's price. Exact per-token API pricing hasn't been published yet as far as I can find, so if you're building on it, budget for the multiple and wait for the official pricing page before committing to anything.

💰 Cost at a glance: ≈ 2.5× the previous generation. Official per-million-token API pricing: not yet published.

The safety noise

It wouldn't be a major launch without drama. A pre-release report in The Verge raised fears of a safety "race to the bottom" between the labs, and Reuters reported that OpenAI told lawmakers it's building "automated shutdown" capabilities for AI tools — which is either reassuring or ominous depending on how your morning is going. I mostly find it notable that the safety conversation is now happening at launch speed instead of after the fact.

My honest take

Half of what we "know" right now is leaked benchmarks and marketing lines, so treat all of the above as a snapshot, not scripture. What seems solid: it's a real capability jump, especially for agents; it's expensive; and the most powerful parts are gated. I'll write a follow-up once I've actually used it for something real — my rule from the Claude Code days still holds: the demo is never the job.

New model launches all sound the same on day one. The only review that matters is yours, three weeks in, when the novelty has worn off and the bill arrives.

Subscribe

Get an email when I publish a new post.

Powered by Buttondown