
1026: OpenAI’s GPT-6 Astra
About this episode
Get every episode summarized
Each time Super Data Science: ML & AI Podcast with Jon Krohn publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
195 searchable segments. Every word is indexed and playable.
Full transcript
Super Data Science: ML & AI Podcast with Jon Krohn — 1026: OpenAI’s GPT-6 Astra. Machine-transcribed; use the interactive transcript above to jump the player to any line.
This is episode number 1026 on OpenAI's GPT-6 Astra. Welcome back to the Super Data Science Podcast. I'm your host, John Krone. Today's episode is all about GPT-6 Astra, the new flagship model that OpenAI began rolling out last week, and that OpenAI's president, Greg Barachman, has gone so far as to suggest might one day be looked back upon as the arrival of AGI artificial general intelligence. That's a big claim, and we'll get to it, but first let me walk you through what the model is, what it can do, what it costs, and then I'll wrap it up with the safety story, which in this particular release is unusually intertwined with the capability story. Let's start with the basics. GPT-6 Astra is the first six series model from OpenAI. It's OpenAI's most capable model then of course, and it's positioned as the successor to GPT 5.6 Soul, which had been the company's flagship since July.
It also, the name seems to imply to me that Astra is even bigger in terms of parameter count than Soul, which is bigger than Terra, which is bigger than Luna, so Moon, Earth, Sun, Stars. I don't know, they don't release parameter counts anymore. But they did say that Astra came out of the largest training run the company has ever carried out, which used more than apparently 100,000 GPUs at a renowned compute facility in Texas called Stargate. In a detail, I found interesting. Astra is the first OpenAI model where other AI models played a significant role in supervising its training. The flywheel of models training models is now explicitly part of how the frontier gets pushed forward, another step toward RSI, recursive self-improvement, which you can hear all about in episode number 1,004. Anyway, the model GPT-6 Astra is available in the API, the OpenAI API.
Whoa, that's fun to say, the OpenAI API. It's under the model string GPT-6 Astra. And as with OpenAI's other recent models, you can dial in how much inference time compute it spends on a given task via a reasoning effort parameter. You can hear all about those in episode number 1,020 from a couple of weeks ago. In Astra's case, there are five settings of this reasoning effort parameter, low-medium-high-x-high, and max. A standard API pricing is $10 per million input tokens, and 50 per million output tokens, which is about two and a half times what Soul has been going for on promotion, but is in line with what Anthropic charges for its top-tier, Claude Fabel 5.1 model. There's also a fast mode, which I hadn't seen before, and the API that delivers up to two and a half times the speed for double the price. So if money is no object to you, you can get these frontier capabilities for faster. Astra is also available through Amazon bedrock,
and if you're a chat GPT user, it's rolling out over the coming days to plus pro business and enterprise subscribers with the three higher tiers of those, pro business and enterprise, additionally getting a beefier variant called GPT-6 Astra Pro. But exactly what makes it beefier, or different at all, from the regular non-pro variant has not been yet publicly disclosed. At least at the time may recording this. One note for people in larger organizations, enterprise administrators have to switch Astra on for their workplace because at launch, it's off by default. Now, onto what the model can do. Across the board, open AI is claiming state-of-the-art results on computer use browsing, software engineering, server security science, and professional knowledge work, and the benchmark table they published is long. But anytime somebody is releasing a near frontier model, we get the same kind of, wow, look, it's absolutely the best on every benchmark that we tested. Wow. So I don't know, grain of salt, but this does seem to be a pretty serious model release, and I'll pull out a few
of the highlights that I think matter most to this audience. The first is computer use, which is the area open AI is leaning on hardest in its messaging. The pitch is that Astra can take over the tedious, click heavy parts of your working life, filling out online forms, updating records in a CRM, organizing your calendar, running front-end QA checks on a website it just built, installing and troubleshooting software, and so on. On OS World 2.0, a benchmark of realistic desktop tasks, Astra scores 73%, versus Soul, which had just 66%, that's actually not that much of a difference. And according to opening AI's latency simulations, it gets there in 40 minutes per task, versus roughly 75 minutes per soul. So it's a little bit more accurate, but it's a lot faster, 47%. And on that particular benchmark, OS World 2.0. On ScreenSpot Pro, yes, ScreenSpot Pro, which tests whether a model can locate the right element on a professional
application screen, Astra scores about 93%, which is a big jump relative to Soul's 77%, that is significant sounding. After computer use, the second highlight is coding. On Terminal Bench 4.0, which covers agentics software engineering, system configuration, and data analysis from the command line, Astra scores about 58%, versus just 37% for Soul. And so that 58% from Astra is about the same as QuadFable 5.1, which came in at 56%. So I'm sure you can notice that 56 to 58% difference. On another benchmark called Frontier Code, the models are bunched much more tightly together with Astra ahead of Soul, and again, essentially level with the top-clod models. So the coding lead isn't uniform across every evaluation. One coding adjacent feature that I think deserves a mention is a new approach to long agentic sessions in Codex, opening eyes,
coding tool, rather than repeatedly compacting a session's history into a single summary as the context window fills up. Astra can keep running notes across context windows and earlier windows remain searchable. So details like why a given fix failed don't get compressed away. And that's experimental for now, but we'll apparently become the default for Astra in the coming weeks. The third highlight and probably the one that will get the most airtime in the mainstream press is abstract reasoning and math. Astra scores 99.9% on Arc AGI3, a benchmark that's Soul scored under 8% on. So opening eye jumps from around 8% to effectively 100% on that abstract reasoning and math benchmark. That is pretty insane. I wish I had the data on how Fable 5.1 from Anthropic performs on math, but just don't have it online for some reason, at least at the time of me recording. The best I can see is that Opus 5 scores 30%. So yeah, TBD on whether this is
actually an enormous jump on frontier abstract reasoning and math overall or just for open AI. Other than Arc AGI3, GBT6, Astra scores about 98% on frontier math tier 4, which is the hardest year of research level mathematics problems in that suite, which is up a non-trivial amount from Soul, which had 83%. So 83% to 98% jump. Open AI has also published two new mathematical results that Astra helps produce. So new math being discovered here. And both of those results concern the gaps between prime numbers. So one of those tightens a long standing bound on how closely together infinitely many pairs of primes can occur. So it tightens that from 240 down to 186. And the other new mathematical results improves a term in a bound on unusually large prime gaps that had stood unchanged for more than 80 years. I can't really don't really understand what that
means or what the implications are, but I guess it sounds like it's a big deal. Whatever you make of the AGI question, a model that contributes to open problems in number theory is a meaningful milestone. It sounds like to me. All right, the fourth category that I want to highlight here is benchmarking on science and professional work more broadly. So on terminal bench science, benchmark, which tests whether an agent can carry out research workflows like analyzing data, running simulations and fitting models. Astra scores about 65% versus about just 22% for Soul and about 53% for Fable 5.1. So that seems like maybe a place where GPT-6 Astra really has an edge on the frontier. And on agent's last exam, a fun benchmark that's playing on humanities last exam, benchmark, agent's last exam tests agents on complex professional tasks in real software, across domains like financial modeling, engineering and media production. And there,
Astra edges out Claude Opus 5, Claude Opus 5, while opening AGI says using roughly 65% fewer tokens. Again, I don't have the Fable figures for you here, but at least compared to Opus 5, that token efficiency point comes up repeatedly in the chat GPT announcement. And it matters for cost because a model with a higher per token price can still actually produce a lower overall bill for you if it finishes a job in fewer steps with fewer retries. Regular listeners will already be aware that I'm obsessed with Anthropics Fable 5 model and it has taken over my working life. I'm writing a technical book that includes latex files, mathematical notation, Python code examples, and Fable 5 in Claude code handles requests I make across whole chapters with accompanying Jupyter notebooks and to end work that a few short months ago would have been dozens of separate requests with way more manual fiddling required. With Fable 5, it just works, essentially like magic first time. Claude is the AI for problem solvers. It's the collaborator that understands your
entire workflow and thinks with you, not for you. Whether you're debugging code at midnight, building a financial model or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. For problems worth solving, get started with Claude at Claude.ai slash super data. That's Claude.ai slash super data. And check out Claude Pro, which includes access to all of the features mentioned in today's episode. Claude.ai slash super data. All right. And the final, the fifth and final capability that I'm going to highlight here is something that OpenAI is making a point of Astra's judgment when instructions are ambiguous. The claim is that Astra fills in routine gaps with sensible assumptions, but pauses to ask a focused question when the answer could change the outcome. And in OpenAI's coding tool codex, here's something that thought was actually really cool. I would love to experience this. GPT-6 Astra can ask a question asynchronously of you while continuing with the parts of the
task that don't depend on your reply. That is damn cool. And hopefully they're working on that for Claude code as well. My typical environment. It's also reported to be better. GPT-6 Astra at incorporating a mid-task course correction without losing track of the original goal, which anyone who has steered an agent through a long task will appreciate. Okay. That's it for the capability story. Now for the safety story. And in this respect, Astra differs from a typical model launch. So Astra is the first model OpenAI's designated as reaching the critical threshold for cybersecurity under its preparedness framework. What that means concretely is that tested without production safeguards, the model can find previously unknown vulnerabilities and turn them into working exploits across well-protected systems without a human guiding each step. This sounds quite similar to what was happening with that mythos and later fable drama from Claude on exploit bench, for example, which tests whether a model
can turn known vulnerabilities into working exploits, Astra scored a perfect 100% versus 79% for soul to address concerns that the model might have memorized historical vulnerabilities. OpenAI also built a fresh evaluation using 20 high-severity chrome vulnerabilities from June through August of this year. And during that evaluation, Astra discovered and used two previously unknown zero-day vulnerabilities, which OpenAI says it is disclosing to the chrome maintainers. Expert assessments found the unrestricted GPT-6 Astra could achieve arbitrary code execution in hardened browsers and build privileged escalation exploits for hardened operating systems. Because of this, the launch of GPT-6 Astra has been phased and gated. OpenAI slowed Astra's release to add safety testing. The version everyone is getting will help with defensive work like secure code review and patching, but it will refuse more advanced tasks such as writing proof of concept exploits. Organizations in OpenAI's application-based security program called daybreak
got access first, and OpenAI says it plans to reload less restrictive safeguards to that group in the coming weeks for workflows like malware analysis and detection engineering. There's also a new layer of misalignment monitoring in production for Astra class models, classifiers that check the model's reasoning and actions and automatically stop activity that looks unauthorized. OpenAI is candid that this can sometimes pause or stop legitimate work. In chat GPT and codec you may be asked to review an action before continuing while in the API the task will simply halt. This all happens, of course, against a specific cybersecurity backdrop. As I detailed in episode 2014 about a month ago, and I'm sure everyone has heard about by now, two OpenAI models escaped their containment during research, accessed the OpenWeb, and breached hugging-faces systems. Opening up pause parts of its research and training in response, including work on Astra, even though Astra wasn't actually one of the models involved apparently. And OpenAI added safeguards before release. Reporting from CNBC indicates the model also went through a formal review process with the US administration before launch, and OpenAI
built a new evaluation informed by the hugging-faces incident that measures whether a model given a difficult or impossible task will exceed its authorized scope. Without production safeguards, so went beyond the authorized target 48% of the time while Astra did so in 0% of cases. So it does seem like they're figuring out how to rein their models in. Yeah, there is one caveat in OpenAI's own materials that I think is worth flagging. Their evaluations found that Astra's written reasoning is harder to monitor than souls, in tests that explicitly ask the model to evade monitoring. So OpenAI attributes this to Astra having greater control over its written reasoning and being able to solve problems in fewer written steps. And it notes that Astra still appears to struggle to conceal the reasoning needed for complex tasks. But a step backward on monitorability like this, at the same time as we take a step forward on offensive cyber capability is exactly the combination the field should be watching closely.
And OpenAI says answer research priority. So finally, is this AGI? Brockman's own framing was that AGI remains a grave fuzzy concept that the term is no longer tied to any contractual trigger with Microsoft and that he personally thinks future observers might point to this model as the moment to this model as the moment. He closed the press briefing by welcoming everyone to the AGI era. I'll leave you to form your own view, but here's mine. As I detailed the years ago in episode number 748, I don't think it makes sense to think of AGI as a binary event. There are varying degrees of AGI breadth and depth. So instead of thinking of this, GPT-6 Astra release in binary AGI terms, we can acknowledge this is at most a small step forward in general intelligence capabilities relative to athropyxfable 5.1 model. But having not used GPT-6 Astra myself nor having much third party benchmarking to work with yet,
this might even just be bringing OpenAI in the vicinity of the Fable 5.1 frontier. So yeah, not necessarily the big deal that Brockman's making it out to be. But regardless, we now have at least two models near the frontier of general intelligence capabilities, including with GPT-6 Astra, that means a generally available model that contributes to open mathematics research, works through desktop applications faster than most people can and finds zero day vulnerabilities in hardened software. That's powerful for sure. And more, along this staggering trajectory is coming in the months ahead, no doubt, including no doubt open source alternatives. If you'd like to dig further into any of this, I've of course put links in the show notes to OpenAI's full announcement, which includes the complete benchmark tables and the two prime gap proofs. Maybe one of you will actually understand them. As well as the
Astra system card where the cybersecurity and alignment findings are documented in detail. Wow. Now go get your hands dirty. Think carefully about what you can now delegate to agents, think even more carefully about what you should, and then go build something worthwhile. Cool. And we do have a new Apple podcast review for me to read for you on air. This one is short. It comes from Ronin 186. The title is simply, I'm a huge fan of the Super Data Science podcast exclamation mark. It's a five star review. And wow, I'm particularly fond of the body of this review. It just says John Crony's great with three exclamation marks. All right. Thanks Ronin 186. I think you're great too. Thanks for all the recent ratings and feedback on Apple podcasts, Spotify and all the other podcasting platforms out there as well as for likes and comments on our YouTube videos, bonus points. If you leave written feedback on Apple podcasts, if you do,
I'll be sure to read your feedback on air like I did today. At the time, I'm only seeing feedback in the US app at some day, especially when I don't have US reviews to read. I will go and look at other countries and read the reviews you're putting in there too. All right. That's the end of today's episode. If you enjoyed it or know someone who might consider it sharing this episode with them, tag me in a LinkedIn post with your thoughts. And if you aren't already, be sure to subscribe to the show. The most important thing though is that we hope you'll just keep on listening. Until next time, keep on rocking it out there. And I'm looking forward to enjoying another round of the Super Data Science podcast with you very soon.
More episodes
More from Super Data Science: ML & AI Podcast with Jon Krohn

1025: Word Gravity: How Transformers Bend Space, with Dr. Luis Serrano
Super Data Science: ML & AI Podcast with Jon Krohn

1024: In Case You Missed It in August 2026
Super Data Science: ML & AI Podcast with Jon Krohn

1023: Agentic AI Skills That Matter Now, with Aishwarya Srinivasan
Super Data Science: ML & AI Podcast with Jon Krohn

1022: CLAUDE.md, AGENTS.md, Skills, Hooks and Subagents: A Field Guide to Steeri...
Super Data Science: ML & AI Podcast with Jon Krohn