
The Most Important New AI Tools from OpenAI DevDay
Get every episode summarized
Each time The AI Daily Brief: Artificial Intelligence News and Analysis publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
The AI Daily Brief: Artificial Intelligence News and Analysis is made possible by:
All brands on The AI Daily Brief: Artificial Intelligence News and Analysis →
“The team at OpenAI said that one of the big changes that had happened internally was that the power of the latest generation of models, like Astra, had increased their speed of development so significantly that many things that they thought were only going to…”From the transcript
OpenAI unveiled more than 20 launches at Dev Day, from always-on Dots agents and the shared workspace Space to cheaper models and new ways to use your ChatGPT subscription across other apps. NLW breaks down the most important announcements, the early reactions, and what they reveal about AI’s shift toward persistent agents, team collaboration, and more affordable intelligence.
Register for our Free Webinar: Build Your Personal AI Benchmark - https://aidailybrief.ai/webinar/personal-ai-benchmark
Next Cohort - Learn How to Build Agents - https://register.besuper.ai/register?program=ati
AIDB Fall Listener Survey - https://aidailybrief.ai/survey
Multiplayer AI Sprint - https://multiplayerai.ai/
Brought to you by:
KPMG – Research from KPMG and the University of Texas at Austin shows the highest-impact AI users treat AI like a reasoning partner — and those skills can be taught at scale. Learn more at https://kpmg.com/us/Sophisticated
Harbor - Invest in the AI ecosystem. https://www.harborcapital.com/aidaily
Hyperagent - Hire a team of always-on agents. New users get $100 in free credits. hyperagent.com/aidailybrief
Rackspace Technology- One accountable partner to build, operate and run your full enterprise AI stack https://www.rackspace.com/
Section - Section turns AI investment into workforce transformation and ROI - https://www.sectionai.com/
Blitzy - Want to accelerate enterprise software development velocity by 5x? https://blitzy.com/
The AI Daily Brief helps you understand the most important news and discussions in AI.
Interested in sponsoring the show? [email protected]
Get every episode summarized
Each time The AI Daily Brief: Artificial Intelligence News and Analysis publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
298 searchable segments. Every word is indexed and playable.
Full transcript
The AI Daily Brief: Artificial Intelligence News and Analysis — The Most Important New AI Tools from OpenAI DevDay. Machine-transcribed; use the interactive transcript above to jump the player to any line.
The team at OpenAI said that one of the big changes that had happened internally was that the power of the latest generation of models, like Astra, had increased their speed of development so significantly that many things that they thought were only going to come in 2027 were actually coming as part of this DevDay announcement. The result of that was more than 20 different launches and announcements, including some big headliners like OpenAI's answer to Muse and Grockbot, and some sleeper hits, like the fact that enterprise accounts can now use OpenAI credits on an open marketplace to buy access to open source models as well. Overall, what we got at OpenAI DevDay does not change the big patterns and trends that we've been seeing in the industry, a move to more cost-efficient models, those models moving to more persistent work, and some amount of that persistent work moving from a solo to a multiplayer experience. Instead, what DevDay reinforced is that these are the trends to pay attention to and the ones that will reshape how we all use AI in the months to come. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.
Alright friends, quick announcements before we dive in. First of all, thank you to today's sponsors KPMG, Blitzy, robots and pencils, and HyperAgent. To get an ad free version of the show, go to patreon.com slash AI Daily Brief, or you can subscribe and have a podcast. To learn more about sponsoring the show, send us a note at sponsors at AI Daily Brief.ai. There are also all sorts of other goodies at AI Daily Brief.ai. On our website, you can get access to all of our free AI training programs. Those include free multi-week self-directed programs like the multiplayer AI Sprint, which is live right now, but also free live webinars like the one that is coming up on Thursday, October 1st at noon, that is all about building your personal AI benchmark. You can also find links to our paid trainings, such as the super intelligent executive catch-up and the super intelligent executive agent leadership program, both of which are registering new cohorts that start over the next couple of weeks. Now last note on today's episode, originally I was going to try to jam in a full accounting of everything that happened at the White House yesterday on top of OpenAI DevDay, but there was just too much to fit, so I've decided to go all DevDay recap today,
and then we will dig much deeper into what came out of Trump's meeting with the leaders of all the major frontier labs, and if and how it changes anything about the future development of AI. For now though, let's get into DevDay and find out what it says, not just about OpenAI, but about where we are with AI in general. It was an absolutely monster day of announcements from OpenAI. In fact, there were so many that on today's episode, we are just going to get through all of the big ones, discussing what the announcement was, whether it was expected or not, people's first impressions, and how it shifts the AI race if at all. Later on in the week, we'll go deeper on some of the most important ones, but for now it is going to take all that we have just to get through it. First up is dots. OpenAI's answer to Muse, Grockbot, and the wave of personal agents that are taking over the AI space. Sam Altman presented dots as quote, remarkably capable, always on agents built to handle, really anything you can think of. They bring AI to a whole new form factor. With dots, users can create a persistent agent, communicate with it through a text message style interface, and then set them to work on tasks using their own cloud
computer with support for 40,000 different apps. Dots can also communicate via voice call or across Microsoft Teams and Slack, with persistent context carried over between services and sessions. Functionally, it shares a lot with Muse right down to the Qt mascot. However, at least at the time of release, there are a few caveats worth noting. For now, users can only create one dot at a time, but OpenAI plans to expand teams of agents in the future. The other big difference from Muse is that OpenAI is only offering dots to pro business and enterprise customers. One of the big reasons Muse has been so successful is making it completely free, meaning OpenAI is naturally limiting their ability to compete on that front. CFO Sarah Fryer said that the vision is to bring dots to the whole consumer base, but for now, this is a pro-sumer product. OpenAI's big selling point is that they're using GPT-6 astra to power the agents, so this is the first time we're seeing a truly frontier model pilot a first party personal agent. Now, this was in many ways the least surprising of the announcements. This personal agent form factor has been basically inevitable since the
launch of OpenClaw, and even the interest in business-focused versions like GROKBOT, OpenAI getting into this particular area was pretty much inevitable. Now, how ready for prime time dots is remains to be seen. The live demo had some issues, although I would put way less in that than the chattering classes on X. Demo's going wrong is basically just a way for you to know that they are actually live. Unfortunately, once people got their hands on dots, they ran into a number of other teething problems as well. Chase Browser posted a session where Dot lost all of its work, and Tenabress said that they were excited to try it out, but found that they couldn't connect multiple computers at once to share their full context. Others had a better experience. Analyst Max Weinbach wrote, My basic test for how well one of these works is candidatonimously do my expenses with access to my email and Google Drive. GROKBOT made a mess and didn't do what I told it. It did the first like three days and then started to mess up. Dots did it after the first time and has been good. Justin Trotter said that he thought that dots was even more work focused than GROKBOTTER Muse. Think a bit less personal assistant he writes and a bit more codex orchestrator and slack collaborator. In fact, Justin writes,
Slack is where it really shines. Your team can message your bot directly and your bot can take actions like spin up new pains. Also getting at a debate, which I think will be everywhere, Justin suggests that he thinks that this should have been a separate app. Chat GPT is great, he writes, codex is great. They are very different products in my opinion and dots are way more codex than Chat GPT. The every vibe check almost leaves it as a TBD, with their users reporting flashes where it was truly excellent, but then other times where it was just extremely frustrating. When Brandon Chu suggested that this was the form factor that everyone was converging on, Nate B. Jones wrote, I think we need to distinguish between form factor and utility here. Yes, form factor is converging at the moment that may change, but utility is not. Muse is good at specific things like phone calls and practical email work. Instinct is good at specific different things like travel. Dots is good at AI context as a work surface, regardless of positioning that's distinct utility. The question with a T for trillions is did you pick the right utility to get right? We'll come back to dots later in the week, but I think the big takeaway is continue to watch the space. While acknowledging that it's a bit buggy right now,
Ali K. Miller argues that there are some big differences here that are fairly significant. Things like each dot getting its own dedicated virtual machine, it being always on, which allows it to be more proactive, and some other changes that are worth watching. In any case, number two on the list is the Decisions API, which is OpenAI's answer to Jeff. Now remember, Jeff is a judgment model. It's not good at outputting text, it's good at classifying things, giving confidence scores between zero and one, making judgments in other words, which are the precursors to decisions, which presumably is where this feature got its name. The Decisions API allows users to call a version of Luna that mimics Jeff's quick classification and decision making abilities. Users to final list of questions and possible answers and the model outputs fast responses. OpenAI said the API could be used for things like classification, routing requests, or choosing an agent's next action. They claimed 10 times faster decision making compared to the responses API. So one of the big questions was how is this better than just using Jeff? One difference in actual use cases comes from Luna's support for visual inputs, which aren't possible with Jeff. Over the past couple of weeks of Jeff Mania,
many users have shown off Jeff making rapid classification for things like visual marketing, but that requires an image to text transformation under the hood. The Decisions API removes that step, and in that way expands its set of possible use cases. Right now, the Decisions API is just in preview for testing with a limited group. And in terms of significance, I think more than anything, this is a recognition that what Jeff represents is more than just a new model. It is an extremely useful and dare I might say soon to be fundamental primitive to have in the ecosystem. Part of why Jeff has hit wasn't that it was flashy or sexy. It's just that as soon as you see all the things that it does better than a traditional LLM. In fact, where you see that LLMs were being rammed like square pegs and round holes into use cases that they weren't great at, it just seems obvious that having that sort of judgment model to sit alongside your generative model is pretty obvious in retrospect. Once again, it's too early to really do comparisons, but I tend to think that this is a space where there is room for more than one model available. In fact, I think that pretty much every frontier lab will have some version of this available very, very soon. Next up on our list is OpenAI Space. Their new shared document workspace that gives human
teams and agents a place to collaborate. You can think of the features sort of like an AI-inhanced Google Drive. Teams can work on shared spreadsheets, slide decks, and other documents, but space also gives teams a place to house workplace automations. Dots can natively work on documents in space, but teams can also set up schedule tasks to produce a deliverable in their space. Many focused on OpenAI going after Microsoft 365, Google Drive, or Notion with space, and while this is certainly generally in those space, each of those services have felt increasingly like a bad fit for agentic work, and it was only a matter of time before OpenAI launched a truly AI native productivity suite. Now whereas almost every other announcement had a pretty wide diversity of opinions, space was one where especially the power users were in love right away. Ray Fernando wrote, I've had early access to OpenAI dots and I'm not here to join in on the glaze fest. Is this a Grockbot or Hermes killer right now? No, but spaces and allowing agents to thrive in apps feels like the right directions for these agents. How IAI's Claire Vowe wrote, in my opinion, dots slightly overhyped
and space underhyped. Every company I know wants an AI native collaboration workspace. We'll be watching to see if this pulls more enterprises to the OpenAI ecosystem. Dan Shipper from every wrote, the obvious win is that you stop switching windows between chat GPT and another app like Notion while you write. The less obvious one is speed. When a dot builds a document through its browser in Google Docs, the edits practically crawl in because Docs wasn't built for agents. Native documents don't have that lag. You can tag your dot inside the document itself instead of going back to the chat and tell it to check the file. Dan concludes, I've long expected the company to do this. In 2024, I wrote about the potential for documents slides and sheets inside chat GPT and I'm glad the product is finally here. I'm already doing most of my work in chat GPT's in app browser and this makes that process smoother. I can tag my dot in a comment on a document and get a revision in line as I'm working. The back and forth feels more like collaboration. The big model launch for the day was GPT 6.1 Sol, which opened AI pitched as near Astra Intelligence for a fifth of the price. The release comes just a week after GPT 6 Sol and provides
some fairly notable upgrades. Coding benchmarks are up significantly across all effort levels, meaning that 6.1 Sol on medium settings outperforms 6.0 on max settings. The models top score on deep suite was 75.2% on high settings, with extra high in max actually seeing a degradation in performance. We first saw this phenomenon with Opus 5 and it looks like top-effort settings are starting to force models to overthink and second-guess correct responses to their detriment. Now that high setting score on deep suite was actually a touch higher than Astra's best performance, supporting open AI's claim of near Astra Intelligence. The pattern was similar across most major benchmarks. A big improvement over 6.0 Sol that landed 6.1 Sol in the same ballpark as Astra at a much cheaper price. One of the notable benchmarks was OS World, which tests long horizon computer use. 6.1 Sol's best performance scored 71.4% on max settings, beating 6 Sol at 64.4% and coming close to Astra's best at 73.5%. 6.1 Sol was also significantly cheaper than the others at a third the cost of 6 Sol and 13% the cost of Astra. What's more, increases in the effort level barely changed the cost,
suggesting that open AI has made some big breakthroughs in the efficiency of computer use with this model. Artificial analysis scored the model at 52, 1.5 Astra and also behind Opus Sonnet Envable coming in in 5th place. They also found incredible cost efficiency that pushes the Pareto Frontier, with 6.1 Sol's benchmark run completed at a quarter the cost of Astra and 31% cheaper than 6 Sol. If you don't need to run it at max settings, 6.1 Sol seems capable of hitting smaller costs like GLM53 Flash, which is the cheaper version of ZAI's open source model. Open AI called it the most cost efficient model for its performance available today. Now as to whether this one was expected or unexpected, on the one hand it's never all that surprising to get a new model from one of these labs, especially on a big day like DevDay. Surprising in that we only got 6 Sol last week? So when it comes to first impressions, let's just say we're going to have to wait to get our hands on it ourselves. Because for every post you can find like this one from Dropout layer on X, GPT-61 Sol just made Opus-5 5 look expensive. You also get one like this one from Bridge Branch. We ran it through our Sunset Ocean test,
same cost as GPT-6 Sol, twice as slow and it barely rendered an ocean. Hello everyone, one big change around AI is we've shifted our thinking from how we rank our pages to how do we become the source that AI trusts enough to answer with. At KPMG they're seeing this first hand. AI generated results now surface answers directly often without a single click. That's why they are increasingly focused on generative engine optimization or GEO, structuring content so AI systems can retrieve it, understand it, and cite it as trusted authority. This is not just an SEO evolution but a visibility mandate. And indeed the GEO mandate from KPMG is simple. If AI is shaping decisions, your expertise needs to show up inside the answer. Read all about it at kpmg.com slash us slash GEO again that is kpmg.com slash us slash GEO. Here's why most legacy modernization projects fail. The AI doing the work can't understand code bases at scale. It sees a small slice of context, examines syntax, and misses years of
decisions distributed across the global application ecosystem. Blitzy solves this the way it solves everything. Grounded in your code before any migration begins, Blitzy's agents reverse engineer the entire legacy system into a persistent knowledge graph. Every dependency, every constraint, every piece of tribal knowledge that used to live in one engineer's head. From that understanding, Blitzy autonomously executes language migrations, framework upgrades, and monolith to microservices transformations all validated and to end. One Blitzy customer modernized a $10 million monolithic insurance stack in 16 weeks against a 137 week baseline with coding agents. That's 9x compression. Retire technical debt while accelerating your roadmap. See how at Blitzy.com. That's B-L-I-T-Z-Y.com. The best teams don't have a single star carrying everyone else. They know their own strengths than each other's weaknesses and play to both. That's the team robots and pencils has built on purpose. Nobody there is grinding through busy work to pat a head count number. People come for the hard problems and they stay because everyone around them is leveling up at the same time. In a market full of companies that are just trying to hire fast, that's worth a look. Check out
robotsandpensals.com slash careers. This episode of the AI Daily Brief is brought to you by Hyper Agent, where you run fleets of agents your team can manage together. Forget local agents and chat workflows waiting on your laptop to be prompted. Hyper Agent deploys always on agents in the cloud, doing real work across the tools your team already uses. Marketing agents turn competitor moves into landing pages. Sales agents in rich leads, draft emails, and updates the CRM. Ops agent chases the paperwork and tracks the budget. Every agent has access to shared context and follows your rules about scope and approvals. It's time you add agents that feel like teammates. Higher yours at Hyper Agent. Get $100 in credits at hyperagent.com slash AI Daily Brief. Hopefully these sole class models are good enough though because we may never see GPT-61 Astra. The Wall Street Journal reports that OpenAI have scrapped plans to release the next version of their flagship model over safety concerns. Sachi Jane OpenAI's head of safety systems set in a statement. For anything regarding safety and alignment, there's a trade-off. You really do need
to find what's the right line between staying within scope but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction. It sounds like OpenAI tried to turn down the tenaciousness that had led to multiple incidents over recent months but couldn't find a happy medium. Jane added that although the model was less lazy than its predecessors, it quote, didn't quite meet the bar in terms of staying within scope and authorization and how it communicates back to the user about the type of work it's done. Now obviously I'm being a bit hyperbolic when I say that we'll never see 6-1 Astra but it actually does seem like they're going to need to hit some technical breakthroughs before we get to some of the next levels of the frontier in a way that they deemed safe enough to release. Now from here we get into what are considered the smaller announcements from the event, many of which people are still noting as quietly significant. OpenAI showed us the latest version of its platform plans, opening Sachi BT to outside developers. Tibo from the OpenAI team wrote, you can now build full native apps with plug-in extensions and ship them right in Sachi BT. We have over 1.2 billion weekly users and will surface relevant plug-ins right in the conversations. This feature is technically called plug-in extensions.
Now OpenAI has been trying to get the balance right on this feature for more than a year, testing native integrations, plug-ins, and MCP is a way to expand the ecosystem. So could this finally be the feature that lets Sachi BT function as an app store, which always seemed like one of the big goals for OpenAI? Certainly there is going to be a rush to experiment here. I'm already seeing things like this personal stylist app from Yana Wellender, but others are wondering if the trade-offs for the app developer are really worth it. MIT's Christian Catalini writes, you bring the app, OpenAI brings the intelligence. Does OpenAI also get the user traces? An interesting partnership if your contribution is teaching your partner how to do your job? To which signal responded, long tail of apps might opt in, but otherwise this is a terrible idea for anyone else. A flip side of this is Sign in with Sachi BT, which is exactly what it sounds like. It allows you to use Sachi BT to sign into other accounts. OpenAI's head-of-applied research Boris Power wrote, Sign in with Sachi BT lets you use any app built with our API. Makes it much better for app developers not needing to jump a huge hill just defying paying a separate subscription. Sachi BT is your one-stop super intelligence juice.
The idea is basically that developers allow people to carry their intelligence subscription with them, lowering the barrier to entry for their apps, and allowing them to take advantage of the fact that so many people have already opted in to using Sachi BT for their super intelligent solution. Jackie Lua writes, Sign in with Sachi BT is a huge deal, and I don't understand why they framed it as off when it's really open AI expanding favorite pricing to third parties and bundling services into the subscription. It finally starts to align their incentives with customers so that individuals stop double paying for tokens and apps stop having to price everything on top of API costs. I'd imagine that expands to many more if not all apps, and Anthropic will need to follow to add value to their subscription. In that world, apps could bypass token costs and charge for the app layer alone again, which is probably good for everyone. In other words, the idea here is that instead of apps having to charge what looks like huge prices to get access to intelligence, they can just charge their $20 or whatever they wanted to price their app experience at and let the API cost flow directly to the underlying. While yes, that means they might not get to scout those costs and add a little premium to them. The benefits of not having to convince
people to pay those additional fees when they're already paying for a Sachi BT subscription, likely outweighs the money that they could make otherwise. For my power users out there, OpenAI has introduced a few new options for their subscriptions. The first is UltraFast mode, which offers 8x faster token production in Codex and 6x speed in the app. Right now, UltraFast mode is only available for Astra and Codex and ChatGPT work, but OpenAI say they will add support for 6.1 soles soon. In addition, OpenAI is introducing a new $500 subscription tier that offers 25 times the usage of the plus tier. It's also the only tier that gets access to UltraFast mode for the time being. Heading into the event, Tibo announced that OpenAI would reopen access to their $200 a month pro tier, which they had turned off a few weeks ago, but with a few tweaks. Usage calculations have been changed, netting out to a 50% reduction in terms of API cost. Tibo explained that this was the best of a bad set of choices. OpenAI prioritized not having to introduce the 5-hour usage limit and argued that API cost reductions and more efficient models would mean users can still get roughly the same amount of work done.
Now, on the one hand, people were sort of expecting something like this, but at the same time, you can imagine how well it went over to have the same model coming back with reduced value. For our purposes here of trying to understand what it says about the state of AI, look man, compute constraints are a real present and permanent. Frontier models are running up against the walls of what they can do with the compute that we have, and new compute isn't coming online at anywhere near the speed that people are increasing their use of this digital intelligence. The upside of that, though, is that we're going to see a big push around efficiency, which in the long run should net out to more cost-effective better experiences for all of us, but it's going to have some bumps along the way. Over in developer and vibe coder land, Codex now has a dedicated cloud environment, allowing users to keep working after they closed their laptop. This one was always coming after people walking around with their thumbs jammed into laptops, became so common across San Francisco that it became a meme. But looking for broader patterns starting with Gropbot, we've seen the entire industry move towards a cloud instance being a necessary part of all-agentic products. This is a continuation in the big shift from AI being something you engage with like software to becoming a more persistent always on application
regardless of how you access it. The Codex CLI is also getting a major refresh with a new look and new capabilities. History goes back further as you scroll supporting the monothread maxis, of which I am absolutely one, while the composer stays pinned so you can always type your next line. Wrote OpenAI, the Codex CLI now gives you a better way to manage parallel work. Use slash agents to see what's running, check progress, and jump between tasks. On the enterprise side, OpenAI has launched private intelligence, which guarantees zero data retention even at inference time. Over the past few months, privacy, which was always a huge issue for enterprises, has become a huge issue for the companies supplying the enterprises, with growing concerns that OpenAI and Anthropic are skimming data from their customers. This is basically the subtext or the main text of every communication for Microsoft these days about why you should be not trusting their competitors. All that means that a stronger, clearer, and-to-end data privacy guarantee is a big deal for those who need it. OpenAI has also launched a model marketplace. It allows users to buy open-weight model inference from base 10 through OpenAI's responses API and Codex.
Now this is one that on another day we could go way deep into because it has some big implications for OpenAI's mode. It protects them from open source disruption and gives customers a lot more choice. It also means enterprises can feel comfortable making large spending commitments with OpenAI, knowing that they can easily use that spend on a multi-model strategy that includes open-weight trust me when I say that we are going to come back to that one because I think it is bigger than people are giving a credit for it at first blush. So we have barely scratched the surface here, but let's try to quickly sum up what all of this amounts to and what OpenAI DevDay revealed about the next phase of AI. We didn't get something blisteringly new. What we got was confirmation of a lot of the trends that we've been cataloging on this show over the past several months. First, with Soul 6.1 in the Decisions API, we got a continuation and a deepening of the trend of models that are good enough and cheap enough that they can be used to do everything. IE, you don't have to turn them off, you don't have to make decisions about what they're used for, they can just do it all. With dots, we have persistent and proactive agents that take advantage of that to actually do
everything. And when it comes to doing everything, plug-in extensions shows that that means bringing everything in as well as going everywhere, which is chatchipt sign on for other apps. Finally, with space, we have more confirmation that the next generation of AI will not just be single player mode, but will also be team and multiplayer mode in a native way, even if that remains nascent so far. Lots and lots more to explore here, but that is going to do it for today's AI Daily Brief. Appreciate you listening or watching as always, and until next time, peace!
More episodes
More from The AI Daily Brief: Artificial Intelligence News and Analysis

Gemini 4 Argon, Sonnet 5.5 and What Matters with AI Models
The AI Daily Brief: Artificial Intelligence News and Analysis

How to Build Team Agents
The AI Daily Brief: Artificial Intelligence News and Analysis

The Real Risks of AI Agents
The AI Daily Brief: Artificial Intelligence News and Analysis

The Rise of the AI Moderates
The AI Daily Brief: Artificial Intelligence News and Analysis