
Gemini 4 Argon, Sonnet 5.5 and What Matters with AI Models
Get every episode summarized
Each time The AI Daily Brief: Artificial Intelligence News and Analysis publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
The AI Daily Brief: Artificial Intelligence News and Analysis is made possible by:
All brands on The AI Daily Brief: Artificial Intelligence News and Analysis →
“Coming into 2026, Google was looking pretty good in the AI race. A lot of the questions inside DeepMind had been answered. They were putting out competitive Gemini models. They were pushing forward with interesting new products.”From the transcript
Google’s Gemini 4 Argon impresses on benchmarks, but does it signal a real comeback? NLW examines Gemini 4, Claude Sonnet 5.5, and what Muse versus Dots reveals about choosing AI tools. In the headlines: Trump’s super intelligence accord, the America.gov AI portal, and an FTC investigation into OpenAI and Anthropic’s rogue agents.
Next Cohort - Learn How to Build Agents - https://register.besuper.ai/register?program=ati
AIDB Fall Listener Survey - https://aidailybrief.ai/survey
Multiplayer AI Sprint - https://multiplayerai.ai/
Brought to you by:
KPMG – Research from KPMG and the University of Texas at Austin shows the highest-impact AI users treat AI like a reasoning partner — and those skills can be taught at scale. Learn more at https://kpmg.com/us/Sophisticated
Harbor - Invest in the AI ecosystem. https://www.harborcapital.com/aidaily
Hyperagent - Hire a team of always-on agents. New users get $100 in free credits. hyperagent.com/aidailybrief
Rackspace Technology- One accountable partner to build, operate and run your full enterprise AI stack https://www.rackspace.com/
Section - Section turns AI investment into workforce transformation and ROI - https://www.sectionai.com/
Blitzy - Want to accelerate enterprise software development velocity by 5x? https://blitzy.com/
The AI Daily Brief helps you understand the most important news and discussions in AI.
Interested in sponsoring the show? [email protected]
Get every episode summarized
Each time The AI Daily Brief: Artificial Intelligence News and Analysis publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
794 searchable segments. Every word is indexed and playable.
Full transcript
The AI Daily Brief: Artificial Intelligence News and Analysis — Gemini 4 Argon, Sonnet 5.5 and What Matters with AI Models. Machine-transcribed; use the interactive transcript above to jump the player to any line.
Coming into 2026, Google was looking pretty good in the AI race. 2025 had been a good year. A lot of the questions inside DeepMind had been answered. They were putting out competitive Gemini models. They were pushing forward with interesting new products. And all the natural advantages that they had always had in terms of consumer distribution and data and all those things had a lot of people very bullish on them as a contender. But then vibe coding happened. And with the increase in coding capabilities paired with the power of the new harnesses like Cloud Code and Codex, those new capabilities unlocked agents in a way that hadn't been possible before. And all of a sudden, Google found itself very behind. Charitably, you would say that in 2026, the company has been playing catch up. But even that's not really accurate to what's been happening. For most of this year, Google has been firmly outside of the conversation as a top model lab. And yet I think that those who didn't have a partisan bias towards one of the other labs would never be fully comfortable writing Google off. This week, the company announced Gemini 4, their first new frontier model in more than six months. By the benchmarks, it looks like Google is so back.
But is that the whole story? So let's dig in to Gemini 4, Argonne, and where it lands in the current AI race. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. [♪ OUTRO MUSIC PLAYING [♪ Oh, all right, friends. Quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, robots and pencils, Harbor, and granola. To get an ad-free version of the show, go to patreon.com, such as AI Daily Brief, or you can subscribe and Apple podcasts to learn more about sponsoring the show. Send us a note at sponsors at aidealybrief.ai. Also note that we have our next cohorts of superintelligence agent training coming up. These are paid programs, which in very short order will get you far ahead when it comes to your agent to understanding and your ability to use agents in your daily work. We have both the executive catch-up program and the agent intensive, which we call the executive agent leadership program. The next cohorts for those start next week. And you can find links to all of that at the very top of aidealybrief.ai. President Trump really looked at OpenAI Dev Day
and said, nope, absolutely not. I don't want that to be the biggest thing happening in AI this week. And invited basically every big AI CEO to the White House for what David Sacks would later call the Bretton Woods of AI. Now, there was a lot of chatter that came out of this meeting. But one of the first things that people noticed was that Trump pulled in Thropic CEO Dario Amade to be the one to speak to the press following the meeting. Now Dario insisted that he's just saying what he's always said, that AI has some incredible benefits and some very real risks. However, in an act which some saw as public difference, other saw as Dario finally playing the game he commented, as the president has said, whoever wins AI wins. I think that's very important. The mechanism how we address the risks is still under discussion. But we all need to work together to make sure that we can win and we can win safely. If we do this right and work with the president and everyone here, we can all win safely. Now, X was absolutely filled with people psychoanalyzing the moment, breaking down the expressions from other tech leaders, Dario's anxious ticks, and even dissecting his fashion sense.
While some thought that Trump and the other tech leaders might be bullying Dario as he stepped forward to speak, the president followed up with some kind words later in the press conference saying, Dario has been fantastic. We had dinner the other night and he agrees with everyone. Whether or not this was Dario truly bending the knee, it was a significant moment for a CEO who previously told staff that Anthropic had been targeted for refusing to give, quote, dictator-style praise to Trump. At least in the immediate wake, the press conference completely overshadowed the subject of the meeting, which was the head of the frontier lab signing an accord on super intelligence. The accord is a one-page commitment to develop internal controls for frontier model testing and deployment, as well as partnering with external auditors for verification of those controls. Trump told the press that the document is, quote, morally binding, suggesting a continuation of the voluntary approach to AI safety testing. Trump added, the biggest people in the world signed that and I signed it as president and it really is a form of protection. He also floated the idea of a 10-member industry oversight committee during the press conference, but for now seemed satisfied commenting, I'm seeing tremendous self-policing.
The general tone at least among the people at the meeting was that this was an important first step. Mark Zuckerberg told the press, the idea isn't that this is the only thing we will ever do, it's that this is a start and an accord, the whole industry could come to. Mostly the media response to this was a complete rush act test about how one feels about Trump and how one feels about the tech CEOs, i.e., whether you're inclined to use the word oligarchs to describe them, but I do think that the skeptical take on this one might be best summed up by media commentator Chuck Todd, who wrote, if you trust the tech companies to police themselves and you want them to decide how AI will impact society without any input from the public, then you'll be fine with this accord. Trying to find the AI moderator, AI realist position in this one, I think a person could agree that it's better that these folks are frequently in rooms together with key leaders in the government, rather than exclusively viewing their role through the lens of competition, while also wanting to see more external involvement, not least of which an articulation of what the independent external auditor or evaluator to carry out independent assessments is actually going to look like in practice. Perhaps now that this accord has been signed,
that particular detail is one that can get built out. Following the lunchtime meeting, President Trump signed a pair of executive orders marking the beginning of the super intelligence age. The first made Trump's name change official, stating, as these capabilities continue to improve, they increasingly represent not merely artificial intelligence, but a new era of super intelligence. The terminology used by the federal government should reflect the transformative capabilities of these technologies, and the limitless opportunities they create for the American people. Much more functional was the second order, which established America.gov as an AI powered single portal for government services. Trump said, with America.gov, the federal government no longer stands in your way, it only stands at your service. We're simplifying it, we're glamorizing it, we're making it what it should be. Alongside that second executive order, we got an Apple-style product unveiling keynote, complete with Secretary of State Marco Rubio playing the role of Steve Jobs. Rubio showed demos of plan future functionality, like being able to use the portal to apply for a passport, change your name after getting married, or enroll for Medicare. The features kind of seem like an MCP for government,
a user can provide their information once, then the America.gov agent will track down and complete all the necessary forms across numerous government departments. How well this works remains to be seen, but certainly it is a real problem. The insane difficulty for example, of changing your name after getting married, depending on which state you live in, is hard to overstate. Rubio said that the functionality will be available from next year, with the executive order giving government departments 90 days to integrate their services into the portal. For now, America.gov only functions as a knowledge base for government services, searching across 29,000 government websites to answer questions based on official sources. Now, a lot of the chatter on this centered around security and privacy concerns, particularly after the president said, it cannot in theory be hacked into, and when they figure out a way to do that, we'll end it because they always figure out a way, right? The National Design Studio, which designed the website, came under fire in June for including activity tracking tools and other websites they've designed. And while some raised questions about providing the personal information required to apply for government services, AI jailbreaker Pliny Deliberator poked and prodded at the chatbot and found its guardrails were pretty airtight.
He wrote, interesting, the America.gov chatbot will flag anything that appears to be personal information, like any text vaguely resembling an address or password, and refused to accept the query until the flag text is removed. Never seen that before. Meanwhile, if you think that them showing up together with this meeting means these companies are going to escape all government scrutiny, it might be worth noting that the FTC has opened an investigation into open AI in Anthropics rogue agents. According to agency sources speaking with the New York Post, the FTC opened an investigation into the two leading AI labs a few weeks ago. The scope will cover the hugging phase incident and the dozens of so-called rogue agent events that have been disclosed since. The investigation has advanced to the stage that the FTC is drafting civil investigation demands, which are similar to subpoenas. An FTC source said, the agency has plans to compel the executives at these firms to testify about their product and about the dangers they allege their products may have to consumers to Americans. Reports state that third party safety research lab meter can expect to demand alongside the frontier labs. Now from very early on in the AI narrative, one big thread of discourse has been focused on the idea
that new AI-specific regulation wasn't necessary. That government could simply enforce product safety legislation that's already on the books. Former FTC chair Lena Khan was of this view writing last month. Law enforcers already have authority to charge companies in their CEOs for creating and releasing dangerous unvedited or defective products. We shouldn't let discussions about new legal regimes to extract from the fact that there's no AI exemption from laws already on the books. This post was co-signed by former AI's art David Sachs and yet another example of the AI safety debate creating very strange bedfellows. Still based on comments from the FTC source, it's not clear that the agency is necessarily pushing for some sort of harsh penalty. They told the New York Post, we need to win this SIRACE absolutely and we are winning and that's fantastic. Getting a political jab in the source continued and obviously the other side, the Democrat party wants to destroy this technology, wants to surrender to our enemies and competitors and that's not something that the chairman's interested in doing and that's not really something the president's interested in doing. With that being said, the laws have to be followed. People talk about how we need new laws for these companies and the chairman's been very clear. We have plenty of laws on the books.
So an investigation, that's good, right? That'll appease people who think that the labs are getting a free pass. That's not exactly what the discussion suggests. Anti-monopoly activist Matt Stoller wrote, it's a protection racket. The ideas Trump investigates and clears them so other investigators have hurdles in their probes. Former director of public affairs at the FTC, Douglas Farrow wrote, it would be good if the FTC investigated these companies properly, better starting two years ago, but the public credibility of this new investigation is 0.0%, especially after yesterday's AICO and Trump love fest. The optimistic take comes from Joel Thayer, who writes, I kind of love this trust but verify strategy. Even though it's mostly self-regulatory, if any signatories fall short of these commitments, it could trigger FTC enforcement under section five of its UDAP authority. Makes FTC chairman Andrew Ferguson's presence at the meeting very apt. As with anything policy right now, it's worth having a whole bowl of salt when you look at this, but things continue to move in some sort of direction. For now that's gonna do it for the headlines, next up the main episode.
If you're leading AI inside an enterprise, you already know that the gap right now isn't capability but execution. That's why KPMG's you can with AI is back with a new season featuring conversations with leaders like Serojit Chatterjee of Emma, Mahabibov Rider, McKesson CIO, Ellery Fisher, and others focused on practical execution. What's working, what's not, and what it actually takes to move from pilots to real scaled impact across strategy, data readiness, governance, workforce, and value. And of course, it's co-hosted by me, Nathaniel Wittemore. Go listen and subscribe at www.kPMG.us slash AI podcasts. That's www.kPMG.us slash AI podcasts. The best teams don't have a single star carrying everyone else. They know their own strengths than each other's weaknesses and play to both. That's the team, robots and pencils has built on purpose. Nobody there is grinding through busy work to pad a head count number. People come for the hard problems and they stay because everyone around them is leveling up at the same time. In a market full of companies that are just trying to hire fast, that's worth a look. Check out robotsandpensals.com slash careers.
If you listen to this show, you likely have a thesis. Maybe it's enterprise adoption, maybe it's compute, maybe it's a specific lab. Harbor Capital's AI lab ecosystem ETFs let you express it via five actively managed ETFs, each seeking exposure to the ecosystem around one major lab, Anthropic, OpenAI, DeepMind, Meta or SpaceX AI. Your view of the AI race in ETF form. Harbor Capital advisors AI lab ecosystem ETF suite gives investors a way to invest in the AI ecosystem they believe is best position for success. Search Harbor AI lab ecosystems ETFs wherever you invest or follow at Harbor Capital on X to learn more. Visit harburetcapital.com for a perspective by maintaining investment objectives, risks, fees, expenses and other important information. Read and considerate carefully before investing. Risks include principal loss and artificial intelligence related risks. Harbor ETFs are distributed by four side fund services LLC. Harbor is not affiliated with AI Daily Brief and the funds are not affiliated with sponsored by or endorsed by any AI lab. This is a paid advertisement and not personalized investment advice. Investing involves risk, including possible loss of principle. When I'm in a meeting, I'm fully in it.
I'm thinking about the iteration and create it back and forth. It takes to actually push a goal forward. What I'm not thinking about is capturing takeaways, tracking to do's or any of that. And that's where granola comes in. Granola is an AI powered notepad that captures what happens in your meetings and turns it into clean structured notes with the decisions and action items pulled out and easy to find. There's no setup and no configuration. It just fits into how you already work. For me, it means I get to stay in ID mode and granola makes sure those ideas actually become action. Once you try granola on a first meeting, it is hard to go without. You can try it totally free at granola.ai slash brief. That's granola.ai slash brief to get your time back. Welcome back to the AI Daily Brief. And friends, it appears that pigs are flying. Hell has frozen over. Choose your metaphor for incredulity because after months of waiting, Google has announced Gemini 4. Although maybe those pigs aren't completely off the ground because although they announced the model, we don't actually have access to it yet.
On Wednesday, Google unveiled Gemini 4 Argon. It's their first frontier model in over six months. Now the 2026 struggles of Google have been well-documented. We never got a Gemini 3.5 Pro with the company basically deciding that it just wasn't good enough to release. And for a few weeks now, speculation has been coming that Gemini 4 was on the horizon and Google DeepMind CEO, Karei Kevaso Glouz, says that the model delivers, quote, Frontier performance in complex workflows across real world software engineering, enterprise knowledge work like legal and finance and cybersecurity defense. Now despite a long period where Google's future as a Frontier Lab has seemed somewhat tentative, to judge only by the reported benchmarks, Google looks to be right back in the game. Gemini 4 Argon is the new state-of-the-art model across many benchmarks. In a giant technology work, Gemini 4 scored 68.9% on the Vals Index, beating Fable 5.1, GPT-6 Astra, and exceeding the score from the current leader Opus 5.5 by a couple of percentage points. The gap was even larger on Automation Bench
and Vals Finance Agent. And on Harvey's legal agent benchmark, Gemini 4 more than tripled the score of current leader Fable 5.1 scoring 19.6%. Agente coding scores were a little more patchy, but Gemini 4 is definitely in the ballpark of Frontier rivals. It's the new state-of-the-art on DeepSwee scoring 77.9% compared to Opus 5.5, 74.2%. On Frontier Swee, Gemini 4 scored 55%, which placed at 10 points behind current leader Astra 6, 7 points behind Opus 5.5, and 1.3 points behind Fable 5.1. On Frontier Swee, however, Gemini 4 scored 55%, which placed at 10 points behind current leader Astra 6, 7 points behind Opus 5.5, and 1.3 points behind Fable 5.1. It was also bottom of the pack on Terminal Bench 4.0 with a score of 57.4%, 9 points behind current leader Opus 5.5. Now you can look at this in two ways. One agent-of-coding is one of the most important, if not the most important use case, so failing to achieve a state-of-the-art score is kind of problematic, or you can look at this as a huge improvement for Google, which it undeniably is.
Coding had been Gemini's biggest weak point for the past year, so to have Gemini 4 be in the same ballpark as the Frontier models from OpenAI and Anthropic, absolutely represents a big catch-up. For computer use, Gemini 4 is nipping on the heels of Astra, which set a new standard for the category. On OS World, it scored 69.2% against Astra's 72.6%. Artificial analysis confirmed that Google is back in the mix as a leading model developer. Gemini 4 Argon scored 53 on the Intelligence Index, putting it tied for third place with GPT-6 Astra and Fable 5.1, one point ahead of GPT-6.1 sole, and three and five points behind Sonnet 5.5 and Opus 5.5 respectively. In terms of cost and efficiency, it's pretty close to the Pareto Frontier, costing $1.99 per task on the AA benchmark run, compared to 326 for Astra and 763 for Fable 5.1. However, that is with Google's 50% launch discount, which will be available for an unstated period, meaning that the full price would make it more expensive than Astra. It's also in that uncomfortable middle ground where it's more than twice as expensive as GPT-6.1 sole,
even with the discount without a super clear improvement on the benchmarks. For the moment, however, none of this really matters, because unfortunately, Google is not making Gemini 4 Argon generally available to the public. Their stated reasoning has to do with cybersecurity. The Model scored 68% on the CWE Bench Cybersecurity Benchmark, tied with GROC 47 and GPT-6 Astra, and slightly ahead of Opus 5.5, Fable 5.1 and Mythos 5.1. This was reason enough, according to Google, to limit access to what they call a set of trusted cyber defenders to ensure the model is not misaligned. They said that they're engaging with the US government's voluntary testing and pre-release access program and intend to gradually expand access over time, beginning with API and Google ultra-subscribers. In their blog post, Google said that the model is already powering internal workflows and boasting strong performance on long-horizon coding tasks like large-scale code-based migrations. Now, as to the choice to announce the model without actually releasing it, I'm not really sure. Honestly, for us old heads who have been watching this for a while, it kind of has echoes of the first time Google fell behind
and felt the need back in December of 2023 to announce Gemini, even though the pro version of the model wouldn't be available for a number of months. Yet, maybe it was a good call because there were plenty of people who were excited to see it. There was an entire genre of post on X that basically came down to Google so back. Dr. Daryl Anouk Mazrout, wow, wow, Google DeepMind is back with Gemini 4 Argon, insane benchmarks. In one fell swoop, they've jumped to the very top of the AI frontier, as I've said before, never bet against Google in the age of AI. Although he did add, and this seems to me to be the key detail, can't wait to try Gemini 4 ASAP. Ethan Mollock writes, and it's a three-way race again. Nathan Lambert, who does open model research and is about as far away as you could be from a hypester, writes, love to see Google surprising people with Gemini 4. More labs at the frontier is wonderful for consumers because of competition and wonderful for the world because of reduction in concentration of power. Excited to see how it goes in real world scenarios. AI Leaker, I rule the world, writes, I've been critical of Google models, but Gemini 4 looks strong on benchmarks.
We've seen this before with Google, but I'm optimistic they can release a first class model with a first class harness like Codex. Great work to all involve with Gemini 4 rooting for you. But then on the other side of the coin is Angel Who writes, I'm sorry, but I can't with all these Google is so backposts. We literally can't even use the model, and it's not like benchmark scores for Google's models have ever been reliable. Do we really need to remember what happened with Gemini 3 Pro? Seriously? Issue Agerwal put some numbers around that, sharing an older series of benchmarks and adding. By the way, these were the official benchmarks for Gemini 3.1 Pro when it released, practically destroying Opus 4.6 across the board. I'm not saying Gemini 4, argon will be bad, but it's crazy that no one has the slightest bit of skepticism given Google's track record. To their credit, Logan Kill Patrick, who honestly should get Google's MVP for hanging in there throughout all of these PR challenges this year, actually engaged with this writing, we've gotten much better at testing our models at scale across Google now. So assume most new Gemini revs go through thousands of software engineers for weeks before getting released, hopefully has helped close the benchmark
to reality gap by a real margin. Now so far, all we have to go on is the reported benchmarks, along with some reports from grumbling inside the company. A Bloomberg piece called Google Grapples with Employees skepticism about new Gemini 4 says, While Gemini 4 has performed well on benchmarks widely used to gauge model efficiency, it does less well when employees actually put it to work. The model struggles to handle certain coding tasks according to people with direct access to the model. Google for their part denied the reporting, commenting that it would be inaccurate to claim that Gemini 4 is underperforming in some areas such as coding. And one of Bloomberg's other sources said that these grippers are a minority, claiming that a quote large consensus internally at the company believe that Gemini 4 is the frontier. What makes the position of Google so challenging is that in the time that they've been away from the top, the nature of the race has fundamentally changed. With much optimism Peter Yang wrote, Google cooked on Gemini 4. Now they just need to be more competitive on coding harness, anti-gravity, and personal agent Spark. In other words, the competition has become more than just models.
And it's not like there already isn't a ton of model competition. One model released in advance of OpenAI DevDay that we didn't have a chance to talk about yet is Claude Sonnet 5.5. And it's certainly if you feel yourself giving a bit of an eye roll or at least a raised eyebrow wondering if you should actually care, the recent history of the Sonnet and Opus series certainly created some reason for skepticism. Opus 5, you remember, was strongly disliked. And Sonnet 5 seemed to deliver even worse results for a higher price because it was so token hungry. However, if you had to pick just one theme from the last couple of weeks on the show, it's the beloved return to form for Anthropic with Opus 5.5. So does Sonnet 5.5 continue that trajectory? The short answer seems to be in most people's estimation, yes. Anthropic said that Sonnet 5.5 is 30% faster and 30% cheaper than Sonnet 5. And the model demonstrates some significant jumps on the benchmarks. For agent to coding, it scored 70.6% on terminal bench 4.0, which is up from just 10.3% for Sonnet 5. And even beats Opus 5.5 at 66.4%.
Scores on frontier code and cursor bench were very slightly behind Opus 5.5. The same was true for knowledge work benchmarks, GDPVAL AA and NA briefcase, where Sonnet 5.5 closed the previously massive gap between Sonnet and OpenClass models. Artificial analysis affirmed the benchmarks, giving Sonnet 5.5 an intelligence index score of 56 that puts it in second place behind Opus 5.5 and somehow ahead of Fable 5.1 and GPT-6 Astra. However, they didn't find that Anthropics claims about cost savings stood during their testing. Sonnet 5.5 cost $7.60 per task and used significantly more tokens than Sonnet 5 at the same per token price. This meant that Sonnet 5.5's benchmark run was almost as expensive as Fable 5.1, 27% more expensive than Opus 5.5, and more than twice as expensive as GPT-6 Astra. Now to give Anthropic the benefit of the doubt, they claimed a 30% cost reduction on similar tasks compared to Sonnet. The main AA benchmark is run at max inference settings, but turning down the settings to extra high
cut the cost by two thirds. Still, the model impressed once people got their hands on it. It is a massive improvement over Sonnet 5 in the rendering tests that grab attention on social media. Matthew Berman made a series of visual tests and game clones writing. Sonnet 5.5 basically Opus 5.5, but 50% cheaper and much faster. I've been early testing it and it's incredible. If this is pacing the frontier, sign me up. But for real world use, Builder Kun Chen found it performed best when it's the right tool for the job. He wrote, Sonnet is definitely not the same as Opus just 50% cheaper. That's true only for problems that don't need much wisdom. There's a clear difference when they are asked to propose plans for ambiguous product problems. Opus is able to approach problems with more strategic thinking, i.e. what's the real goal here? While Sonnet is more just looking at tactically, how do I get this done? Now Kun's current approach is to use Anthropic models as a complete system with Opus for planning, Sonnet for implementation, and calling on Fable when things go sideways. Pavel Heron tested Sonnet on his bug fix benchmark
and found it outperformed all other models including Fable and Astra. He wrote, The secret? Sonnet 5.5 Max might be cheap and fast for most tasks, but it's the least lazy model I tested. It leads in bug hunting by spending turns. Fascinatingly, YouTuber and entrepreneur Theos review was that Sonnet 5.5 is an incredible model that you probably shouldn't use. In his view, the model is a big improvement over Sonnet 5, but on Max reasoning it over things and risks getting things wrong because of it. And on lower settings, there's no point where Sonnet is more cost effective than Opus because of how token hungry it is. This changes a little when using Sonnet as a sub agent, where it can be more efficient on certain tasks, but it's usually not optimal as a standalone model. So will these results stand after a couple weeks of testing? With Sonnet 5.5 being a technical marvel given that it can produce the results of Fable 5 just four months after the release, but being too token hungry to make sense for most, that's certainly what I'll be watching and I'll report back as people get more reps in. Still going back to the new world that Gemini 4 has to compete in,
the point that Peter Yang again was making is that just being good on the model isn't enough. In fact, one really interesting question brought up by the success of Muse is whether an AI product with a great user experience, but a less than state of the art model beats a product with a less good user experience, but a state of the art model. Muse has hit three million weekly active users. Daily users, who send at least one prompt per day, have reached one million. Now remember, the product has only been available for a little over three weeks, so this is a strong positive early indication. At the end of the first week, Muse had 500,000 weekly active users and 250,000 daily users. Now on the one hand, this is a tiny fraction of Meta's distribution potential, which reaches half the people on Earth and is nowhere near as strong as Chatsy-Bee-T's growth, but a better comparison might be the adoption curve for Codex. Codex took about three months to reach three million weekly active users, so Muse currently looks like a faster growing AI app. Now one might expect that, given that it is a personal agent versus Codex, very developer-focused audience,
but whatever the case, the early success of Muse continues on. And yet now that OpenAI's dots is here, we're going to start to have some direct answers to this question of good UX plus worse model or worse UX plus better model. Yerika Co-Founder seen her rights. After using Muse for a while and loving it, I'm starting to notice the underlying model isn't smart enough. I'm not sure it's context overload, but it gets stuff wrong a lot recently. Maybe they've had such huge success that they've had to switch to a different cheaper, worse model. Anywho, Muse is great, form factor is great, but as a consumer, my loyalty is to none. If OpenAI will drop Muse with better models, I will use whatever is best. Intelligence is still the most. And yet on the flip side, a couple days in, I'm definitely seeing some gripes with the way that dots works. How BBs slap on X rights? Overall, the vibes are not good. This feels directionally wrong. I just want to talk to the model. This tool wants you to be a plumber in a house where you're not allowed to touch the pipes. That's has so much friction built in that it feels like it's not meant to be used. And unless you think that this is just a general gripe
or account, they then go on to include about a dozen examples. The big one that I've seen repeated comes in their final thoughts when they write. They've launched a new product that wants to help you manage your work and chat GPT and codex without having sufficient access to any of the tools and data it needs from those surfaces. I am certainly at least not even close to being willing to declare which is the winning strategy between focusing on the model versus focusing on the product user experience. One last note is we look across the state of the competition that Google finds itself in as it gets ready to release Gemini 4. In one surprise launch this week, DoorDash has announced that their AI agent is getting the Muse treatment and can now take orders via text. DoorDash already had an agent on their site, but they're using some of the features of personal agents like Muse to take it a step further. Customers can now text the DoorDash agent with a natural language prompt and the agent will figure out the rest. DoorDash gave the example of a customer texting, order my usual protein bowl to the office. The agent that matches the phone number to their user profile looks up the usual order and delivery address then handles the payment. Now in light of Amazon's decision to block Muse agents,
there's a huge open question of whether the best agent strategy is to partner with the category leaders or build a platform-specific agent. The logic for an agent from DoorDash seems pretty clear at this point. The company has said that agentic orders have a 50% higher basket value in buying groceries. And by building their own, DoorDash can attempt to hold the user relationship close. There's a risk, of course, that external agents like Muse could order directly from restaurants and cut DoorDash out as a middleman. At the same time, user preference might reject platform-specific agents because they have much less control. It's unclear what the optimal strategy will be, but for now, DoorDash seems to be doing a bit of both, keeping their platform open to third-party agents while also building an internal agent and for that reason it's worth paying attention to. Even if you don't much care who wins, how you order lunch to the office. Anyways, bringing it back to Gemini IV, Argonne, I'm with Nathan Lambert. When he argues that Google doing well is both better for consumers and better for competition. So put me firmly in the camp of people who are hoping that this release is great. Let's just hope we get it soon so we can actually decide for ourselves rather than dining from the table scraps
of self-reported benchmarks and gripey employees talking to Bloomberg. That's gonna do it for today's AI Daily Brief. Appreciate you listening or watching, as always. And until next time, peace. Made better team. 是我 Die Wanted. Learn So Much oined Yijiang
More episodes
More from The AI Daily Brief: Artificial Intelligence News and Analysis

The Most Important New AI Tools from OpenAI DevDay
The AI Daily Brief: Artificial Intelligence News and Analysis

How to Build Team Agents
The AI Daily Brief: Artificial Intelligence News and Analysis

The Real Risks of AI Agents
The AI Daily Brief: Artificial Intelligence News and Analysis

The Rise of the AI Moderates
The AI Daily Brief: Artificial Intelligence News and Analysis