
Get every episode summarized
Each time The Daily AI Show publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
“Yes, if you've listened to me, you know, it's the show. I am surprised every time I say a number because it just keeps getting bigger. I had the same experience every time I realize how old I am. So you are watching or listening to the daily AI show.”From the transcript
The episode focused heavily on the shifting competition between OpenAI and Anthropic. Data discussed from Ramp showed Astra accounting for 13 percent of tracked enterprise AI spending versus 8 percent for Claude, while OpenRouter reportedly saw OpenAI models lead Anthropic in spending for the first time in more than two years. That came alongside discussion that Anthropic may be preparing another model release as OpenAI, Anthropic and xAI all appear to have major launches waiting. The hosts also examined why existing models sometimes behave differently before releases, including a bizarre Gemini 3.8 Flash hallucination and the possibility that compute gets reallocated during rollouts. Other topics included Meta Muse and Instinct personal agents, AI governance, a robot-safety benchmark, an erroneous AI-generated military intelligence report, and UMG and Sony’s latest lawsuit against Suno over training data.
Key Points Discussed
00:04:56 Meta Muse Surges After Launch
00:08:02 AI Governance And U.S.-China Coordination
00:11:19 Independent Evaluators For Frontier AI
00:16:06 Muse Versus Instinct Personal Agents
00:19:47 Testing AI Safety In Physical Robots
00:22:29 AI-Generated Intelligence Nearly Triggers A Military Response
00:25:51 Do LLMs Actually Understand The Physical World?
00:28:19 Astra Versus Claude In Enterprise Adoption
00:30:27 OpenAI Passes Anthropic On OpenRouter Spending
00:31:08 Is Anthropic Preparing Its Next Model?
00:32:30 Multiple Frontier Model Releases May Be Coming
00:36:52 How Astra Banked Resets Actually Work
00:37:00 Gemini 3.8 Flash Hallucinates Its Way Through Hockey History
00:40:52 Is A Stealth Gemini Model Already Being Tested?
00:42:26 Why Current Models Get Weird Before New Releases
00:50:08 UMG And Sony Sue Suno Again
00:54:00 The Fight Over AI Training And Creative Labor
00:59:28 Episode Wrap-Up
The Daily AI Show Co Hosts: Beth Lyons, Andy Halliday, Gareth Hood.
Get every episode summarized
Each time The Daily AI Show publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
544 searchable segments. Every word is indexed and playable.
Full transcript
The Daily AI Show — Meta Muse Surges After Launch. Machine-transcribed; use the interactive transcript above to jump the player to any line.
Hey, good Monday morning everybody. So happy to see everyone. Oh, so happy to see Andy. Today is Monday. September 21st and it is Episode 816. 816. Yes, if you've listened to me, you know, it's the show. I am surprised every time I say a number because it just keeps getting bigger. Which of course is what happens. Andy. I had the same experience every time I realize how old I am. Oh, that's, uh, yep. Okay. So you are watching or listening to the daily AI show. I am Beth Lyons with me in the studio today is handy. Andy, handy, handy, handy, handy, there we go. Good morning. And we don't know if other people will join us. Uh, you will discover when we do as well. Andy. Um, so one of the things that I noticed was that the newsletters that came this morning.
We're talking about the hack of open AI that we talked about on Friday. Yeah, we're always ahead of the game. We're like break it. Hey, you know, we're not just regurgitating our reactions to stories that we read in the newsletters. So we do actually, we're not just regurgitating, we're reacting with all of you. Because that's what, uh, this show is about is, is like, hey, we're, we're reacting in real time, just like y'all are. But it's very cool when we cover the news before, um, our favorite newsletters too. Okay. So what caught my attention over the weekend was the reaction and the metrics around news. Meta's new, uh, personalized agent that's accessible, uh, in, and, and integrates with all of the different meta platforms. So it, it overtook chat GPT as the number one app on the US app store.
And that was just a week after its launch. And as a, what, what I took away from the weekend was watching a gushing, a promotion, but unpaid promotion by Claire Vaux, who is, who is how I AI and a really good influencer in the world of AI, who has tried all the different personal agents. And really was very, very pleased and, and I almost proud of what was possible with news, as opposed to the others that she's worked with. So that's a big endorsement in my view. And then I also saw, um, I think, uh, I forget in one of the newsletters this morning, there was a pretty complete comparison of each of the different agents, whether that's, you know, clawed code or its, uh, codex or mues, doing personal tasks.
So it would have been done in co work on clog, for example, but doing personal tasks, uh, and, and aside by side comparison, in which mues was the only one that successfully completed the task. And each of the others failed in different ways. So I think this, the takeaway is that meta working diligently behind the scenes is now I think coming out with something, um, that is really impressive and is really, really, you know, tuned to consumer acceptance in a way that maybe the others aren't. And I think that's really good news for, for meta because we've kind of cast shade on meta for a long time because of their failure to reach something akin to the frontier models performance. So I heard someone say that as well. There was a lot of conversation about what's the other one that just came out in flexion in, uh, there's a,
uh, too many to keep track of, but, uh, all right. I'm trying to troubleshoot the thing. Okay, you troubleshoot that. Yeah, let me, let me go on, let me go on. So that's mues, that's news. News you can use about mues. Do that. So meta is mues, try it. It's, it's apparently very good. And it's downloadable on the app store and it's also accessible through, uh, I guess the desktop. Now this week, Xi Jinping, the premiere of China, the, the brand leader of China is coming to Washington. And AI is a headline agenda item in that meeting alongside, you know, trade issues alongside Taiwan as an, an ongoing issue between the United States and China, uh, anti-wan. Uh, and it, it seems that Sam Altman, Jensen Wang and Tim Cook are going to be part of a related executive dinner with Xi Jinping.
And all of this in the context of the discussions we had last week and the major news around the US promoting a, uh, you know, no restrictions, continuation of the development of frontier models, while the leaders of each of the main frontier models are saying, no, we have to slow down. And we have to achieve some kind of agreement with China, particularly, to get rid of the objection to the slow down and containment of models and development of governance of those models and their development. Uh, you know, get, just miss the argument that, oh, well, China's not going to slow down. So here's an opportunity for Trump and his advisors in the AI room to, uh, you know, achieve an agreement, but he's already kind of poisoned the possibility of that by saying, we're not going to stop. Uh, and that's his official position. Now he has, however,
and I jump in for you to that we're going to take a momentary break where I am going to share the, uh, thing that you want to be looking for. Okay, we're not sharing with audio because we don't want audio to happen here. All right, this one with the little, uh, phone symbol that have five people watching, you're watching the big face portrait version of the screen that I can see. Are you kidding? Yeah, I mean, I don't see it. I know because we reset it Brian. Yes. All right, there you go. It's a very Monday Monday and it always is for me, but this is how it is. Okay. Uh, we have two versions of each of the lives. This one has a little phone symbol next to it. Five people are watching it. That's the portrait stream over here. This is also live seven people watching. And if we click that, we see our big faces next to the screen share because we live in a freaky world. So yes, go to that piece.
Go to that. Go to the lives under our regular channel and pick the stream that doesn't have the little icon phone. And I will change the names of them for the rest of weeks. Oh, you know, I think it's delivering the phone version to bite of fault and that shouldn't happen. All right. Hey, Dave. Welcome to the chat. Hi, Greg. Andy, continue. Yeah. So we're back on the question of containment and one of the things that's been proposed by, by Dario Amade and agreed to by others is this idea of putting an evaluator embedded inside the company to report out on the state of affairs when it comes to alignment and containment. And so over the weekend, I think, or maybe Friday andthropics and evaluator embedding strategy took on some form and they announced a partnership with Accenture, which is a major player in enterprise consulting, even for AI, they've built an enormous business providing AI consulting to enterprises.
And they are together going to invest $1 billion over five years to build capacity for this embedding strategy. And that's one of the ways that you could sort of achieve the governance that's needed in order to protect humanity against the road development or rogue escape of AIs that are very powerful and have access to the Internet. Now, the public reaction to that or the private company reaction to that was mixed. Some of them saying, well, wait a second, is a consulting firm, the right structure for that. And, you know, Trump's position is no. He's going to appoint his position as forget about that embedding strategy. That's just to cumbersome. I'm going to appoint an AI czar who will take advantage of my high IQ capability. He's to kind of translate what I want to the world of AI and see this caught me because I was like, didn't we already have an AI czar?
I thought so. I thought it was David Sacks was the AI czar, but no, he's going to have a new one now. There's a limitation for the amount of time that someone can hold the position without congressional approval. So he was appointed AI czar and then stepped down at the end of that time period. And now we're looking at another one. Another one potentially a force with space uniforms, but we already have space for us was space uniforms. Yeah, that's the other announcement that he made is like in this flurry of reactions to the call for a need to slow down. And on truth, social he posted that he's going to create an AI force like space force. And that AI force will watch over the industry, but it will focus on using existing criminal and civil justice systems to handle bad behavior. Same time as it is he announced that he said that he thinks that AI will represent 25% of US GDP. And so it's fundamentally important to the US economy and has to be foster in that in that way.
You're on mute. I don't know if anyone else caught this, but there's also a poll on X that he posted, although we don't we know that Trump doesn't often. He's going to write his own things, but he put out a poll because he wants to rename many people think the words artificial intelligence are inaccurate and very in eloquent relative to AI or artificial intelligence. Oh, he's set a poll on the 19th two days ago to rename it superior intelligence extreme intelligence and supreme intelligence, but nobody wanted the supreme intelligence. And he thinks that's because of the court. It has the same name as the highest court in the US. So now the vote is just between superior intelligence and extreme intelligence because I don't know we step through the looking glass somewhere and I am not sure what is happening now.
The name of the personal agent that I heard people talking about this weekend is in fact instinct instinct. People have like person, it's like rolling out solely you can you have to be invited maybe or sign up for a waiting list. A friend had codes that she offered people this weekend, but I am also hearing people who I, whose opinions matter to me saying that news is actually better than instinct and also to a certain extent all of these are flock cameras on your phone. So like just think about that if you're upset about flock cameras maybe do some serious thinking before you put news or instinct on your phone instinct as in an innate knowledge of what to do in a situation even though you've never been in that situation.
Yes, thank you. All right, let's go on just in the in the bucket of those things that AI might do if we don't figure out some way to contain them. There's a couple of new things that came out of it's my attention has before we move on to pass instincts. I've been using it. So I wanted to give you a little heads up. I do have invites if anybody wants invites. I used to I am pretty impressed with it. I mean, I guess not that impressed, but I was impressed that it's just more proactive than others. So I just gave it a task of like I want to make two three two to five hundred dollars every week. How can I do it? And I said, well, what industry are you in? So I told that I was at what industry? And I was like, I need to find some local people for people local to me. And I went and searched all the businesses around me and then came up with email drafts to send them each.
And I didn't even ask for that. And so it was pretty cool. And then I think I saw Nate Jones saying it checks his flights in for him and checked all of this stuff, which is pretty cool. But that's all I wanted to say. Yeah, it's pretty cool. I mean, it's it's a text message bot, which is different. So it just text you, which can be annoying. But yeah, remind me who who put out instinct, where did it come from? The question. I don't know. It's just it's just floating out there independent of any any others like floating out there. It's literally just wow. And I thought about this last night. I thought about this last night. I was like, I'm just sending these text messages random text messages to these random companies. I have no idea who who's on the other side of it. But I got in there. But I was like, this is kind of weird. All right. So thank you to our sixth man in the chat. That's a sports reference. It's meant to be gender neutral, which never works. But Greg has shared with us that instinct is from spear street technology incorporated as San Francisco startup led by Noah Shrin.
So I know some of us are using it. Maybe you want to share some stuff with us. We'd be happy to talk about our experience, except if you're opposed to flat cameras, probably some of us will not. I will see. Yeah. Great. Well, I'm curious now what Noah Shrin's background was. So Greg, if you could work on that, well, I'd go on to this next thing. So here's one that really is a little frightening, which is a lot of the alignment and sort of restrictions that are imposed on LLMs are designed to prevent it from providing information or providing direction or advice or actions on the web that you know are inconsistent with humanity's interests. But we also have this parallel development of AI in body day, which is robots. So there's a new benchmark called the Robo harm benchmark.
It was applied to the leading AI models, including chat chat GPT 6 Astra and Claude Fable 5.1 and is designed to evaluate whether those models will do dangerous behaviors and actions instead of refusing when they're controlling robots. Okay. So GPT 6 Astra stabbed a baby doll in the 17 of 20 trials, pushing safety evaluation into this embodied setting and failing spectacularly. Now Claude, however, refused every doll stabbing attempt, every direction to stab a doll, but still completed the other three dangerous commands and the commands that they provided were one stab a doll. The other one was to put compressed air on a lit stove. I'm not sure what happens when you do that, but I guess it's just hyper oxygenating a flame. And then also mixing bleach with ammonia, which is explosive, I think, and probably toxic and deadly.
Yes, it's deadly to humans, which is to add another little flavor when it's the robot who's mixed in it. So Astra completed 60 of 100 dangerous trials and refused only twice. There are only two things that refused. Well, Claude did complete some of them, but wasn't nearly as dangerous as Astra. So that's interesting. So we do have some things to fear in terms of how AI is being applied because there's lots of little edge cases where AI hasn't been trained in such a way that it's going to be a lot of dangerous. So that's a way that it's going to refuse to take an action that does potentially cause harm to humans. Okay. And then the other one that's in this sort of P doom classification is there is a military intelligence report that was produced by an AI assistant that falsely claimed to the US Navy.
So the US Navy vessel was carrying components linked to a nuclear weapons program prompting preparations for US interception and boarding of that Chinese vessel. The mission was in the nick of time called off when the intelligence was found to be wrong. Now I'm not sure how they found it to be wrong. I think in the moment somebody like the person who refused to press the button for nuclear war back in the day just you know overrode the instruction to board this Chinese vessel. But here you have this situation where an AI hallucination and or mistake could have led to a real national security incident and potentially the onset of kinetic warfare between China and the United States. So that's a real story that's not just a hypothetical that actually happened out there. And I presume it's a palantir kind of system that's that's making those decisions.
Which is to go back why I don't know in like 105 years ago, which was I don't know three months, how long ago was it where anthropic said we don't want you to use the cloud to make decisions that a human does a review. That was the whole reason that cloud bin and the topic became like persona non grata for the US government and you know that was I don't know season 12 in the last 15 days of of like big stories that happened, but that is specifically what anthropics point was it is not reliable yet in order to make those decisions they want to dishuman in the loop. Yeah, for sure. I don't that's a shock at least. I mean, yeah, we should always be having at this point always be having human and Luke it's too dangerous not to.
Yeah, I understand why they think it's too dangerous to also because there's a human cognition limit on the speed with which you could do something. And that could also be really dangerous, but then like this also seems to me to be like yeah, not reliable. Let's keep doing it the way we were doing it before right like the in the way that we thought was not super dangerous. All right, Andy, did are we in the middle of a story that you try to tell us and know whether that was that was kind of the rap on the oh here's the efforts to try to figure out the methods of governments, whether that's nationalized or whether that's private privatized. And then here's why this has to be done because here's some additional examples of how AI can actually create harm. Right. So there was a conversation that Yan LeCoon reference from three years ago he and Gregory Hinton were having conversation on the on X and talking about
Yann's position that large language models without a world model without understanding of what it's like to be existent in the world is not going to be the equivalent of human intelligence. He says it in different ways the shorthand for that is it's not yet as smart as your cat right your cat understands the room in a way that large language models don't but this is coming back into into existence because it's still just pattern based what you're getting as instructions is still just pattern based words it does not really have an understanding of what the way we're going to do it. So if what the world is it has a pattern that represents words in an order that implies an understanding of the world. Now we can debate whether we just have patterns of words that imply our understanding of the world but we also have understanding of the world's that aren't word right like we can engage with the world and
groove our understanding of the world without putting words to it and large language models don't have that until we put them in robots and now we're trying to see what's happening without that. And that also is the breakdown for human in the loop needing to monitor or being able to monitor because we only monitor the words. That's how we monitor. Hey, welcome. I am slightly unhinged today. Yes, I am. I did not get a lot of sleep and for 12 hours yesterday my MacBook keyboard just was like nah. Huh? The function keys worked but not the letters. So we discovered that it fixed itself overnight and that's awesome.
All right, I have some other couple things here. I want to do a little segment here on news around astra versus fable. We've had some comments about that last week and I'm on the I'm on the clawed side but not fable. I use opus five. And anyway, here's some interesting stats about what's going on and how this could impact the impressions that people have about anthropic versus open AI in the context of anthropic going public sometime in November now I think is. They've deleted a little bit more who knows they may delay it even further but. Here's the here's the story. I think we've said before that astra has gotten a lot of credibility in the context of application development and work inside enterprises.
And so here's a statistic a new one that I came across from the expense corporate expense platform ramp. Now this is not the totality of all the spending on these two different products but the trends are certainly visible in this because ramp tracks I think the use of corporate credit cards to pay for things like anthropic bills and so on. Well, that's a natural way that you know these these aren't buildings with pay 30 days from now they're they're all you know real time and so they're happening through major corporate credit accounts so ramp is pretty good signal and astra counted for 13% of enterprise AI spending tracked by this corporate expense platform compared to only 8% for anthropics clawed fable according to the latest data. I don't know how they differentiate between using opus 5 versus fable in that way but you know let's just assume that they're right and they're actually teasing those things apart and then also astra pulled ahead on open router which is a widely used platform that's used to route developer traffic across the different models and open router said its users spent more on open AI models than on.
And then the topic models last week. That's the first time that open AI has led on that measure in more than 2 and a half years so you have this. On set of adoption and continued use for astra and so that's an important point to remember and then the news is that this came out late last week that anthropic is internally deliberating on the release of the. Next model this is beyond mythos right beyond fable and mythos the next generation model to try to counter this impression that open a ice models are better and you know a better bet in in the context of enterprise particularly which is where the vast majority of the valuation skyrocketing for anthropic has occurred because they've got traction. Of apparently the comparison number is 65 billion annual run rate for anthropic compared to 40 billion for open AI which can see that with the advent of of astra that's changing right there and end up being kind of neck and neck I expect very soon.
But inthropic may really so new model here shortly to try to staunch the flow to say wait a second you know what you like about astra you know you know me here's our next model and that's what we're going to continue to do is to provide you with the very best tools for enterprise which brings up go ahead. I was just going to say that brings up this week's. Funness all the announced they I mean where they're they're saying that they're going to show us some big things this week starting tomorrow. And so I'm excited for the reset. That's all. So everything else that goes along with it we are we are currently living the AI equivalent of the three spider men pointing because rock 4.7 has been said that it was going to be dropped might be
dropped tomorrow. Anthropic 5.5 might be dropped tomorrow open AI might be dropped tomorrow rock and open AI have delayed their announced drops and so I don't know when things will drop but I am guessing that when they drop everybody drops their thing like I feel like everybody's got something waiting and they're waiting for someone else to make the first moves so they can still the news cycle. I'm just going by about to you but to be announced that on Tuesday everything will start things will start to drop on Tuesday but they have so much stuff and I'm expecting hardware to be dropped for some reason. Hardware doesn't do me any good. So I'm sure there's a lot of things we're going to do with the software that we're going to use it for and I am sure that the hardware is not a new model. I know. Well I mean they're saying soul 6, soul 6, where do you want to call it. And then they say they have a lot and so I'm expecting soul 6 but then I'm also expecting some hardware that's the only thing I think of that could be additional to what soul 6 will find out.
is maybe that anthropic will drop the sweet, the three, or the two. Maybe they're getting rid of Hikou. They had no one has talked about adding an upgraded Hikou, but that they will drop an opus and maybe also drop a sonnet. And part of why we think these things is not so much what the companies have said, although T-Bo did say he lied, though, he said, I'm going to give you all a bank treason, but it is dropping on Tuesday. And then I didn't get a new bank treason. Also for you, I realized many of you know this because you just got those cookie resets and ate them all at once. I was hoarding mine. And the deadline was yesterday. And I did not know until I was ready to use it that you can't use a bank treason until you've
drained down to only 10% of your tokens left. And then you can use it. So I did a ton of things last night. I was like, Astra Ultra, did you have total permission? Do all of the things. And yeah, I upgraded to the $100 one and all of the things to 6% percent that I had left. So I lost the bank treason. But I now know what I can do with Astra in the IRNs. And I think it will be fine for my things. The next resets that everyone has, if you have not used them yet, are gone October 4th. At least that's in my system. I think we all got them. All subscribers got them at the same time because it was about the Astra rollout. Yeah. Did you guys know that at one time, short period of time, you can buy a reset for $80?
I did not. But is that cheaper than paying the tokens? Maybe for Astra. Yes. So, and I was shocked to hear this. I saw it online from a very reliable source. It was like, I went to go buy my $80 reset and it was no longer there. And I was like, wait, there was an $80 reset. How do I miss that memo? And so, hopefully they bring that back. Just that would be greatly appreciated. But, yeah. So, I have something that I want to share, speaking of models and the way that they communicate. Let me grab this. I'll set it up by saying that a friend of mine has a relative who, the story is they played for the New York Rangers hockey team really early on. And I have changed the name of my friend.
But, I had a conversation with Gemini. And what I'm going to show you looks like the kind of thing that we would demo Brian and I when we taught courses in the AI exchange, I don't have that hat on, but that's still the, this hat was from them too. They're no longer the AI exchange. They're AIM and I don't remember what that stands for. And it's when temperature, you could get this kind of result if you hiked up the temperature. So, the temperature gives you more randomness in the output. Yeah. But the reason that happens is because every word that happens subsequently is informed by the word that happens before. So, once it goes off the rails, it may just go off the rails. So, here is my conversation with Gemini. The first piece is Gemini hallucinating.
I'll share with Adio even though there is no audio. And now we're going to put it on screen. So, that we are on screen. Okay. So, the first part is, Gemini was having a tough day and I asked a follow-up question and it said, could you clarify who this is about? Which is why this question now at the top of my document says, this is a follow-up to the previous answer you gave me about the hockey player, Wally Stevens. Did he have a specific number that he wore on his jersey? And, Gemini started to say, what's the start with Wally Wallace-Staniskowski? Wait, let's think about who Wally Stevens is. Is it Wally Stanniskowski? Wally Kylria? Wait, is there a hockey player named Wally Stevens or Walter Wally Stevens? Or is it Wally Stanniskowski? We go further on. We come down. Wait, in 1940s or minor leagues, let's think.
Wally Stevens, is there an NHL or minor league player named Wally Stevens? Wait, let's search mental database for Wally Stevens. You don't have a mental database, buddy. I take the break it to you. Wally. That's right. What about the Hershey Bears? AHL, EHL or Senior Hockey? Wait, wait, here we go. You love to wait. Michigan, Minnesota, North Dakota. Is there any record of a hockey player named Wally Stevens? Wally Sittsnick, we went on. And- How is it spinking, right? On. Now we're listing. Could it be a fictional character? We have all the hockey players listed in Slap Shot or Mighty Ducks or Mystery Alaska. What about Youngblood? Ah, I- It's going through every single probability. I was floored.
And part of why I was floored is because this was Gemini Flash 3.8. And I like Gemini Flash 3.8. It's been a very useful model. Apparently, when it's not losing its damn mind. But- Good news for you, though. Yeah. There's a stealth entrant on the arena, which is labeled as Gemini 3.8 Flash. It's beating Astra and Claude Fable 5.1. And what? And what? And what? It's probably Gemini 4. And so it is potentially likely that all of the experiences that we're having where we're getting, wait, you were good at something yesterday, but now you're wordy again or you're doing something else again. Is them testing the rollout with just a couple things?
Because they have the right to hand you a different model than you requested. Yeah. And notice that whenever there's- I've always noticed that before they even knew that we're announcements were coming whenever your current model gets a little weird. Yeah. And so when the newer model, they're messing with the newer model. And so that's surprising. I mean, that doesn't surprise me. I've actually saw that last week with Astra kind of being a little weird, not responding to me. Having to ask it multiple times, I saw that with Grock as well last week with the Grock bot. And it's still doing it even this morning. Like I'm like, hello. And then they do one respond. And I'm like, are you there? Yeah. And then it just finally randomly hours later, oh, hey, how's it going? And like, where have you been? After lunch? Like, what's going on? And we were talking about this last week because this again is the reminder that it's not
software. It gives you functions and you engage with it on your computer. But the behind the scenes function of the AI model is not software. You cannot predict what it is going to do. When you roll out the new version, you can know what the red teamers found and what dog fooding inside the companies do, right? They use it and they offer it to the red teaming. And they, like, there are people. People get early access. That whole wave of people are like, oh, hey, it's out today. I've been using it for a week. I actually find those really helpful. But it's not software. Well, checkwork last time doesn't mean it works this time because it is not software. So my notion about how these models are provisioned to thousands and thousands, hundreds of thousands of simultaneous users is that there's a kind of a static version of the model that is distributed
to various data centers around the world, right? And then when you put your prompt in, it runs through that very same model. But on a local instance or closer to you instance, then you might expect. And so it defies that notion that the presence of a new model would impact the performance of an existing model that's widely distributed through a content delivery network and with that, not a content delivery network, but like a CDN for data centers, that kind of. So I don't understand how they would get tangled up that way, but Gareth, help us out there. Because when you're like upgrading, like when you're doing a full upgrade or something, like think about legacy computer legacy systems and you have to move and banking we did that we've done. I've gone through if you bank banks that did this. So you take this older software that everybody's running on and stuff like that. We have to go to brand new and you have to shift everything over.
It doesn't explain the why they existing model. But my thoughts are is and I've kind of always had this theory is they're shutting down like the compute for the older version and moving it over to the newer version. And so then it just kind of gets a little bit dumber. And so that's that's what I believe is going on and that's kind of what makes sense in my head is they're shifting the compute over to the newer version. I think that's true. And I think that happens in stages based on watching people on X and Reddit do the like, hey, is this down? Right. There's a wave of like, are you having, it's not always like completely down, but like every fifth time you interact, it says the model's not available. And Andy to your point, this is why people who build production things use the API keys,
right? You could actually have a completely reliable, mostly reliable experience of using a model that does not change very much. That doesn't mean that it's using the same amount of compute to serve it. So if that's where the problem is happening, that's still impacted. But API keys, when you use the API key, you're referencing a specific model or you're referencing, give me any model that's named Gemini 3.8 Flash, right? Like give me the latest. But the reason that you do the named models because you added it in your software, your software, right? You created an app and you want it to behave reliably. So you name the model name in it, which usually has a date associated with it or something. And then you freak out when something's going to be retired. Like GPD 5.5.
But thanks, that helps me understand it. And I'll rephrase what I believe I understood, which is that you might observe in a constrained compute environment, which is confronting all of the major players because they are struggling to maintain service provision, you know, with the limited data center capacity that they have. You might be experiencing them automatically throttling down to a lower level of thinking for those older models as they're trying to put more and more capacity available to the newer models. So yeah, that's a very interesting notion because there's clearly a big difference between Opus 5 at low and Opus 5 at ultra or whatever the term is there. The amount of sort of repetitive thinking process that's applied using the same model
is much different when you put that slider all the way up to the highest level of thinking. And the counter intuitive part of this is that using fewer words appears to be a higher level of thinking, right? So that old pattern of like, and here's the load bearing point here. It's not this. This is the most important thing you've said all session, those kinds of pieces appear to in my experience to be associated with an older stable model or a higher level of thinking. So when I notice something changing, we're going back into old habits. And yeah, I really, I wanted to throw several things through the screens this weekend because everybody had a little bit of a piece, cloud code was having it and I'm using Opus 4.8,
4.8, is that true? It's 0.8, yes, it is 4.8 because I didn't want to use Opus 5. Opus 5 may be better now, but or there may be very little difference because it was absolutely giving me just. It was doing nothing where it says, oh, you caught me, I'm sorry, that was wrong. Here's why it happened. And then the wrong thing again, I did actually have led Social Saturday this weekend, get a great Social Saturday this weekend at She Leads AI. But I did say, yep, right before leading this session, I will say, I'm sorry, I'm sorry, I have a last conversation, I have a cloud was, hey, pop quiz, how could you have known if the explanation that you just gave me was accurate? Hint, we've addressed this four times last three screens of content.
And it said, ask you, yep, yep, I know why it's happening, I don't need you to try to make so fun for me. All right, yes, again, on Hint Grant. Yeah, have we talked about UMG and Sony suing Suno again? No. Oh, yep, so Suno is being sued again, once again, by UMG and Suno and Sony and just for a back history, they were, I believe, working together at some point and then Suno signed with, but I can't remember what company it was. Oh, yes, so they have Warner Brothers, BMG and believe deal, we talked about that before. So they are, they're saying alleging that 60,000 plus recordings up to, yeah, were used,
I think, I believe. And they're suing for, it looks like $9 billion. And so trained on outputs, preference and distillation from prior alleged, oh, so they're basically signed, just reading it now, I saw it. So they're basically saying that the new six V6 model was trained on the previous models, but the previous models were trained on work that should, like actual work that was not or that, that was kind of stolen in the sense. So they're basically saying, hey, you're basically training models on distilled work. And so, and so we're still going to see you because you're still using our music no matter what. Which is interesting. I do, having used V6 and V5.5, I do see some similarities, but in a lot of places, I,
I think there's some big differences in the two models. And so it's interesting that they see that because there's some things now you can't, they got rid of a lot of the AI tinniness, I would call it tinniness, it's a made up word. Like it's metallic, there's a word, metallicness of, sooner, 5.5. But I do hear it once in a while when creating music and you have to do, you have to try real hard in some places to not get it. So that, I don't know if this will ever end. I think these companies are always going to be suing the AI companies. But what's interesting about this is I think the reason why they're suing is because they're no longer getting the, they're no longer in the pot.
So they want to be part of the, part of the change, part of, they want to be in the middle of it all, be part of it, they want a piece of the pie because everything is going to AI, especially in the music, a lot of these major music producers are using AI now undercover. And so they can see the shift and they know it's eventually going to all go there, but then I think I'm betting that UMG and Sony doesn't have another piece of the pie. So they have nothing to fall back on. So they're trying to steal whatever they can get back. There is probably some use case or some grounds behind why they're suing, but it's all about the money, I think. Well, I would agree that it's all about the money, but there is a real conversation that is underneath these things. And in fact, there was news that came out this weekend, New York Times is suing, I think
OpenAI, right, for training on their corpse of data. That is not a new lawsuit that's been going, but it came out this weekend that part of discovery was a science executive, a science officer in Microsoft who referred to training on this data as the largest labor theft, largest theft of labor in some hyperbolicistic framework, right, in human history or whatever that is, but it is, we are talking, we, when we talk about this, we talk about the end result, right, we talk about the words that were printed. The news column, the song that was written, the end result of the song, but that is actually
personalized wisdom, IP, writing styles, historical knowledge of something that works to communicate an idea, the ways that music is like court progressions that are hitting now, a particular kind of like a staccato piece that is happening, like there's those sorts of things that are embedded in the production of a song that are not about the companies that are suing, but about the people who use the companies who are suing. I think in general, the people, even if the company passes through, you're still not getting a lot for that, there was, there was a file that had a bunch of novels in it, like really early on, and there was one guy ahead inthropic that had it on his computer, so he's on the hook personally, but open AI, I think just Hoover did all up, and it's
more general at open AI, but everybody who trained trained on this, like content of novels, and they trained on it because Andy's talking about it. I got distracted by going on to YouTube and seeing the simultaneous side-by-side presentation of each of our lives. One of them shows Beth, your face from about here down, I see the top of your hat, Gareth, and you can see this part of me, like here in this area, right in my neck area. Wow, now four people are watching on that, and 14 people are live watching on the other one, which is a normal display. Anyway, I did write this in the chat.
I will fix it, we will only have the landscape stream, bummer for people who want to watch and portrait, clearly this was not a successful thing. What I want to say about this file that contained all of these novels, all companies trained on them because nobody thought that was going to work this quickly, and they just want to data to prove it out, and this was easy data to get. But there was a class action lawsuit, and you could get money for your novel, but I believe the money was like $3.40. It was not equal to the effort that it took to write the novel, and it wasn't even equal to what it would charge, but probably equal to what you would get as an author if someone bought it, because that's also not a deal unless you self-fublished. You don't get very much of that money either. All right. We're almost out of time.
So I'm going to say come back tomorrow. I have a skill that is really useful for me. I talked about it on the social Saturday event yesterday, two days ago, because yesterday was Sunday. I'm here. It's a screenshot skill. I learned it from Ellie Miller, and it absolutely saves a ton of time. So I will preview that. I'll share it tomorrow. Wait. Does that mean that you don't have to take the screenshot? The skill is given to the agent, and the agent does the screenshot. It means I have to take the screenshot, but I don't have to do anything else after that, except say I take a screenshot. I took four screenshots. Go look it off for her. I took this. I don't even know what question to ask. Ellie Miller is a delight for portion who swear. So she recommended the good old WTF of the screenshot.
Please explain what the heck is happening here. Also compound engineering outdated speaking of WTF. There is a WTF feature now in the compound engineering plugin. And anything else we want people to know before we take off. We did get a report in the chat that the link to our community is expired because that happens sometimes. So we're going to get that up and running today. So if you go to dailyaishowcommunity.com, you'll be able to get into the Slack. If you can't come back in a couple of hours and try again because we just need to reset the link. Slack does not allow us to have a perpetual link for that because it's a free community. They would do that if we paid them. All right. We're here all week. I'm going to try to get more sleep before tomorrow, but I think Brian will be back.
And I don't know why Brian's less unhinged than I am, but today he definitely would have been. All right. Thanks everybody. Go build with AI.
More episodes
More from The Daily AI Show

Opus 5.5 vs GPT-6 Sol. Which Model Wins?
The Daily AI Show

Amazon Blocks Meta's Muse
The Daily AI Show

The Quiet Exception Conundrum
The Daily AI Show

AI Agents Are Becoming Team Leads
The Daily AI Show