
Get every episode summarized
Each time The Daily AI Show publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
“And I can already tell long enough issues. I should not have cloud running in the background, but it is. So I will get that situated here in a second. But yeah, what a weird 24 hours as far as model releases.”From the transcript
OpenAI and Anthropic released new models within 90 minutes of each other, shifting the conversation toward an AI price war. GPT-6 Sol and Luna arrived with lower prices, while Claude Opus 5.5 showed a substantial improvement on the Artificial Analysis Intelligence Index. But cheaper tokens do not necessarily mean cheaper work. Brian shared a direct comparison from AJOVA Journeys: Opus 5.5 cost $2.99 to complete four research and planning steps, versus $1.05 for GPT-6 Sol. The initial evaluation found that Opus included more passenger quotes and better captured Amanda’s voice. The hosts discussed whether higher-quality output justifies the extra cost, why medium reasoning effort sometimes performs better than higher settings, and how businesses should evaluate individual steps rather than commit to one model.
The discussion expanded into AI agents and commerce. Meta’s Muse reportedly reached 500,000 users in its first week, Stripe introduced MCP-based checkout tools for AI shopping agents, and Amazon’s restrictions on outside agents raised questions about who controls the future of online shopping. Other topics included DeepSeek’s rising usage, a simulated economy where AI agents struggled to adjust prices, social media content farms, and Runway’s experimental interfaces that generate and adapt interactive scenes to different screen sizes.
The Daily AI Show Live_ September 23_ 2026.txt
Key Points Discussed
00:00:20 Episode Intro And Three Major Model Releases 00:03:28 The AI Model Price War Begins 00:05:43 Opus 5.5 Versus GPT-6 On Artificial Analysis 00:09:18 Why Medium Reasoning Might Beat Higher Effort 00:11:46 Changing Reasoning Effort Without Losing Cache 00:14:03 Brian Compares Opus 5.5 And GPT-6 Sol 00:15:39 A $2.99 Versus $1.05 Production Test 00:17:11 Which Model Better Captures Amanda’s Voice? 00:20:19 OpenAI’s Model Roadmap And A Deleted Post 00:22:04 Why Gemini Still Matters For Video Analysis 00:26:14 Early Reactions To Opus 5.5 00:30:55 Does Model Quality Outweigh Token Savings? 00:33:44 DeepSeek’s Growth And Specialized AI Workflows 00:37:06 Testing Opus 5.5 On Automated Thumbnails 00:43:12 Meta Muse Reaches 500,000 Users 00:44:15 Stripe Introduces Checkout Tools For AI Agents 00:46:10 Amazon’s Restrictions On Outside Shopping Agents 00:47:50 What Happens When AI Agents Run An Economy? 00:49:58 Why Faster AI Work Doesn’t Always Increase Productivity 00:51:45 Inside A Social Media Content Farm 00:55:55 Runway Demonstrates Interactive Generative Interfaces 01:01:10 Using AI To Operate Unfamiliar And Legacy Software 01:03:16 The Debate Over Renaming Artificial Intelligence 01:05:27 Episode Wrap-Up
The Daily AI Show Co Hosts: Brian Maucere, Andy Halliday, Beth Lyons, Gareth Hood, Karl Yeh.
Get every episode summarized
Each time The Daily AI Show publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
891 searchable segments. Every word is indexed and playable.
Full transcript
The Daily AI Show — Opus 5.5 vs GPT-6 Sol. Which Model Wins?. Machine-transcribed; use the interactive transcript above to jump the player to any line.
Hey, what's going on everybody? Welcome to the Day of the Eye show. We're live. It's September 23rd, 2026. And I can already tell long enough issues. I got a stop-clone in the background. But with me today are Andy and Beth. I know better. I should not have cloud running in the background, but it is. So I will get that situated here in a second. So I stopped freezing. But yeah, what a weird 24 hours as far as model releases. Not to say anything from GROC 4.7 was released on Monday. But then yesterday within 90 minutes of each other, Anthropic and OpenAI released new models, which there were some resets. And I know Beth you talked about that on the OpenAI side, as well as some things that Anthropic did. There was just a lot that happened between when we were alive yesterday and today. And it seems to be all everybody wants to talk about. There's definitely like price wars. No, it's not a coincidence that they were announced 90 bit, 90 minutes apart from each other.
We know that. So lots to get into here. And Andy, we were just talking right literally before we went live. And we were saying, you know, usually there's some good graphics to put up and usually you're on top of that. I know you said you're also opening up. What's the website that you just said you would know about? It's artificial analysis. And so it's worth sharing. I'll do that in a moment here. But just let me talk about the drop of the models yesterday. And I'm putting it in the context of a term that Brian used, which is price wars. And that's really the dimension that's being played now in the competition for supremacy in the world of AI. So the truth is that as of yesterday, now, shocking levels of intelligence are available at prices that would have been unthinkable just a few months ago. And it's just a few months ago that this story was about token maxing. And the accidental $500,000 or a million dollar bills
that enterprises were, you know, hit with, as a result of the high costs of API tokens. And so yesterday, OpenAI released GPT-6 sole. So this is GPT-6 Astra was available earlier. But now the sole and Luna sort of smaller scale models are now available in the GPT-6 lineage. And they bring down the cost to top customer business by 50% from the 5.6 versions of that same model line. So like if you were using 5.6 sole, now if you want to use six sole, it's half the cost. But it's better, OK? So that's what's happening. We're really getting to a price, you know, to performance ratio that's really much better than it was before. And then Anthropic yesterday released Opus 5.5, which the company says costs around 40% less than Opus 5.
While maintaining top intelligence, speaking of top intelligence, let me share my screen and show you where Opus 5.5 comes out. And, you know, it's interesting that sole and Luna aren't even shown on this display. I don't know why. They still show GPT-6. But let me share my screen. So it's an interruption to the process because it takes about five clicks and various, OK, here we go. So this is artificial analysis. And artificial analysis is doing an independent analysis of AI, as you see here. And they create this artificial analysis intelligence index. What's stark about this is you see GPT-6 Astra on the max setting was 53 on their scale alongside Fable 5.1. And remember how recently these models seemed
like they were killer. Well, Opus 5.5 jumps from 53 to 58 on their scale. And there's on artificial analysis, if you go down here to intelligence, that's the index here. Let's see where sole max comes out. Only at 48, at the same level, this is GPT-6 sole, which was released yesterday. It comes out the same level as mu spark 1.3 on that artificial analysis intelligence index. So on this expanded thing, you can see that here's where six max on sole was at 48, it just came out at 48. 5.6 was already at 47. But this is less expensive and at the same level of intelligence. But this is really pretty impressive on this particular index, which covers all of these various benchmarks that you see here.
Artificial incorporates a briefcase, GDP-VAL, automatic bench, A, automation bench, terminal bench, side code, humanities, last exam, and a bunch of others, omniscience, et cetera, to build this index. Now, if we scroll down here a little bit, we can see that Opus 5.5 is at a pretty nice position on the scale of intelligence index versus cost for task, as well as showing well below, Fable 5.1. And Opus 5 is about the same level. This shows Claud Opus 5.5 being cost per task a little bit higher down here on this X scale. And so that's in opposition to the assertion that Opus 5.5 is actually cheaper. Because if you look at the cost per task,
that's the combination of the token cost and the number of tokens that have to be used in order to deliver a response. Anyway, any questions about that performance and where Opus 5.5 came out? Here on this particular display, you see that the... This is again, just scaling their own artificial analysis index on the y-axis and over here, time. You see that there is a big jump by anthropic that's ahead of the other players when it comes to this intelligence scale. When coding, Fable 5.1 seems to still be there. Here's Opus 5. I don't see 5.5 in here. So they haven't actually done that since it just came out yesterday. They haven't completed their analyses.
So I saw something and I'm sorry, I thought it was gonna be that one where 5.5 medium scored higher on coding than the higher extra high. And that I think we're seeing more and more. You have a deterministic thing that you want done. Don't give it go, think lots of thoughts very hard. Like no, just give me what I asked for. And medium is also an easier or a lower cost for it uses, it's not a lower cost. It uses fewer tokens, I guess, which is why it's a lower cost. Yeah, that was a big call out on in the original X post I saw from anthropic, I guess that would be right from 5.5. So and Gareth is called that out too, saying 5.5 medium is the way because that's what everybody seemed to be saying. So 5.5 medium is just use medium.
Why would you use higher extra high? It has the graph superior to say it degrades. I was trying to figure out if there was something more to that because that seems odd and that's what a lot of people said to and there wasn't any definitive answer. But I was wondering if that also had to do with how many concurrent, how many parallel tasks were happening at once because it seemed like the graphs are also pointing out to that. But to be honest with you, I was with everybody else. I mean, everybody seemed to be saying the same thing, which is like, I mean, I'm looking at the graphs and it looks like 5.5 medium would be your best option. You know, if you're going to be running this through. So I guess maybe more to come on that is people test it. But that's, I heard, I saw the same thing Gareth and you were talking about. I was hearing talk like that about strategies for using codecs, same kind of idea. Like if you have a deterministic thing that you want codecs to do, don't give it a ton of room on the top.
You don't need ultra max, you don't need because codecs is max, right? I don't know. I think anthropic is extra high. Maybe I'm there. Of course you don't use the same terms. Yeah. But what I do think we're moving into that place. The other thing about what anthropic shifted a couple of weeks ago what is that you can change the effort midstream and still use your cache. But it used to be that if you switched, like you're going to go from high to medium, it's a whole new thing and you're going to get the full cache hit. Now you can switch within the effort. So if you've had a conversation where you wanted it to, like you wanted a little more in an interview kind of thing, ask me a couple of questions. Let's imagine some things. Great. Now we've got a plan. Go down to medium, execute it. Don't overthink yourself. Yeah.
And so, OK, so I'm just going to re-re-say what I think I you just said because if I'm in either of the tools and I move between high medium low or whatever the names are, the comparable names are, there is no re-cache hit. I know that that's not obviously the case when I was moving from table 5.1 down to 5 previously or up. Anthropic would actually throw up in the cloud, code desktop would throw up a thing and say, are you sure you want to do that? Because it's going to have to re-read the entire conversation. And that's definitely going to hit you. And it just gives you a quick warning. And you could say, yeah, it's either worth it to me or if I really do want to use table 5.1 in last week's examples or even earlier this week examples, I might choose to do a compact. I might choose to do a handoff. I might choose to do a brand new session. Then move to that other model. But just to be clear, Beth, you're saying, no, if you're moving between high medium and low or whatever the naming conventions are,
that does not impact the model having to re-read or cache more information or anything like that. You're just literally using variations of the same model. But the model has obviously not changed. Right. You can't go to a fable in the middle of something. Correct. And that is a recent change. That's something that wasn't announced, but it wasn't like everybody went, woohoo, this is amazing. But now that there is a significant difference and you might use a higher level to do one part of a conversation and then move to medium to execute it, now this seems way more relevant. So I'll just put this up on the screen, but it's really text heavy. So I'm just going to call it a quick couple of things. I literally, I mean, just quite literally this morning, I had it doing some club code doing some background analysis because I got, I was using 5.5 and I got a email that said, you're being charged $10 from Anthropic
and then another one that said $11 from Anthropic. Now this is for what I'm doing for Amanda for Joe Vigerni. So it was for that project. I was like dang, $10 immediately by 11, what the hell's going on? Well, it was rerunning some things. I was, I was headed in the background this morning before I started scaled work. And so I investigated it and I said, well, okay, we were using 5.5 grade through the API since yesterday, but I had to do this side-by-side analysis of using Soul 6, GPT 6 Soul versus 5.5 for cost. But there is, there is something I definitely want to call out that Andy and I were talking about ahead of time. Now if we just look at these first four steps that you're seeing on screen, if you care to even look at this screen, it's fine if you don't, are the first four steps that happen after an initial idea is presented to Amanda for a new video about curses, right? For our Joe Vigerni's, if you've been following along, if not, it's okay. That's what it is.
That idea now has to go through what is now a 16 step process. That's 16 step is like literally push it to YouTube. So that's the end of the road. But the beginning is what's the idea, what's the outline? Where does Amanda be the human in the loop, all the things? So these first four steps are side-by-side, comparison and just looking at this part alone, Opus 5.5 costs $2.99 to GPT sold $1.05. So one third the cost in these particular four steps. Something to keep in mind though, is web searches are included at $10 per 1,000 for both. So what I was saying to Andy, and I was saying before the show is, if you are using your APIs, your whatever models, and part of that as it is in my case, is heavy research. Because what happens is like in this case, she wants to do a video about the Marguerite Abilities Islander Cruz, she was just doing a ship tour of that Cruz. But it needs to go out and it's validating a bunch of information
on Marguerite Abilities website. So there's a lot of web search in there really. Now, I was able to automatically knock these prices down really quick by basically saying search the website, find the facts, hold it in the schema in the database, and that actually solved a lot of my problem right off the bat for this. So I have solved for the pricing problem of the web search by just doing it more efficiently. But this is just a side by side comparison. We don't need to go through it. But basically, it's just like you would, if you're doing an evaluation on the left, is Opus 5.5 on the right is the output of GPT 6 sold. And at the very top here, it gives us not only the cost analysis for where it actually thought one did a better job than the other. It says like one of the things is that Opus actually went through five fact checks. So when you're somewhere, it's talking about Opus is outlined and carried 12 passenger quotes to GPT 6 this sold five, which is actually for this particular use
case problematic to me. Those quotes are actually really important. So I'll be curious to see if that holds or not. And then it says here, Opus sounded more like Amanda. And it gave a quote as, I wouldn't thinking the shortest trip would be the easy first pick. GPT sold reads more like an analyst. Mary Scaries, that's a user name. Preference doesn't conflict with those averages. This is something that we'll be getting better with Amanda's voice as we go, right? So it's just give me a side by side. I have to actually read through all that. And what I would imagine, as we always said, is it's not all of one or the other. What I imagine is maybe step one will be Opus. Maybe step two will be sold. Maybe step three will go back to Opus or some version of that. Because truthfully, if one particular step is doing a better job in a different model, let's go ahead and use that model. But there's no doubt about it. To do this research is costing roughly a dollar versus $3. If I'm doing, if Amanda's doing three a week, I don't think that's going to break anybody's bank.
But hey, that's still $9 to $3 a week. And that adds up over a year for the same potential quality output. So anybody that's out there looking at these APIs and what you're doing. Be sure you're doing the Vals. And truthfully, after this, what I really should do is say, what's the other one that came out, Luna? Did they skip Terra? Is that what they did? Look, it's a G36 Luna. That's right. It's just sole and Luna. So the next thing I'd have to do, too, is actually look at it and say, could I potentially get away with Luna for some of this as well? Which would drop the cost even more? That takes time. You have to take a while to figure that out. It would take multiple runs. And it's probably not something I would really know the answer to as a definitive cost for maybe a month or two doing this multiple times per week. And I actually just said that to Amanda, which is why I built that cost page that I showed yesterday. It's because she was like, hey, I do want to know what this business is costing me, because it's not bringing in any money to the family. And we don't really know when her first client would be.
And it's probably not going to be enough to cover our cost to the business. So she's like, I just want to know what our loss lead is like, and for how long. And I said, yep, and then you're right to ask that. But also, I don't know the answer and won't know the answer definitively for quite some time because it just takes time to find the efficiencies in running tokens. And she's like, yeah, fair. But tell me what we're paying. Which is the right answer. So that is the codex equivalent or similar strategy that I was referencing because people would be like design it or like plan it in Astra and then use Luna Max. Yes. Right? And Luna Max, you lost a little, but it was perfectly capable of executing. And it was so much cheaper. Right? Yeah. And I do think that I don't know if they're going to bring out Tara.
Sam also had a weird post on X. And I don't know if anybody else saw it, but he made a mistake. And that's not Sam's X strategy. So his post said, we want the OpenAI API to feature the best model at every price point and to be the best at every modality, text, code, image, video, etc. And then we all want you to come up with great ideas and build them and get to be happy users and the best ideas come from y'all. And then he said, ooh, and it's gone this time. I saw it earlier. That's very interesting. He deleted a post. He deleted his response to that was, oops, not video, we meant voice. And now that's not the post that is the next from him.
Now it says, we heart developers see you next week. So some of Sam's happening. We don't know. I personally don't go to video again. Please, like, hmm. Well, I mean, look, I would love it if OpenAI had something as good. As what Gemini puts out, I mean, still thing Gemini has to standard in a lot of ways for that stuff. So yeah, I would, I'd be thrilled if they do better with video. Let's, all right. So I'm going to revise my request. Gettos, good at Gemini at watching video. Yes. Especially when it's here. And definitely don't talk about generating. Yeah. Like, and by the way, Gemini still does the like screenshot thing. I we're still not at a model to my knowledge. There might be an internal model somewhere where it's watching a video at anything close to 30 frames per second. Like, that's the, that's your standard for animation unless that's changed over the years. But or 24 might be the animation standard 24 frames per second when, when you used to have
animators hand draw. That's what it was. It was, it was 24 hand drawings per second of which is why they took forever, you know, the Simpson still takes forever to produce like six months is the lead time for a Simpson's episode or something like that. But, you know, obviously digital is a little bit different or our CG is a little different. But we're not at that point, even with Gemini. And you can imagine how, how like, like a resource taxing that would be to take snapshots of a video of this show at 24 frames per second per hour. You know, that corpus of knowledge is pretty rough. So we haven't seen that yet. Even Gemini is great at taking screenshots. The difference that I see over and over in like Beth, I don't know if you agree is like, Gemini's phenomenal at telling you all the things that were going on in that scene or in that particular screen. I mean, it's just so, so good at it. Even still I pushed and pushed and pushed for the project Bruno. Like, hey, are any of these flash models comparable to like 3.8, you know, 3.7 flash,
3.8 was alive. The 3.7 flash is comparable to 3.1 pro, which by the way, I'm still using on, because there's been no pro model to come after that. And it's like, no, 3.1 pro is still the standard. And it's like, it's still doing a better job than the flash. Now, we probably aren't that like a flash versus a pro that makes sense. But 3.1 pro to 3.7 flash, it wouldn't be weird for 3.7 flash to actually be better than that older pro. But Gemini just simply hasn't come out with that model. So that's what I'm really waiting on. I mean, if everybody's releasing models this week, hey, Gemini, you know, please, please give us a 3.5, a 3.6 pro. I would love that because I bet you it's going to be phenomenal on the multimodal side of video, which would be super, super cool. So yes, anyway, I feel like I feel like every time they let it, they release a pro, it's like, you're not as good as the others. Like, try again.
Then they release a less, it's like, oh my God, this is amazing. You're so good. Like, I don't know that we've given them incentive to come out with a pro that people are going to be disappointed. I would love it because the other models don't hold up to it. You know, I mean, anthropic look. Inthropic in-clog code, you guys saw what it was able to do. It, it, it, it, FFF pegged and did the whole thing. And I think back to what that used to take me, what that used to entail for me to do with Project Bruno, ah, year ago, whenever I started working on that process. And it was complex when I was paying render through a worker and stuff like that. And clogged code shrugged shoulders at me and was like, I'll just do it. Push back from the chat. Jeff has never said, Gemini was so good. Never. I'll give it Jeff specifically for multimodal video. You know, like, we got a narrow in music. We're not saying, we're not saying globally. We're saying specifically for video and multi-mode. And sound. Yes. Like if you, and sound like if you give it a sass down, yeah.
Jim, you haven't done much with sound, I guess, but yeah, yeah, yeah. It's really good. It sounds as well. So I think I was 5.5 got points. No, maybe it was, maybe it was codex. Sorry. There's just like all of the things that were. They're all interchangeable now. And I'm like, the reason I'm saying that is that, oh, I think the person who post, who like retweeted it, repoved, I don't know, re-exed it. So I didn't want to. Game from is a, is a open AI fan. That's what I'm saying. I just scrolled back for this one, but I didn't want to miss it because Cisco did say, I used 5.5 a lot last night. It's back. That that that so far. Hopefully it won't be. It won't start to degrade in a month like the last big club release. Now, you know, Cisco, you've been right there with me and, and, and others, you know, on the show and stuff like that. I'm saying like that, yeah, there was just that like that three week period where Opus was. We were going back to 4.8.
Bet you were like, I'm just going to use 4.8. I did the same thing. I don't like forget about it. I would see hints of glory at, at, at, at Babel and then it, or, if I was just, I'd rather use 4.8. I think it was because I was using all of Opus 5's intelligence and you guys were like left, with the drags. There you go, Cisco. Yeah. Sandy, actually, it could have been me too because I ran out of tokens on open AI. So I had to go to cloud. Yeah. I had non-stop. But I will say for the slowdown, can I just say I make so excited because I ran on a token on Fable, but now I can use Opus 5.5, which is smarter. And I don't have a limit on how much I can use on it. So like, well, the five hour limit, but, yeah, well, five hour, but I don't have like a max limit on like Fable does. And I was like, I feel like I'm hacking the system right now, but like, that's what it felt like. Because I was, what do you guys think is like, so look, I, I build in, I do, I do prefer. It's also where my account stuff. So I do prefer to being cloud code, but I can't
ignore the cost difference. And so for you guys, now that 5.5 is what it is, we've talked about the prices. We know it well, we, we hope that it's a capable model. So far, so good. I agree with Cisco personally that yes, since yesterday, I've seen hints of former greatness from 4.8. You know, I feel like maybe there's been some moves in the right direction there. But I mean, we know soul is a damn good model. And 6 is probably even better than I haven't even tested it yet, really, other than that side by side comparison. I haven't even read that yet. So I'm not even there. Like what's going to keep people coming back to I entropic? You have a more costly model of 5 hour limit. The, uh, yeah, I don't know. You don't have a 5 hour limit on, on, uh, on soul, correct? They brought it back. I don't know if they let it go again, but they did bring it back when they, the way I'm looking at it though, like, look at up, at least for me. I can't afford, and not afford like as in like money wise.
I can't, uh, what's the word? I guess I can't afford to, to, this sounds really snobby. I can't afford to have the lesser model that is might have more mistakes. Um, because my project is so, not to like hype up, it's very complex. And there's so much, there's so much data going every which direction, um, and getting parsed out and broken up and used in different ways. Um, uh, if, if I have a model that goes in there and is lazy, um, and screws one thing up, it's just a chain reaction. It'll just start breaking things. Not that it's that fragile, but like, I mean, it is with any, software development. If you break one, but one thing, you're going to find 20 more bugs. And so, um, or when you're developing a feature, you're going to always find bugs.
And then sorting out those bugs and making sure that when you're fixing those bugs, the splash radius is not breaking other things. So, um, I really want to, to use the less, the cheaper models, but I, we're so close to the end. I'm, I'm like, you know what? Gareth's talking about his project. We're so close to the end of the project. The project. Yes. That was not an expression about P doom here. That's right. Yes. We're so, I didn't think about that. We're so close to the end and launching this, um, my project at work. I don't need extra work to pile on. And so, um, yeah, that's kind of where I'm at. And so I'm here. I will get, probably get to a point where I will say, okay, I can use soul, um, six, um, and, but right now I'm just, I can't afford it. But I've seen some great things with, um, a lot of people talk about the models and
talking about, okay, jev up to these models. Um, and so let me share a screen again about the GDP, Val leaderboard. Yeah. And this is the argument for why an enterprise might not choose to go after the cost optimization here because of the vast difference in real world work performance tasks. This is the, this is the leaderboard, the test of how well these models perform on real tasks that you put them to get to just like the, you know, a Jova processes being, you know, laid out by Brian and as Garth's doing his workflow for his company, this is where the difference is, look at the difference in the score between GPT 6 Astra here at 1542 and GPT 6 sold under 1500
and Opus 5.5 at 1846. Now, what if you have mission critical work that you're doing and a lot of these are in the flow of impact on revenue, are you going to save that little bit of money and get this, you know, no, because it's more money fixing the problems that miss in the first place. That's the way I look at it. Right. And so that's what Aldo was just saying. And Aldo Dev was just saying in the, he was just literally saying that, yeah, I mean, Hikou is cheap, but if it costs four times as long and a hundred follow ups to get anything that works, that's not affordable. Yeah, exactly. Yeah, exactly. Comment above that, I guess they, sorry, well, do I, I'm assuming gender, I don't know. We are in just the new release, right? So, hey, everything's really good now. We want to see what's happening next week after everything shakes out because all the comments,
Opus ages like milk, it's nice and fresh and many grittles and runs through much. And we don't know, but there is, and I would say across the board, I don't know about rock because I haven't used it when it comes out and I only use it to explain expose to me. But, but yes, they come out there like, we can do all of the really good things. And then everyone uses it and it's like, all right, now we know what it's going to cost or takes. So that got itself. So we're going to, we're going to, we're going to make it work for at a level we can provide everyone. And I thought, I felt like that, open AI was overshadowed by an anthropic yesterday. I was like, oh, this is all you got. This is all you gave us this week. Oh, thanks, open AI. Like, I don't, I'll keep, keep going. Still on that before we leave the subject of models, just wanted to mention that deep sea, we haven't talked about a lot.
But this month, they released 4.1 flash, which hit number one on the open router leader board. Meaning that's the model that among those who are using open router to determine which the best model is to use for individual component tasks. That's the one that's, that's winning. And it jumped 172% in usage this week, this week alone. So deep sea is an alternative out there that doesn't, doesn't score on the benchmarks, you know, anywhere near where the frontiers are. But it's, it has practical use. And I think that gets to my underlying point about all of this, which is most of these miles can do what you want them to do. If you go through the process of detailing your harness and your prompts and your evaluations and determine that they're sufficient to that task. And we've talked about Jev, which is really very simple.
It's just making deterministic calculations about the probabilities of whether something is or isn't, right? And that can be slotted into workflows in a way that's incredibly cheap, almost free. So there's a bigger process that I think you could observe in a Jova that Brian's working on where he's building these side by side evaluations and determining whether or not there's a qualitative difference sufficient to justify spending $3 instead of one, right? Does it like if that's just a one time cost to build something that's going to have a long term asset value, but its quality is higher. Maybe that justifies spending an extra $2 up front to do that. And that's the kind of just thinking has to happen on the human side when you're building out the systems. And what we talked about yesterday that Brian mentioned he's doing, but this makes that that process even more valuable is training setting up the Jev portion.
So it's good enough that it doesn't need your cognitive load to judge, right? So now you're training Jev to be able to do that first round of, did you, are there bugs that need to be fixed? Is this complete, right? Or whatever that is? And the part part of that when we were teaching for, yes, the AI exchange, we were teaching prompting for the AI exchange was trying to bring, it trying to help people break things out into system thinking. This is one little function stuff goes in something happens comes out. Now there's another little function. We don't have to, we don't have to help people do that anymore because AI can do that break it into discrete functions. Let's create a test and then let's compound on it. Let's improve. And then it may be so good that you can combine some of those now, but yeah, exciting stuff. Cool. Sorry, sorry, this is just a, if you're going to move on, this is one last thing, because I forgot to mention it was I did test Opus 5.5 on our thumbnails that we've been doing for this show.
I think I talked about that, but I was only able to use Fable 5.1 previously, Opus 5 just was falling down on the job just a little bit, but enough that it was causing revisions and that just costs tokens of money and my time and all that stuff. So I've been using 5.1 daily to do this show now what it does just for anybody who doesn't remember talking about is in here every single day. It literally looks at this show so this hour long broadcast that we're going to do it looks at the transcript. It looks at the transcript to figure out what was talked about in order to figure out what the couple words should be on a thumbnail and so on and so forth. But it also takes screen grabs speaking of multi model. I'm using in traffic on this. I'm not using Gemini and it looks at the screen grabs and looks at our faces and tries to determine out of like it does like 200 a piece. So for the four of us, it'll look at about 200 face tracking a piece and it'll try to figure out who's looking at the camera. And what's the best screen grab of me, Andy, Garrett and Beth to use to then go create the thumbnail.
It's complex. It's like it is and now Carl's jumping into so be Carl to rather than the by bus. So who goes where? How does everybody gets stacked up? If if Carrot, and I'm only picking on you, Garrett, because you come back to most and he's the easiest and he's always like, Andy's great. I'm like, I know we've talked about this. It's never problem with Andy's face. You know, but there's very little there's very little change here. It's all the same. I guess it always likes you, Andy. It's that hasn't once given me a problem with you. But like, finally someone likes me. Wow. Yeah. Opus 5.5 and then fabled it as well. But Garrett, because of the backwards hat sometimes and then the reverse, like my my hat says hardly on it, but I have my mirror reverse so it can read it. But like, Garrett, the other day you had a I know that might been Oregon had on actually and it was backwards in your in your mirror. Right. So it had to figure out to flip that. And that was like complex because it was like, Oh, it's Oregon backwards. Let me reverse that because it's supposed to be what we're on here. So the only thing it did yesterday was it said to me what what messes us up more than any other thing with this thumbnail is these damn title bars underneath here.
Because it obscures part of our shirts as you can see on the lot. Well, yesterday, and I'll be curious if it does today. I get one that you probably won't be able to read them. My shirt says vipers are Andy says, but now it's in the transcript. So maybe it will. Or maybe it won't know the Carl has a Colorado hoodie on. But Garrett, he couldn't figure out your pattern of your shirt yesterday. And so it said it was the best. It was like couldn't figure out the pattern on Garrett. His name was covering it. I put him in a black hoodie and I was like, I mean, that's normal. I mean, I don't know why. But like, and it was a good looking black hoodie. I was like, wow, even alone. I don't know. Nobody gets to put them now. You know, as long as it gets the rest right. And the other thing it'll do to is like right now Carl usually is looking off camera. Yes, I'm saying on your car, but you're usually looking off camera. Well, well, the reason why I'm looking off camera is I have a like like my camera here. So it's weird just staring at the camera the entire time. Well, I know all of you are staring at your computer with your camera up here.
Right. So I'm just like, okay, sounds good. So it's just that's a line. Oh, no, that's not the reason. Carl knows that his from the right profile is preferable to the wrap. That's right. So it's harder. Well, yes. He's hard to tell. Oh, turn back Carl. Turn back. So anyway, I got into whole sighting to tell you that like 5.5 has been able to handle this. But I guess, but it's also it's an interesting conversation about the complexities of what it's trying to do. And actually, it's a statement to how good it is at actually doing the thing. I mean, the idea that it is finding the five of us, finding which way our faces are pointing. So I'm not looking, I mean, if I stop for a couple of seconds, maybe it'll grab this. And if I smile nice, I'm looking at the camera, but mostly I'm looking down at you guys, right? So now yesterday, best face was kind of like, go do what you were just doing that. Just turn a little bit to the side. Just a little bit. Okay.
So beth- oh no, but now you're looking like Carl, you're way off to the side. You're not good to me. Now that now that has to be off to the left. But if you slightly look off, where should your eyes be pointing just off or back at the camera? And so sometimes we have to fix that with AI too. So yeah, you guys are all fucking with my suit of language. You guys are all messing with my post-processing today. This will be a mess. But my point is 5.5 handled it down. Three other shows that I had done in the past. Thank you. Thank you. Wow. Wow, I can't wait to see my- hey, that'll be the game now. It's like, what will tomorrow's thumbnail look like? Yeah. There'll be four wrong pictures in Andy looking amazing. That's what it's going to be. It's going to be four-shitty pictures, four-shitty images. I'm going to look like big chubby cheeks or whatever. And in Andy's going to be like, buddy, Jesus, all up in the thing looking great. So you wait. I'll show to you tomorrow on the show. Yeah. I'll show you tomorrow.
God forbid we ever actually beat in person and then we'll see who actually looks great. Okay, sorry, sorry, Andy. I went way off top of it on that. No, that was my process yesterday. So it's interesting. Yeah, but it's a peak into just how complex and possibly dithering a AI process can become. You know, why are we spending tokens on that? That's a good question. So I want to talk about agents now. So Metis Muse agent surpassed 500,000 users in one week. So 500,000 and that's a function of the ease of implementation that was built into their product design. And of course, they're enormous catchment of users that they have through Metis platforms, Facebook and what's that? I'm sorry, not with what's that, but Instagram in particular.
Now that's all I have to say about Muse. Muse is very good. But in the chat yesterday, the question comes up like, what kind of personal information have we now added to Metis accumulation? If you're already a Facebook user and you've given it everything about yourself in your lifetime, maybe you don't care. But I don't use Facebook. So I'm not going to use Muse because I don't want to give Meta more and more ammunition on that front about my life personally. Just an attitude and a personal decision. I want to talk about Stripe. Stripe is the most AI capable and most integrated MCP based integration for payment processing in e-commerce. And that's relevant to agentic commerce. And what they just added are Web MCP based checkout tools. So Web MCP, I guess, uses the browser interface rather than the behind the scenes thing.
But it then accesses an MCP server and they produce checkout tools across 7.8 million businesses on the Stripe platform, giving AI shopping agents a structured way to complete purchases. So instead of having to have the agent click through a rendered web page, now there's a behind the scenes server that can process the entire checkout flow all using something that's faster, 40% faster than even document object model automation built onto the page. We talked long ago, like last year about DOM or document object model being a way to avoid the problem of having to have an agent or a computer use kind of tool look at the screen and try to first interpret it and then try to move a mouse to the position on the screen to take an action on it. Instead, document object model says on this web page there are these things including this clickable object
and you can call that clickable object and act on it. Now the MCP server has a structured process that uses far fewer tokens than even that than DOM model. So that's pretty important with respect to agent agent actions in e-commerce, which is a big piece of where we spend our money now. But remember that agents are blocked by Amazon, which is the big gorilla in the room. And that's that story that we talked about yesterday or before. One thing that I what I wonder about that is yeah, you can block it on from a muse or instinct or whatever. But like well if an agent can use browsers, but I don't know how you'd be able to block that. They're apparently doing it from IP. So muse comes from a cloud-based yeah right. So they're blocking that IP entirely. You can't you know it can't be in that. What I mean is like if I give my agent, cloud, codex, hey open five tabs in browser or use my
chrome plugin, go find block. There's no way you can block that. I don't I haven't seen anything that I think it's the agent agent interface that they're trying to block. They don't find if you sitting at your computer have a computer use tool that's going and doing it. They what they're trying to block I think is the agent agent communication that's that's being promoted by stripe. Right, but what I don't have to be at my computer for the agent to be using my chrome. Yeah, right. That's what I would tell people is it was like your blocking doesn't I was like okay, how are you going to block me then if I decide hey 10 agents 10 10 tabs go to town and then tell cloud the same thing 10 agents 10. I was like oh and when you guys are done I do have another meme that I want to share with you pretty funny. Okay, about exactly what I'm just talking about stupid. Let me just quick throw out one last agent piece of news and that is
this was reported in the near on this morning. There's a new study that ran 100 AI agents through a simulated economy where they could transact and and that those agents were able to transact successfully with each other. So they're doing transactions in a simulated economy that's not an e-commerce thing that's oh I have a service you want that service here's what the price is here's wages I have to pay to my you know sub agents to do that sort of thing in that internal simulation wages and prices didn't move which is not normal for a market a market if there's a shock that's introduced that's like oh here there's a supply constraint in this area then prices change wages change in response to longer term price well weirdly in this agent economy it those didn't move it was like the agents were kind of frozen around a certain price structure so the conclusion of that
particular study was that agent individual agents while competent to do these transactions don't automatically produce a good collective coordination on the price dimension like price elasticity in response to other things so that's weird that's concerning because we expect that the agent agent economy is going to be you know in some ways more efficient and there's the possibility in my mind that okay maybe that creates stable pricing and wages in a way that inflation you know it is the opposite of right so maybe it's actually better but I don't think that was the conclusion drawn by the study and I haven't read the study so well there was that seems to be one of the stories of AI being used anyway right like you see oh I saved five hours I saved 10 hours but generally across the board the there was a study that came out the about law students they tested law students doing
real-life law tasks and they did it in like 20% less time but they made the same level of errors that were made when doing it longer so it's it didn't actually it increased time efficiency but it did not create value and that's there's another thing that's come out in recent research which is companies which implement AI they do the coding step faster but they don't deliver products to the market faster because the bottleneck shifts to the downstream requirements to review and and you know finalize testing and implementation and evaluation all of that's taking almost as much time additionally as the efficiency gains at the coding step and so we haven't seen an overall end-to-end
you know conception to product delivery improvement in the sort of supply chain of delivery of products so there there are all of these things that have to get worked out in order for this to be a true gain in the economy so now I'm curious about what you've got Carl please tell it that's something funny you always do well okay so what I wanted to show was just two things and I thought it was it it's funny but also it's like a little scary at the same time and I think we kind of understand here this over we got the share oh you need an admin to share it for you yep we're all doing a thanks here there we go so I got Chinese man as AI running 50 social media accounts 24 seven which is kind of crazy so whatever and this is why whenever you know x instagram whatever it is
I think it's I think it's mostly x two you're like who is actually responding who is posting and this is why it's like you gotta be super super like every single post is now critical on why is that person posting is that a troll is that a bot what does that look like right because something like that that's a job for jive yeah that's a job for jive exactly this is a troll this a bot so look at this this is a social media farm yeah like and that's the thing is like I now I don't even know if the the the post or the description matches the the video but it's like upload AI slop across 150 TikTok accounts using the as I go so whether this is true or not it's like but you can wow you can do that right I'm pretty sure this is true yeah no I've seen I've seen many videos of
AI farm like social media farms and they look something similar to this or much much bigger this one that's really tiny I've seen like warehouses yeah where like where and if you're just listening to us there's a wall with hundreds of phones on it yeah and they're all plugged in and apparently in order to manage these accounts it's not as simple as having them happen behind the scenes you've got to actually have a real phone in order to have that phone be on the network and identified by whoever it is TikTok or or you know Instagram as a real account so wow that's that's crazy yeah it's five rows and it's probably got 30 to 40 phones or rows at least yeah yeah yeah this is why it's so hard to break through from a content side but also because it's so easy to create right the just these two you can see okay it's like you got bots literally or farms
creating so much content I think like that's why like how do you be like human experience kind of I think how do you still connect that because that is so easy to do but the human connection is very hard still so right and and without even without this like any of us can just pump out tons and tons and tons and tons of content with really not thinking about anything either well in the business use case that is being shared is that they grew their uh whatever their tool was to 37k right monthly um what what most people want uh like I would like to we would like to grow the channel to 10,000 subscribers and then we'd like to grow it more than 10,000 subscribers we're not doing that obviously because we're not at 10,000 subscribers yet please tell your friends subscribe to the thing if you're listening that would really help us out
but most people will feel good about results that they can get with like running five API calls every three hours or something right yeah we do not endorse uh should say um okay uh I want I think uh it's Garrett had something about dark trace you know I think best is better okay so we go that okay I keep saying I got a thing oh sorry sorry well and I'm glad you're here for this one girl because I think um that this is actually really interesting it's a convert it's a coming from um runway labs and yeah share that's fine um and it is talking about responsive generative AI interfaces okay so you're not making a portrait one or a widescreen one or anything else these are just
interfaces that automatically change like that that one we shared like three four months ago that was like it you could just kept diving into it and it generate as it was as you clicked and it understands the um the dimension of the window so you can see in this example where it's a storage unit it smushes everything right understands there are three-dimensional thing where it wasn't smushed before and you can interact with it are you sharing something about oh come on yes am I not sharing no yeah all right here we go that's fine oh oh we're sharing okay all right let me go oh yeah um hey okay so let me go back to scrolling down that page uh which means I need to be there okay so that we see people coming down uh ski slope um then um these are the the hot air balloons you can see how it shifts
um and this is showing us the shift between um between a landscape phone and or like a landscape regular computer view and then a portrait phone view but you could do this with any size frame you want it right like you instead of just moving pictures you could move uh what is focused on this understands the 3d aspect of things and then you can interact with it and this is going to take the towel off the off the clothesline and put it in yeah let's go back and put it in the basket right because things can you can you can you can interact so this for me Carl gets back to like what is the new website and what is it going to give you the ability to do that's yeah this is a runway lab thing it's not out yet but they're talking about it in like yes websites are
now just going to be interactive video yeah we saw this kind of starting I think three or four months ago um with that one I figured what the company's name was yeah I think you and I were thinking about the exact same thing where you move the person's like clothes and I think I forgot is like a world vision thing world thing is it I can't remember the company anymore yeah I don't remember but you could like basically click it showed like a architecture like a city and you could click on the city and then it would zoom in and generate yeah yeah yeah yeah yeah yeah yeah yeah and you can keep diving in like you could go into like a clock and then it would tell you like the internals of the clock and it all came for it was like an endless generation of of things and it was really cool but those are things that um I think those are things that have a static back end so you can go deeper and further out but I don't think you can move stuff around in that correct and that seems to be what
runway is coming to it's it's an interesting thing and this goes back to the Sam Altman post that was deleted because Sam the post that got deleted was Sam saying sorry I met voice not video and I found this post because uh Cristobal uh who's the CEO when the co-CEO runway said oh no worries we have both voice and video over here and look what we're doing that so yeah wouldn't that be interesting if actually open AI came out with something similar to that runway I don't think that's their intention yeah I don't know if they're out of the video aren't they he he wrote a post Sam wrote a post that reference that he wants to be the best at video and then he amended it he didn't amend it he posted a thing underneath that said I didn't mean video I met voice
and then that got deleted in the time from when I saw this morning and I mentioned it on the show and now it says we love developers so we don't know what's happening yeah well I think there's still we have all the tools regardless of what you have right yeah like any any tool right now can use any program good it's your description whether it's good or not but we all know one generation two generation from now so it's gonna get it won't matter it really won't matter about the the program I'm just excited that we can use AI to use programs that none of us had any business using or would have taken us like months or we used to train on I'm like I have no business using a blender or on real engine but now I do I'm like okay I can just use this and then any program any software anything that you have so for anyone who's still using this
like really old software which I was working with my dad like the past couple days on yeah you like we're working on a database in a virtual machine right so we had I just had AI to go into the vert use the vert open the virtual machine open windows 95 I think and use windows 95 to go into said program and extract said data and I was like it can do blender it can I forgive forget blender can I just point can I just point these new tools at Microsoft 365 and say please don't ever make me go in there he has I don't want to go into your dump software ever please don't make me go in there you know I don't know I just go find the thing just go find the thing I need to do don't tell me where it is just to go do the thing for me I don't know anything about Microsoft and I don't know anything about all this weird stuff that Microsoft puts out like yeah and yeah I think what
a prison on our team is like we had to implement this thing Carl I don't know how to do it so I just point to codex it is so awesome I know we have to touch Microsoft. Here's the sad part Carl Microsoft could do this with co-pilot but we all know that it'll just make it more complicated Cisco was at a training for co-pilot yesterday Cisco report back I know you're doing a meeting right now but report back I'm not a so if you he was here today so you know I won't go I don't know but then he he was here and then he had to he had a bounce he will be back I know you're listening to this afterwards yeah let us know well speaking of we're probably a good time to wrap today's show up I'm realizing our our two times are off again because of whatever the issue was we were live we realized it at the beginning it's only because I said to Andy literally before we went live today as like today's the day everything's gonna be great and then it was it or kind of was it though I shot ourselves in the foot immediately so I won't say anything tomorrow okay I do have I have to have one thing that we have to we have to let the audience know
about and that is that yesterday at the you know the United Nations Donald Trump went through a long list of things but they also declared he declared that we are no longer going to call AI AI it's not artificial intelligence he said there's a new term for it and we're going to use it's super intelligence and so now from this point forward this has to be the Daily ASI show yeah and that's that's what one on the x-pole that's how that choice got made yes it was it was a it was a poll it was extreme intelligence super intelligence superior intelligence I voted for a superior intelligence it was it was supreme intelligence not supreme yeah that's right and he thought that he took that off because he thought it wasn't popular because of the court yeah well somebody forgot to tell him that super intelligence is a term of art in AI
that that is already in existence and is being used and should not be conflated with AI so anyway that but you know he's in charge of worldwide branding now so alongside lake america you've got certain intelligence Ontario lake Ontario wow the way it says super high intelligence talk show otherwise known as s hi yeah it's the it's the shit show uh that's a welcome warning I come to the the super high intelligence talk show uh with your hood yeah yeah you can take that pun along that's great although you're the best for operating it up at the end of the show I appreciate that kind of comedy all day every day uh all right um we'll wrap it up for today but thanks rangin out with uh stay tuned for tomorrow uh the figure out what the thumbnail is going to look like uh for today's show um we'll I'll share that tomorrow also curious uh whether hopeless uh 5.6 see stop messing with it guys uh stay tuned to maro to see whether one of the shows
Garrett I feel like you should do this too um see if um uh soul six or opus 5.5 so I can get that right right to better song out of the gate for suno I feel like that would be a good test we've talked about uh all the models and what they can do but can they write a banger of a hit for 2026 and not have it sound like absolute crap now they'll get it Garrett I know you you pre-training you to like you cannot give it any information you just have to say right right of 2020 26 hit you can't tell it all your secret squirrel stuff like how to write good lyrics and you're on mute it's pontificating my problem is is my secret squirrel stuff is in the memories nope nope doesn't count it knows nope it's not to structure a song it nope api it sorry doesn't count all right all right that's it we'll see you guys tomorrow for a lot more um shi fun super ian diligence tuck
More episodes

