Skip to content
TrackPodcasts
technologySep 22, 20261:00:41

Amazon Blocks Meta's Muse

Get every episode summarized

Each time The Daily AI Show publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

About this episode

“We are 15 minutes later than we wanted to be because we ran into all sorts of fun Tuesday problems and we are trying to get it all sorted out. And I just, sorry, that's me replaying it.”From the transcript

The episode focused on JEV, a specialized decision model that could change how businesses build AI agents. Brian demonstrated its potential for moderating live chats without removing constructive criticism and discussed using it to score client projects, test different scenarios and check LLM outputs. Gareth shared results from 60 test cases in which JEV ran roughly 7.5 times faster and cost 95 percent less than Gemini 2.5 Flash for his decision-making tasks. The hosts examined how to combine specialized models with LLMs, while raising concerns about JEV's data terms and adopting it in production. Other news included Alibaba's Qwen appearing in the U.S. Federal Register's search interface, international calls for frontier AI oversight, Grok 4.7, anticipated OpenAI model updates and Amazon blocking Meta's Muse shopping agent. The final segment featured Brian's AI-first AJOVA Journeys command center. He demonstrated a system that manages video production, generates thumbnails and Shorts, tracks leads, plans client communications and monitors costs. The first completed video required substantial recording time, but the system aims to learn from Amanda's feedback and improve with every production cycle.

The Daily AI Show Live_ For Real September 22_ 2026.txt


Key Points Discussed


00:00:17 Episode Intro And Streaming Problems

00:03:33 JEV Filters Negative Comments From Live Chats

00:06:39 Building Comment Moderation Into AJOVA Journeys

00:09:20 Why Specialized Decision Models Matter

00:13:21 Using JEV For Client Project Health Scores

00:16:34 JEV Versus Gemini: Gareth's Speed And Cost Tests

00:19:20 Detecting Conflicting Information With JEV

00:21:24 Data Privacy Concerns And Early Adoption

00:23:14 Combining JEV With LLMs In Agent Workflows

00:26:54 Alibaba's Qwen Appears In Federal Register Search

00:29:32 International Leaders Call For Frontier AI Oversight

00:31:58 Grok 4.7 And Its Electrical Engineering Results

00:34:20 Anticipating OpenAI's Next Sol Release

00:37:30 Amazon Blocks Meta Muse Shopping Agents

00:39:09 Shopify Embraces AI-Powered Shopping

00:41:00 Brian Tests Muse For Personal Shopping

00:43:42 Inside The AJOVA Journeys AI-First Business

00:45:21 An AI Command Center For Daily Business Tasks

00:46:54 Turning 33 Recorded Clips Into A Finished Video

00:48:44 Automated Content Ideas, Thumbnails And Shorts

00:49:21 A Built-In CRM And Client Follow-Up System

00:50:30 Tracking AI Production Costs

00:51:00 The First AI-Produced Cruise Video Goes Live

00:56:30 Why The First Video Still Took Hours To Record

00:57:38 Building Feedback Loops Into Every Workflow

01:00:16 Episode Wrap-Up


The Daily AI Show Co Hosts: Brian Maucere, Beth Lyons, Andy Halliday, Gareth Hood.

Hosts & guests

Transcript ready

1,450 searchable segments. Every word is indexed and playable.

Amazon Blocks Meta's Muse

The Daily AI Show

0:00
1:00:41

Full transcript

The Daily AI Show — Amazon Blocks Meta's Muse. Machine-transcribed; use the interactive transcript above to jump the player to any line.

Hey, what's going on everybody? Welcome to the Daily AI show. We are 15 minutes later than we wanted to be because we ran into all sorts of fun Tuesday problems and we are trying to get it all sorted out. We've been here, we're y'all being. It's just what's happening. Jeff says he found it, which is good. And I just, sorry, that's me replaying it. I was just validating that we are in fact live all that YouTube, which we are. Okay, I gotta let Garth know he's waiting in the shows channel that I had to go up to see it as a reply to today's show. Yeah, yeah, so sorry for being late. Thank you for being here. Hey, hey, Cisco, hey, Jeff, thanks for being patient with us. And also letting us know that we weren't live. We were actually live on LinkedIn. We weren't live on YouTube due to some errors with which stream yarded how they set things up. We were laughing and saying it's been over three years,

800 and what, this is episode 817. And this was the day. This is the day of minutes after all the time. So we're here, we're ready, we're resetting. We're just slightly later than we were to get going. Okay, so today is, just not, nope. Today is September 22nd. We could do this. Focus. Hey, Garth. Red day. And today it was like we match September 22nd, 2026. And with me today, now our best Garth, Andy and Brian, it took us so long to get this right that we actually waited long enough for Garth, join us live in Suffolite. So glad you're here. I'm glad this wasn't the one. I was just saying on the other live that Ann was saying she was happy to make it. I'm so glad actually she didn't on this one because it would just been more confusing to get the links out and stuff like that. So we're good, we're reset. And as I was saying on the first live earlier,

was you, while three at least, you actually were here yesterday. I wasn't, but I was at the beginning of the show when we were dealing with other issues. Just on Monday, Tuesday. And I would say Andy or Beth, one of you were saying to the other one, like, oh, look, we're really early on AI. And I was saying that I wanted to be on the show because I wanted to say, yeah, and Jeff too, we were early on Jeff. And so Beth, I know you had played with it. I had done some demos with it because I had gotten that, not early access. I just got fast access. Like you did Beth. I had it within 24 hours. And I had some time in the morning to play with it. I was showing some like what I was getting excited about in terms of like just basic stuff like transcripts and digging through it. You know, I think there's been a lot out there on Jeff and what it is. But if you don't know, it's not an LLM, but it is a model that basically you bake into the backend exactly which are criteria are. So it's not actually, I wouldn't call it revolutionary by any means. There's been plenty of systems where you give it yes or no's or you give it a criteria.

Or I forget the third, it's Nol's criteria and I forget what the third one is for how Jeff is set up. But, and now it seems like all the other newsletters are like talking about it a ton, which is great. One of the things that came out of that that I wanted to start to show with is that somebody on X-demode running Jeff to see how fast it can remove negative comments from chat. So literally like we have live chat right now here. Now ours are manageable. We're able to talk to everybody. Everybody's pretty good. We have removed comments and people and blocked them in the past because of obvious reasons. Not everybody is here for the right reasons. But you can imagine if you're a Twitch game streamer or something to that effect where it's quite literally just a stream of comments, just running comments. I think everybody has seen some sort of version of that. And this guy, Dev Ed, DeV space and then the name Ed Ed. He was showing that he was able to basically

get this dialed in enough that it was getting rid of negative comments but not constructive criticism. Sarcasm is actually really, really hard. But you can imagine you can write enough criteria into Jeb so that it's able to filter and deal with all those different fringe cases that aren't exactly just a hate speech or something that would be easier to pick up. So what I found interesting about this guy is like, he, this guy, Dev Ed is responding to people. And a lot of people are saying, hey, this is a great use case. What I find interesting is people going, there goes free speech. There goes, you know, like you got the doom, you got the doom's errors or whatever. And I guess the take I take on this is like I do with this show. This show is not a democracy. It's not nor is any social media that you independently run. This would go for anybody. And what I mean by that is you have posts here who have a Thursday of control to say, nope, you don't exist in our world. That's part of social media. And so I find it funny when people say,

we're interesting I guess, when people have comments, not that every comment was like this, that was like, oh, free speech. As if the goal of having people being able to spam you in a live chat was somehow the best we have to offer from free speech. And so I just find it interesting because if you're the person running, you have every right to run something like Dev in there and filter out the stuff. In fact, I love this example because another demo I want to get into other news and I'll show you later some updates from what I've been building with my wife Amanda. That's exactly something we were in thinking about Jeff, she wasn't I wasn't because it wasn't around yet, but we were absolutely already thinking about this. Amanda and I are both kind of the same in this sense which is funny because we're so different in some of the other ways, but I wear my heart in my sleeve and I don't do great. I know I don't do great with like a bunch of negative comments. I feel them, I think a lot of people do. Some people are able to see that in social media and rolls off their back and they go, what do I care? I'm not gonna give them another second of my life.

I'll think about it for the day. So I always kind of felt like with this show or anything else or social media, I'm not the type of person that can just read a whole bunch of negativity and it would have an impact. So we were already building a solution for Amanda where it was sort of doing some pre filtering as she started to put out content about, cruises for our Jovich Ernie's and things as we've been talking about in the show as sort of side hustle. That's something that's important to us. And so we had already thought about like, hey, when she has this command center, she will interact directly on YouTube. She will interact directly on Facebook. She will interact with the comments that should have her interaction with them. We're already filtering that out. But I like Jeva a whole lot more. Now we don't really need Jev for real time because she's not doing lives like we are right now. But I think it's a great use case and I think it's a great example of how this bottle can be used to do some really quick, let's call it sorting. It's basically, it got all its criteria built in.

You have to work to do that. It doesn't magically happen. But as I showed last week, I was instantly able to take Jev connected to Cloud Code and Cloud Code ran and wrote all the JSON for the criteria. I didn't have to write any of them. So if anybody's worried about like, oh, but I don't know how to do that, how to write criteria, well, you can use a language model like a cloud or any of the language models to explain with really good details, actually like a PRD. This is what it means to grade. This is what it means for a yes or no answer. This is what it means for my criteria. And you get really specific and you get the LLM like you do that, ask me back, interview me, do all the things, then write the criteria and then have an interrup process to get better at that. And I think that's a really fast way to work with Jev. You're on mute, Beth. Fine, Gareth. Thank you. Everybody's done well. What are these is going to be the last straw for Beth today? That's all I'm showing you.

I'm going to show you. The daily I show, but as mine, there we go. What are we doing? This is a really good point. And I had the conversation. I actually think it was with Anne. But I think people get the impression because of what we're doing with AI that we're really good at the thing outside of AI, right? Like Anne was saying something about, well, I'm not, maybe it was obsidian. Like I'm not touching obsidian. I was like, neither am I. I don't touch blender. I don't code Python by hand, right? Because AI offers the opportunity that you don't have to. But you are still talking about the actual thing that is the result. It just doesn't mean that all of us could do it without AI at the same time, right? And I think that's important, Brian. I want to just comment on what I see as this explosion of interest in something that is really clear in my mind now

as an evolution of our understanding of how all the different components of the AI work together. And we've made it with the appearance of Jeff, which was tuned and created from its inception to do something very different from what a generative LLM is doing. It's really kind of opened our eyes to how that all fits into the agentic workflows that we're now crafting and have the tools to do. And what Jeff is in my mind is a discriminator model. It's not designed to write an essay for you. They can't do that. But interestingly, when you talk about the criteria that you put in, it's pretty sophisticated in respect that you don't have to write in code language. You can actually write sentences into your description of what the state and what the judgment calls are that you want the model to create.

That happens in a way that doesn't require a deep understanding of code, although I would still code a cloud to actually generate those entries in order to create those discriminations. But basically, what's happened is, before we always thought of AI and the tools that we use as, OK, we're going to have a message delivered either through text typing into the keyboard or we're going to talk to the model. And then the model is going to process that and it's going to come back with text. And embedded in that text is a decision or a recommendation or suggestion and all of that. But what Jeff does is it says, there's a lot of things where you don't need a whole lot of reasoning. You just need a discrimination to be made. You need to make a choice. Pick one option or you need to have a score, which one of these scores highest on a certain result. Or just give me an answer.

Yes or no is this true. And all of those things together create this highly efficient and very inexpensive model that's really incredibly capable in the vast majority of things that we want an AI model to do. Because the deep research process has a bunch of intermediate steps, for example. And many of those are these discrimination judgments that have to be made about what to do next in the process of arriving at the collection of research. And then finally, the LLM can take all of that and build out of that as enormously valuable text output, that is the culmination of all of that research. But the fact is that all this time, we've been for all of those intermediate steps, we've been paying for the whole essay every time, sending all the context in just to get one number back in that step.

So anyway. And it could be hallucinated. That number, when you get back, could have been wrong, which is a bigger problem with it, too. So it's like, we're using a tool. We were presented this tool three years ago. We said, oh my god, this tool can do a lot. And LLM is just amazing, for 1,000 different reasons. But what we're finding now, and this is, Jeff is just one of many examples of this, where models have been altered or a language model has been altered. This one's a little different, obviously, because it's not a language model. But it's using a screwdriver to hit a nail. I mean, you're going to make progress, I suppose, but it's not a hammer. It's not exactly the right tool for it. But you may get along with it, but there's going to be a whole lot of inefficiency that goes within. And where I start to get really excited with this, too, is you don't have like you go online, and you have like a financial calculator. Or you go to some website, and it's like, can I afford this home? Can I buy this home, whatever? And it's like, well, listen. Where's your credit score?

There's a couple of drop downs in there. It's pretty basic. All of that is prebuilt in the back. You know, when you go to that website, it's like, well, if they say 750 to 780 credit score, put them in this bucket of an APR, right? If they're here, here, put them in this bucket. And that's all it can do. Those are great for what they need. But what I'm excited about is now I know that I can go in, never mind the stuff with the show that I showed last week. But we do a lot for our clients, as well as internally as scaled. We do a lot when it comes to scoring, meaning that we're scoring currently, retroactively. Where is this project? Where is this project in terms of a health score? We've defined what that health score is, but we've been leaning on LLMs to sort of break that all out. And I'm thinking, no, no, no, no, no. That can be done by Jeff, right? That would actually be better done by Jeff. We have the criteria. It would actually be a better, more accurate health score, more consistently. But then I want to go one step farther and go, well,

now, because it's so fast and so incredibly cheap, I can also say, well, what if we took, I want to consult it on my team to go, well, what if I'm here, I'm here with this client right now, what if I took these next three steps? What if I choose my journey, choose my own adventure, was I did this this week, I delivered this next week, but I still made the milestone and the completion date of December 15th, right? Whatever that particular thing is. And they go, and then Jeff goes, well, based on the criteria, boom, this is where we think the health score would go for this particular client. And you go, oh, okay, what if I choose my own adventure and I did it slightly this way? And I moved up some of this data. What if I worked 30 hours this week on this client, instead of the next three weeks to 10 hours? What would that look like? And you can imagine, I don't have this built, but you can imagine you could do the criteria for this to get it so that it's not only giving you a health score of what has transpired retroactively,

but proactively and programmatically, what you can do going forward with a client to keep them on pace and you leave it up to the human in the loop to say, give them my judgment and my understanding, and we have really smart people who are consultants for us who deeply understand different parts of Revoptix and all sorts of different things with sales, and let them be the human, to be the expert and go, well, what if I pick this path? Ooh, it looks like the score adjust. What if I pick this path? Oh, okay, now I'm gonna make a decision based on this and move forward, but I have some data behind it that says, this is probably the way this will go based on my criteria, whatever was built into that. It gets really exciting and it becomes something that is really cheap for businesses to run. I mean, really cheap. Yeah, I was at a minute stage where I, sorry, Beth is on mute again. Go ahead Gary, I'm not laughing. I didn't want you to think I was laughing at you. I just saw that start talking and it, that's it, okay. So I was at, and my work, one of my work builds,

I have to switch out an older model for a newer model, I'm using Vertex, Gemini, a Musin Vertex, and I'm using Gemini 2.5 Flash. They're retiring that so I had to replace it with 3.5 or higher. And so it was like, oh, this would be a good chance for me to kind of check out Jev to see if this would be a good replacement where I could slip Jev in, instead of using Gemini. And so I did like 60 test cases, and I close up kind of like final results in this track, but basically it was interesting to see the numbers. Jev puts out more tokens, has more token, oh, reported more token inputs. It has more token inputs, but it was still cheaper, even with the token inputs being much like almost double, oh yeah, more than double, then what Gemini 2.5 is. But overall, Jev was almost seven and a half times faster

than Gemini 2.5 Flash and 95% cheaper, which is crazy, especially, I mean, this was only 60, just expected or the decisions basically had to make. And so if you scale that to larger amount, I mean, that's a lot of money that you're saving and the big scheme of things. And so it was the full response of what it took was like one way, 163.2 milliseconds, where it was once and you're not paying for output. Yeah, and if you think about it, it's flipped the model, it's flipped what we're used to on its head. So what we're used to is that we input something and then AI does a lot of process in building the thing that's gonna output to you.

That's not how this works. We input something that's actually made in a way that gives very definitive answers. It meets that criteria, the confidence score is 0.7. This is true, no, no, yes, right? Like those are the outputs. And so the actual output is so small that they don't charge you for it. And so it doesn't use more tokens. The other thing that I wanna say, Brian, you referenced hallucination. And that is also why hallucination is not possible for Jeff, because what Jeff is giving you is not based on is not being communicated to you in words that are made of tokens based on probability that lit up in terms of proximity in the wild, wild thing that is the, what happens when AI identifies probabilities in terms of next word or next token.

So this is a completely different process, which means no hallucinations. And again, the reason for that is because a hallucination is only called a hallucination when it's wrong. It's doing the exact same thing as not hallucinating when we don't notice it because it seemed right. And in my, one of my tests I was doing, I actually did this one on the Jeff site, not separately using the API. And I gave it some test scenarios where there was conflicting information and its job was to call that out and not decide either one and it did that. And so when I'm part of this system, I'm building your comparing people's policies versus their audits. And sometimes they can be conflicting. It says, yes, we have this, but do they actually have that? No, they don't. So that's a conflict. And Jeff was able to accurately identify those conflicts

100% of the time, which was good because I have a challenge with that in using LLMs because they just want to make it, they just want to decide. And so there's just that slight, just a slight little thing that can make them trigger one thing or another. So I definitely, I'm going to use it, I'm going to implement it. But I'm also like, it's so early to adopt something at such scale. And so I'm like, how do I do this delicately? Delet? That's not worth it. That's what I'm talking about. That's also brought up the point too. They're terms of use of data. I looked at that. Isn't great. So they can't train on your stuff. They can just use their, use your, your, whatever you put in, right, to better, like better their model, or better their system.

Because I did a deep dive into this. So they can't really steal it, but they're kind of do, but they're not going to go and take your ideas and run with it. They're just, I had some, I had it give me some, I had to give me some examples of like what it could do. Like give me a life, real life scenario of like what, what it would use based on the terms conditions. I had this conversation too and, and cloud attorney at, not law, said, actually what you're, what you're responding to is pretty boilerplate for an input into an AI that is actually going to do something with it and give it back. But damn, like when I was like perpetual all time, we're not going to train on your data, but we are, we reserved the right to do almost everything else. I have a, I have a question for you guys. Yeah. And, and that is given this notion that there is a,

you know, a process by which a combination of LLMs and Jev work steps are pieced together in order to arrive at a workflow. And we're going to take advantage of the speed and inexpensive nature of Jev to reduce the overall cost of that workflow process. What do you use as the method to stitch those things together? Where's the orchestrator and what, you know, like we think, in my mind, I think of N8N as being able to do this kind of thing. Okay, when you finish this step, Jev, then it's going to go, N8N is going to send it over to this one, which is a different thing. It's a NLLM. It's going to do this to it. And then it's going to go to another step, which is a Jev. But I don't think we're using N8N anymore in that way. What, what do you think of using, particularly Brian, because I know you're starting to think about using Jev for your clients and inside scale? What are you thinking is going to coordinate and orchestrate all of them?

I mean, the command centers I built, coordinate and orchestrate all of it on their own. They all have all the skill sets behind them. So yeah, it is a, if this happens, then that type scenario where it triggers the next part of the process, same with what I built for Amanda. So there's, you know, multi-multi-step processes happening behind it. And I would be plugging Jev in as either a controller or a stopgap if it didn't meet certain criteria, obviously. But what I'm really doing is thinking about how I use Jev to check some of the outputs coming out of the LLM in some areas. Again, I think you can go to transcript. Ben's bite actually calls out and says that he, on top of the one I just showed, skip sponsor segments on YouTube, gmail by intent, or filter to YouTube to do's, a new smarter copy, paste, drag and drop, and command, F experience. That's interesting. Controlling your Mac with voice.

So I don't, like I don't think I've expanded my mind yet, Andy, to get to where all it can be done, but to answer your question, it's, I'm not doing NAD because now like, in those command centers, it's sort of like it's own baked in NAD. It's already running through those processes step by step. And it's a, it's a reasoning model that's doing that orchestration for you. Correct. Okay. Yeah, that's how I'm using the algorithm. It's not the reasoning model, right? What's that? Give it to Claude or Astra or like you're giving it to another AI to build that criteria. And that's right. In my case, I'm using Anthropic, but yes, correct. I found some very reliable guides on Jev, or this last week, I'll post them in the Slack. I saved them because, and they were from reliable sources on X. And by Slack, you mean our Slack communities are you referring to? Yes. Yes. Reason to go to the daily eye show community. Actually, I don't think I fixed the link. Tomorrow, go to the daily eye show community

and join it. Slack, Slack unfortunately kills that link every 30 days. It's super annoying. And honestly, I set reminders to do it, but you have to literally go in. I have to go into our, the main service and change them. It's annoying. So, I've been able to daily AI show.com. You still won't bring up the Slack, the right Slack community page for joining the Slack community on it. So, unfortunately. It goes, it forwards to an invite link and the invite links expire. Yeah. And they, they roll it literally every 30 days. I don't know why Slack does that. It might have something to do with paid or not paid or otherwise regardless. It does. All right. I'm going to move on to it because we already started the show 15 minutes later. So I'm going to move on here because I do, I know we have other news to get to from everybody else. And also, I do want to show something from the project with Amanda because I think if people really like that as well. So I'm going to bring this one up really quick. I found this slightly humorous.

US Federal Register caught using Chinese AI model for documents, sir. I don't know if it's the right word. It sounds like it did something nefarious, but I'm going to scroll down here to where somebody got a screenshot. It's no longer live. But basically, there was a, where was this? It was a bit of a, it was a Switzerland-based commodity portfolio manager that squared this, squared shared this screenshot of a document search interface on the federal register, the official journal on the US Federal Government. Far down on the site's interface is an option to choose a search mode which default to semantic. But it also contains two options for Quinn, three.0, six billion, a large language model developed by the Chinese firm, Alibaba Cloud. So you can actually see it in the screenshot. It's not there anymore. It was quickly taken off the page. But as they say down here, and this is from futurism, by the way, I like how they wrap this up, futurism.com. Their inclusion does not imply that a Chinese model was

put in place maliciously, or it's hoovering up American secrets or just batching them back to Beijing. Rather, it shows that whoever was in charge of maintaining this search feature on the Federal registry site simply chose to run with a highly capable and low cost Chinese AI model, issuing bulkier home drone models that have major contracts in place and various government agencies. So it still elapsed obviously for that to happen. And it's sort of embarrassing for it to be on one of the major US government pages. But I just found it, a note to share on it. It's gone now. But low and behold, people inside the US government are also like, hey, you know what can do this cheaper? Quaint. Oh, just good results out of it. And then somebody was like, hey, don't show that publicly. Don't put that out. Don't put that out to everybody. Use it internally, but don't tell anybody. Don't tell anybody we're doing it. So anyway, wouldn't the share that. But let me throw it to you guys as well for other news as well. Something for Xi, Xi, Xi, Xi, and Trump to discuss.

You're using Quint. So my guy. My guy. OK, speaking of discussions around China and international cooperation in AI, right now the UN General Assembly is underway. And today Donald Trump is going to deliver some high IQ remarks to the General Assembly. So we're looking forward to what he has to say about AI. And the brilliance will shine through. But yesterday there was a separate meeting at the UN General Assembly called the UN Digital Cooperation Day. And it was hosted by the UN's Office for Digital and Emerging Technologies. And they, the adherence to a cooperative approach to the development of AI, they issued a call for control of frontier AI models.

And it's a very short one-page thing that was signed by presidents and premieres from dozens and dozens of countries. The US was not a signatory to it. But I don't think China was either. These are all largely, I think, the other countries out there that are saying, hey, we need to do something about this. And this is our plan for it. I'm not going to read the whole thing. It's worth reading, though. But if you look for the, just search a call for control of frontier AI models, that came out of the UN yesterday. And where I saw this is in Gary Marcus' newsletter, because Gary Marcus, a major proponent of skepticism about all the things that are coming out of AI hype and a very well-heeled and very smart commentator

on these things, coming from the technology development of AI, he spoke to that UN digital cooperation day assembly. And so he has his own remarks in his newsletter about what he said to them. And then immediately thereafter was the release of this call for control of frontier AI models by all of the presidents and premieres of multiple countries. And he did not know about that when he spoke to them. But there was a large degree of alignment between what he was saying and what they said. So it's a good direction. It's worth reading. And so I commend it to you. Definitely. And we'll get that in the community as well, like we said. I'll get that link, taking care of. Beth or Gary, with any news you guys want to share from today, they're not, I'll be happy to share some demo stuff, but I want to make sure we get through the news first. Just something real quick. Grock Point 4.7 got released yesterday.

And it is supposed to be specifically trained for Grock bot. And so I am out of my usage in Grock bot. So I'm grounded for two more days. I can't test it. So I'll tell you in a couple of days, once I get used to it. But it sounds promising. Benchmarks that they put up seems pretty good, much better than what Grock 4.6 was. I'll see what you want. And it outscored everything in the world on electrical engineering, which is a separate thing from computer engineering. Electrical engineering is, you know, like, think more in terms of the chip design and manufacture. And those aspects of electronic computing. And so yeah, Grock 4.7 really aced that benchmark way beyond all the other frontier models. Which is so weird and interesting.

Like why? Why, what did they train on for it to go like to jump like that in that specific area? That had to be a reason for it. I just saw that too, Andy. And I thought that's an interesting field to just get drastically better at, you know? Well, it could be that XAI is actually in pursuit of design of their own chips for it. I mean, yeah. And so they wanted their latest model to be expert at that kind of engineering. Right. I mean, go ahead. Well, and Elon came out and sort of laid out the next rollouts before this one happened. It was like 4.7. It's going to come. Not going to be as good as the frontier models. The next thing is whatever the next thing is. And then they move up a number. So it's going to be 5. And 5 is going to be bad ass, right? Yeah. I'll be interesting because I have this grand scheme idea of building this drone that paints murals for me. I've always like scoped it out like with different models.

That's a good idea. And so like, yeah, I've scoped it out and that one of the websites that we mentioned a couple of weeks ago, I went and did it through there and saw what it could do and still above my pay grade of how to build things. But I'll have to check it out. Also one thing we're expecting sold today by all sources. Oh, right. And maybe Luna. Maybe Tara's going away. Yeah. So somebody at OpenAI. Maybe Luna's going away. Sorry. I don't know. One of them is going away. Words, maybe expecting to. Angel, I think it's her name. Angel, she posted a thing about sold today. I forget she kind of made fun of it. Or I'm so excited. It's what's true. Which was on a comment for the testing catalog account saying, hey, maybe we're going to get sold today. And Angel works at OpenAI.

And Tiibo wrote, ladies and gentlemen, start your engines. We are almost Tuesday. And I promised a reset for Tuesday. So we got a reset last night. Among some other things, see you soon. OK. But the promise, the conversation that happened back and forth, the commenter said, you owe us a banked reset because you promised stuff that didn't happen. And he was like, OK, but it's coming on Tuesday. Yeah, no, this was this. What was given was not a banked reset, right? Yeah, I don't know. No, it was just a not a banked reset, but you're saying it's a reset. I don't I'm looking at my. You have an extra banked. I don't have either. That's usually my escape. I was like, oh, great. This is coming. It was like last week of the week before. And I was like, and it's not there. By the way, I'm excited about any sort of updates to soul. I've said what I've said about Astra.

That's where I'm at with it right now. It is what it is. I know lots of people love it. And I get it. I just haven't had the same experience. But I have had a great overall experience with Soul 5.6. I mean, absolutely. It's a better chat should be tea experience. And I love how seamlessly it'll move from a chat into work. I think that's like super smooth. And there's no loss from the handholding from the user standpoint, which I really like that happened twice yesterday as I was doing some stuff. I had to do some really deep dive stuff for me yesterday. And it really hit the ball out of the park. And that was all soul. That was an Astra at all. And so yeah, I'm excited if there's any sort of updates, whatever they might be, because just generally when I'm using chat should be tea and then when it bounces into work and does stuff, it is really, really good. I love its integration with image 2.5 and how it's able to quickly bring images into it as well. So yeah, I hope it's a great update today, whatever it is.

Now, my day's ruined. I didn't get my use that. Sorry. I'm going to start getting through. I promised something today. But I'm super excited about what Brian's going to show. So come back tomorrow, Maro, for real. And I'll share those. OK, before we go on to Brian, I have one last one. The next item that I think's important. And that is that Meta's Muse agent is getting further and further kudos from people who really admire what Meta has done with that. It's like the choice among agents out there for many people. Today, I read a commendation by Ben Thompson, who writes Stratitory, which is a business and technology analysis place. And if you really want to read somebody who's incredibly articulate and erudite, that's Ben Thompson. And he mentioned specifically in an article that he wrote yesterday about Meta's Muse

and that he's using that personally now, which is interesting, because my inclination is not to go to Muse. But Amazon has similar problems with Muse. And they blocked Muse agents from accessing shopping on Amazon, which is a very important step here that says, OK, the front door of commerce sites can be a blocker for the use of your personal agents. And there's an open question about whether that's fair. So what is it we do? Elon actually wrote a response to this instead of it's coming from your IP and using your cookies. It's you, right? Like they can't block that. Good. But this is where Meta is like, well, we support small businesses. Your our Muse agent will only buy from small businesses and will not support Amazon. Well, that's they did have a huge Shopify

jumped all over this. The CEO jumped a Spotify rounder to Spotify. Jumped all over this and was like, yeah, we have a deal with Muse. Just like why wouldn't then and talk about supporting small businesses lots of companies I used to use Shopify for a business. We're about to use Shopify. Yeah, I mean, they're they're great. So more or at least it's an easy, it's an easy interface to build. There's headless headless Shopify. I did research last week on it because some over over other one choices and it looks really cool, looks really promising. I only thought Shopify was like a hosting website. So I used to use it for my art. And so they have an API that is like headless. And you can connect it to the back of your systems which is interesting. I didn't realize that. All right, right? Right. Yeah, show your stuff. I was, yeah, I just I didn't want to, if anybody's a Salthana, but I saw the same thing, Andy, and I thought that was interesting. And like Amazon's approach, like you said on that was, I read somewhere that said, well, the agents didn't announce themselves at the front door,

which is a funny visual to have, you know, of agents like knocking. I knocked three times. And then I just let myself in, you know, and there's anyway. So that'll have to be all figured out. It's going to be, it's going to be a big, I mean, I know you're just saying it's one new story, but I think this is a good, be a new story that we hear a lot about because Amazon, you know, same thing as Salesforce or whatever, as Salesforce really just a database that we go through headless and we don't ever go into Salesforce anymore. And if that's the case, why are you paying Salesforce prices, right? There's a lot going on in this space. And I think this is another one where Amazon is saying, hey, it's only really great to us if you're on our system, shopping on our system. And maybe we don't want anybody else's agents that doesn't mean Amazon will have their own, or their own muse, they kind of do. And, you know, we're going to block others from doing it. So I'll be really interested to see where this goes. And the last thing I'll say is, I've been using Muse like once a day, honestly, like that's about where my use case is. It's fun for me last night.

I had some fun with it. My Instagram account, which we don't do much with, it's sort of just a carbon copy, right? But like way way back, it started, I guess we started in 2017, so says Muse. And I was like, oh, just for fun. Like tell me your top 10 funniest things that I've ever posted about that Sophia, my daughter, said, you know, going back quite a bit, this 2017 she would have been seven. So really, from anything seven on or whatever. And it was just fun. It was like, there was nothing bigger, grand scheme of building or anything. It's just Muse made that really easy for me to have this fun moment where I was just laughing and remembering these things that Sophia had said as a little kid or just thought little music, whatever little musings, no pun intended, alone the way or whatever. And so that was super fun. And I also have it basically shopping for a car for me, which I thought, you know, look, I could do that other places too, but I use Muse makes it very nice. It's a nice user experience to do that. I've been, I've been saying on the show, we're giving my daughter our, our EV, our bolt.

And so what do you replace the bolt with? Well, it'll be a used car. And I think I have it narrowed down. And so I basically have it once a day just going out and checking to see if any of these particular vehicles have come onto the scene anywhere locally within like a 50 mile radius of me. And how do they compare and contrast to the other ones we've already seen? And every day still gives me the top pick and says, nope, hasn't moved. This is still a top pick. I think it's priced went down $150. And it goes as far as to say, like, do you want me to pursue this for you and me to reach out? They also had that same feature with Facebook Marketplace. Muse will go and shop for whatever the thing is you want under the price point that you want it. Go in contact, negotiate, and then come back to you and say, okay, this is what I found. That's pretty powerful. I mean, I go on, I don't know, on Facebook Marketplace every day, but I've certainly built and sold things on there because it's easier. And yeah, sure, if I'm looking for a particular thing, a new guitar, I don't know something, you know.

Yeah, why not? Why not have it go do the work for me and come back to me and go like, hey, I negotiated this person down, they're willing to take $50 less and they're only about 20 miles from your house. Do you want me to tell them it's go? Sure, thanks, Muse. That's helpful, you know? So anyway, I just wanted to share my experience with it. Improve your Facebook Marketplace rating. Yeah. Rate it on your transactions. Yeah. Okay, so I'll show this really quick. So let me give a quick, yeah, I know not everybody, most people are actually not here every day. So let me just tell you really quick. So I talked about last week or whatever, but I've been working on this for months. Wife had decided that, my wife had been and decided that she wanted to be an independent travel advisor. I think there's a difference between travel agent because she's not a travel agency. She's working for an agency. I don't exactly have all that verbiage correct, but anyway, she's a traveler, right? Not a full-time job, not looking to do that full-time, something she wants to build. And we sort of devised a plan and said, well, look, if you want to start this, like, where could it be in three years when my daughter is college, right?

Getting out of high school. Okay. And so I said, well, look, I really would love the opportunity to build this as an AI first company. If you're cool with it, my wife said, yep, great, go nuts, right? So over a call last couple of months, I've been really pinging and figuring out what I could do. And a week or so ago, I shared how what my point was to it was how powerful I think HTML pages were. Well, at that time, I didn't have anything visual in terms of the command center to share with you all. But I thought today, I would show you where I am. And it is absolutely a work in progress, but I don't mind sharing as I'm building. So this is actually nice because I do very similar. They look very different, but similar and scope in size with scaled, but it's client work. And so I can never bring it on screen. I can just talk about it, which is not as much fun. So I will share some of this really quick. And so this is the start. I'm sure it'll look different in time, but this is the start of the command center for a Java journeys. And so what this does is I will in time do a lot of things.

But I just want to show you a couple of things. When you're on the today page, basically, that's almost like her chief of staff, her command center. And if I come up here, it's like four things that need you today. So the goal here is that Amanda could be on her phone. And checking this, she can come to a desktop and check this if she needs to like a lunch break at work or she could just do it when she comes home at night. And it's like, where do things stand here? Like what do we need your approval on? What are we working on? So like what I can do here, as you can see here, this is a video that we is actually going out today at like now. It just went out live at 11. And so this is the one that we've sort of been testing with this Utopia the Seas video. And so I can actually click into that, but I'll actually go through it this way. That would take me directly to it. Here are the two videos that we've sort of started working for. And you can see there's a nice color code that shows how far along this full process Amanda is in developing this video. So then if we click into the video, you can see on the left that it's got all the stages and steps that need to happen from what was it done? The review checks, the evidence, did she approve the outline?

It got rechecked again. Now we're moving into titles and thumbnails. Then there's a full packet, which I kind of explained last week. Now then it gets down to the recording package and the record. So Amanda, actually, I know there's a server error. That's not great. Can I go back? Please let me go back. OK, that made me nervous. Let me see if it'll at least get to the product when the show. So then what it did for thumbnails is it actually not only did it record and do all the steps. That's why I wanted to show you the record. But if I was on the record one, what it did was it took the entire transcript. It broke it down into steps. Instead, OK, these are the ones where you're on screen, Amanda. These are the ones where you just have to read something off screen. I'll show you an example of that in a second. So it gave Amanda it took a 11-minute video and turned it into 33 clips. Amanda just sat at the computer directly in the command center, not anywhere else, and she recorded her 33 clips. This system brought everything together.

All 33 clips together. It used a new tool that I'm using through an API to enhance the audio. And I'll put it in the video. I'll show you what that video looks like in a second. It also created all the thumbnails. What process? Exactly the process I just built for the daily AI show. And we've been using now for five or six days, about a week now, with about seven shows that I have now building thumbnails. Well, you also need thumbnails nowadays for the shorts. So it built all four shorts that came out of the original video and did all the thumbnails for those as well. So Amanda looked at these. And Amanda's first response yesterday was, my eyes don't get that big. And if you know my wife, you know that she has, we call her smiley eyes. And when Amanda smiles, her eyes disappear completely. Sophia has the same things. One of my favorite things about Amanda is, she smiles. And you can't see eyeballs anymore. And so she's like, my eyes are way too big. I was like, we can fix that on the next time, babe. But like, you know, so funny enough in the office last night, there's me and Sophia saying to Amanda, with this picture big on the screen, we're like,

no, you're not doing it like the girl in the picture. And we had Amanda trying to recreate it. And we're like, eyes bigger, eyes bigger. And Sophia's standing right next to me going, no, mom. You have bigger, bigger eyes. And Amanda's like, my eyes don't get that big. So anyway, it was a funny family moment yesterday on that. So then you can see publish in live. And that's not exactly live on this site yet. But I just want to show you some of the other things on here. You have ideas. In time, right now we're pushing these through manually. But every day this system will actually be going out as agents and figuring out new ideas for new videos for Amanda. She can have her own ideas, but she doesn't need to. So she just happened to do a tour of the Margarita Vol. Islanders, so we actually did this manually. You can see it says, soon here, because I like to manually and build out everything and wireframe it. But if it doesn't exist yet, I just want to know where its placeholder is. It will create all the socials for her. It will create all her newsletter sections for her. And then if you get down into people, this is the CRM. So I built out a full CRM for her as well.

So now you can see here, this is where Brian TestMe have come in through the system. And I've come in through various different ways. I downloaded a magnet, whatever it is. If I click on that and it says Brian Test Cruise Planner, it's how I came in. Here's the part that Amanda needs to do next. Hey, you need to respond to him. Here's where she was able to put notes down about it. And it's got all your general basic starter CRM stuff in it. And it'll probably get, there'll be a lot more fields in here. It'll get bigger and better as it goes. But the idea here is that it's tracking everything and everyone through the system where they need to be. When they get into trips and client journeys, this is where it's going to be doing three and four and five part email sequences. If Andy's taking a cruise and that goes out on February 27th, there's five or six things Amanda probably wants to tell Andy about in preparation for a cruise. And it's setting all that up in there. Same thing with client journeys. Then you can see all the library of all the clips. These are all the clips that Amanda did. See, there's the 33 clips as well as stuff that came from the individual cruise lines.

Here's all the activity. This is sort of like my log of everything that's happening in the system. And then I'll jump over connection so I don't share too much there. But I also built out a cost screen that's just letting me know on any given day, what are my total costs? Because Amanda said, hey, I need a P&L for this. I need to know what this is costing. And I said, I don't know yet, but we can track it. If we can track it, then we can have an opinion about it. If we have opinion about it, we can set a budget. So right now I have it at $75 a month. And you can see so far with all my iterations and all the things that I don't, right now, we're at $1.55 for the month, for all the images, all the videos. Most of this is being done off of LLM. So like FFM Peg is being used for editing all the videos. Been using that for a year now. It's a great tool as well as the audio. And then if you just go here, then you can actually see this is the video that just went live. And you can just edit all the cameras, but you can also see here. It's enough for, where it did all the, is it good for you?

Which show to reserve first? I was trying to get to like Royal Caribbean has on sale for Utopia. So all these graphics were all done by the system. So for the Navajo. Everything. The, the C3 glasses, a popular in the iPhone right now. So you book this, all of that. And you can kind of see where the clips would have been. It also, down here at the end, you had the three night sailing. Figured out how to use the night. Sunday, November, did a demo. I'm going to close it. Amanda what she should say about it. And when she comes back on camera here, it also created a workout and made it work. We've used to figure out everyone else hearing Amanda at the same time as Brian. Oh, sorry if it's her and hit me at the same time, isn't it? It's a hard to hear me or her. Yeah, you're at the same level. Oh, sorry. That's annoying. Thanks, but I'm trying to hear Amanda Brian. And most people prefer Amanda talking over me. That's not it. That's not the find you site. So you can just see here, my point is just to show you guys, like all this was created through that process. So like Andy,

when you talk about earlier about like, you know, the whole system running together, that's what I'm building. Like that's the goal. Like I've been building 99% of this in clogged code, right? And just iterating. And a lot of this was generated in the clogged code desktop browser. And working. Care, let's just let your comments. Sorry, same thing you said on here. Brian stopped talking, I'm Andy here Amanda. But my point is like, so once I got that sort of structured, now it was time to start building this towards a actual command center. The goal is for Amanda to be able to like, this is a single point of source truth, everything. It knows her emails, it knows her contacts, you know, it knows how everybody came through the funnel and it's moving everything through for her in keeping track of that. So when she comes in in the day, she's got sort of that chief of staff that says, Amanda, these are the top four or five, six things you need to be worried about right now. You haven't responded to, you know, to Gareth. Beth is waiting, she's post-cruise,

but she hasn't done her review. Andy needs this or that. Andy, Andy doesn't want to do the cruise at all. Maybe you should follow up later on that. Whatever the case may be. So I just want to share that so that people can kind of see where my brain is is I'm sort of building this ecosystem out. It's going to take some time. But the goal is a full command AI-forward business that's built from the ground up so that every decision, as I said, has been being made. Where does AI fit in this? What can it do? How can we replace other human steps to take up time? Where does Amanda need to be spending her time? And we build the business around those concepts. And honestly, I think this is something literally anybody can do. I really don't think this is above anybody if you just put your mind towards it. And you said, I want to build something like this. You have the same ability. I think I do to do this. I hope that kind of maybe inspires people to go build bigger things. Well, I agree that it could be done. I think you are an exception and Amanda

and you together are an exception that have created something that you understand clearly that most people would have a difficult time fascinating into being. So you're setting a great example and I would not make it seem like, oh, this is what anyone can do at this point. Well, it's true that you can do it. But there's a lot of expertise that's been developed by you and by Amanda in this particular domain that is showing up very clearly in the work that you guys are producing. So Andy is wanting Brian to understand. As always, we all want Brian to understand how incredible he is when he does these things. They just toss it off and worked a little bit. Look, and we're all like, damn, that's very complete and it's just getting better. But Brian's point is also to be taken, right? Brian is working with AI within what he wants

and other people can absolutely do this, right? It may take you slightly longer. You may have to have more conversations of what people want. Sure, maybe. But Brian's point is also true. Andy's is true, but you can do this. So I do have a question for Amanda though. What is the cruise operator and or the cruise ship that has the lowest incidence of viral plagues and X that go through the staffs? I can't hear you, Andy. Break it up. You're breaking up. I have posted the link to this video in the Slack or which you can't go to until tomorrow. But definitely when you see it, go and give Amanda some love. Yeah, that would be great. I appreciate that. If you can go to YouTube and don't just follow Joe Vigernes if you don't want to. But if you want to give this particular video that I will tell you this, it's not all smooth sailing.

The goal is that Amanda can do three of these a week. Just think about it. You saw everything that it was. But if you see if the system ramps up, it's plausible. 33 clips, Sunday, not a straight seat, over four hours for Amanda to do those 33 clips. Not because it was hard, but because like it was just, well, the transcript says this and I have to think about it. And it's not in my voice and like all these things that need to be wrinkle, you know, iron and wriggling out over time. But it was still very laborious. That doesn't mean that the AI is failing. It just means we're learning as we're going and we'll take improvement process along the way. But it has been very laborious. Like we've spent 10x more time on this one video. But of course, the power is 2, 3, 4, and 10 will all go through much, much smoother and faster. So that's what we'll pick up time and the efficiency on it. Hopefully, you know, over time. And it's everything is a CPI, right? Everything is continuous process improvement. So there's not a button that I built on the command center

that doesn't have a learning behind it. Because what good is a button for Amanda saying, I don't want to do this idea or I disapprove of this transcript. If there's not a learning to it, what good is this video going out that gets views on it, that gets comments on it. If we're not throwing that back into the feedback loop that says, how do you get better? So in time, the whole system is getting smarter as Amanda gives her feedback as there's analytics on the post of the socials, on all socials, that gets fed back in. So the system is dumb by comparison of where it will be in six months because it will have all that feedback data to say, oh, next time, I won't make this type of recommendation, that doesn't work for Amanda. So I really look forward to where this is in six months because I think it's going to be much smarter than where it is today. It's just going to literally take all the inputs and the manual work to get there. But it will get smarter over time. So yeah, it's a lot of fun to work on. And I'll keep sharing as I, not every day,

but as I feel like I have something worth viable, that's that I think you guys would think is cool. That I'll certainly bring it up again and share with you guys. There's a tweak that I heard from Ellie Miller on a webinar that she did for Gen X Women last week, Gen X Women in Tech. And she said that she actually makes the, makes cloud, makes the AI score it for her. How much do you think I'm going to give this? A one to a 10? What do you think I'm going to say? And that actually adds a little bit more in what we would call fourth thought. That's not how I AI works, but sure. It helps get to the place faster. But just like, I don't like this try again, do it better, make me proud. And you're going back to the beginning of the conversation. So we grab things up. Again, a use case is for Jeff, right? How would you score this? How do you think I would say this is going to do? Maybe throw that through Jeff and then use that criteria or that greeting of it to re-informed future ones.

Because if Ali is saying or Amanda was to say, it comes back and says, I think you're going to give this to seven. And Amanda goes, that is my score. Good job. Or I think I'm going to give this to seven. Amanda says, that's a four. And here's why. Back at the feedback loop. And it gets smarter. And it gets smarter. So that greeting hopefully gets closer to what Amanda would actually say every single time. Now for that to work, Amanda needs a score first in her head so that she's not being influenced by the score he gives. But it can be done. And I think that would be a fun way to do that. And another thing that Jeff would possibly be really good at. OK, we're actually at the one hour mark, which is to say we're 15 minutes over the hour, but we start at 15 minutes late. So we will wrap that up for today. Thanks for hanging out with us. Thanks for having to jump over to another comment. We'll get this all sorted by tomorrow. Just lovely improvements from string yard that threw us off there a little bit. But it should be all good for the rest of the week. So thanks everybody for the comments. Thanks everybody for hanging out. Thanks to Beth, Garrett, and Andy.

And yeah, we'll be back tomorrow for another great show. So we'll see you guys then. Bye, everybody.

More episodes

More from The Daily AI Show

View all episodes →