
How to Be "Agent Native" in 2026 w/ Every CEO Dan Shipper
About this episode
In this episode of The Neuron Podcast, Corey Noles and Grant Harvey sit down with Dan Shipper, CEO of Every, to talk about agent-native engineering—the framework his team uses to build and ship AI-powered products at a pace most companies can't match.
Dan walks us through what happened when his AI document editor Proof went viral (and then went down), why he believes the way we build software is fundamentally changing, and how Every's small team manages to ship and maintain an entire suite of AI tools: Spiral (automatic style guides from your writing), Sparkle (AI writing cleanup with custom folders), Cora (AI research assistant, now on iOS), Monologue (AI-powered journaling with notes), and Proof (the agent-first document editor that broke the internet for a day), as well as their new to be revealed on Friday: Plus One (a hosted AI agent for Slack).
Whether you're a founder, developer, or just someone trying to understand what "agentic" actually means in practice—this conversation is the real-world playbook.
Subscribe to The Neuron newsletter: https://theneuron.ai
Products mentioned:
• Every: https://every.to
• Spiral: https://spiral.computer
• Sparkle: https://sparkle.computer
• Cora: https://cora.computer
• Monologue: https://www.monologue.to/
• Proof: https://proofeditor.ai
• Plus One (the new one!): https://every.to/plus-one
Get every episode summarized
Each time The Neuron: AI Explained publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
1,795 searchable segments. Every word is indexed and playable.
Full transcript
The Neuron: AI Explained — How to Be "Agent Native" in 2026 w/ Every CEO Dan Shipper. Machine-transcribed; use the interactive transcript above to jump the player to any line.
Hello, we're live. Whoa, there was not a countdown. That's bizarre. There was, there was a countdown, but you've missed it. I never saw that. It's 30 seconds, isn't it? Oh, well, hello everyone and welcome. Welcome humans. Thank you for joining us. We started a minute early here. We're, we're just want to kick things off now that people are joining. Yeah, yeah, just getting, getting set up here, going to be a good one. We're going to have Dan Shipper from every in a few minutes and Grant tells us a little bit more about who he is in just a moment. Coming up before we get started, just a quick reminder that our Nvidia contest is still active through this Sunday. So if you attended a session last week and took a screenshot, go throw it in there for a chance to win a free DGX Spark that we'll be giving away next week. That is an extremely good computer, y'all. Like you can win an extremely good computer for like $4,000, the $4,000 computer at least. And have the envy of us both because neither one of us has one.
And it is sad. I would love to have one. What are these days? But yeah, make sure, how do they sign up for that? What's the link? Do you know? I assume I'm gonna drop that in the chat for us here momentarily. Okay. Okay, and I just saw. Looks like our guest is here. Grant, would you like to take just a moment and introduce Dan? Yeah, definitely. So today we're doing something special. We're live with Dan Shipper, CEO of Every. If you're not familiar, Every is a 15% media company and maybe more. Maybe less, I don't know these days, Dan. You can tell us that publishes a daily AI newsletter, ships multiple AI-powered products, and runs a consulting arm and their engineers write virtually zero code by hand. You guys are awesome. Dan, welcome to the show. Really excited to have you. Thanks for having me. I'm psyched to be here. Excellent. Yeah, it's great to meet you, Dan. I really appreciate you taking the time out of what I'm sure is a busy schedule to come join us. Of course, anytime. Appreciate it, appreciate it.
Dan, so I just want to start. We kind of had the hook for this episode as your saga with Proof. We actually promoted Proof in the newsletter as well. So I don't know if we contributed to your headaches there, but I appreciate it. It's good problems. Yes, exactly. But before we get into that, because I do want to talk all about that, for people who haven't come across every yet, what is it you're building over there? Give us your pitch. Every is the only subscription you need to stay at the edge of AI. We have three parts of the business. One is a daily newsletter about AI. The other is a app studio. We build AI apps that we use to work and live better with AI. We build them for ourselves, and we release them to our audience. And then we have a training part of the business where we do live streams for subscribers and courses and all that kind of stuff. So you pay one price, you get access to ideas, apps and training, all bundled together. The apps that we make are things like Cora,
which is an AI agent for your email, Sparkle, which is, organizes your files with AI, monologue, which is a speech-to-text app, sort of like Whisperflow or Super Whisper Spiral, which is an AI ghostwriter, now Proof, which is an agent-native document editor that I just built, and we just launched something new today called Plus One. Oh, you launched it today? We launched it today. So this is my first time talking about it. Awesome. You can check it out, every.to slash plus one. Plus ones are one-click hosted open clause that connect to your Slack, that connect to all of the every apps in our ecosystem, so they natively are connected to your Cora, to Cora, so you can do your email, to Spiral, so they can write for you and your voice to proof so they can write documents, all that kind of stuff, and then we have a bunch of skills and workflows that we've built into them in our working with open clause. So we started using it open-call all the time internally and basically found that it totally changed
all of our workflows, and also working as a team is really interesting together when you have many many clause altogether in a Slack. And so we just took all the stuff we learned from that and then turned it into a hosted service where you just click a button and you get everything that we think is good. That's awesome. That's awesome. That's very cool. Thank you. So much to unpack there. We'll definitely talk more about plus one, but the thing that jumped out to me was the open-call, like coordinating a team full of open-claws. Yeah, how do you do that? It's a really good question. So the like big unlock for me in using open-call and realizing that it was a serious thing that's different from other types of agents. You see, anthropic is shipping super fast with clause and people are being like, oh, they're like adding all the same features, which is true. And I think that they're obviously being really smart about how they're building that product.
But there's something very different about having a clause in your Slack that everyone uses versus having a clause with no D in your Slack that's yours. And the difference is when it's yours, so my clause named R2C2, my plus one, it's really a plus one, plus one, is named R2C2. But anything I say about a plus one applies to clause. And R2C2 is a little bit of an extension of me, right? Because I'm using him all the time. And he's modifying himself in response to what I want and what I like and what I need. So he's writing his own code to better serve me. And in that way, it becomes a little bit of a mirror of me. So R2C2 knows all about proof. This app that I've bived code a couple of weeks ago, he handles all the bugs for proof. He's got opinions on YouTube headlines, but he also does my book notes.
So he's really into quantum physics. And so what's really interesting about using them in a team is people see me use R2 publicly, because it's in a Slack, so they can see me use him for certain things like a bug comes in, and I start talking to him. And then they see that if I'm using him and they trust me, then they're going to trust him for the same kinds of things. So I sort of like transfer my trust and reputation to him, both by using him all the time. So we're pretty sure if he's modifying himself and responding to me, that he's good at the things I need him for. And then I can transfer him because I'm using him publicly. And so then people start using him for that. And we found that basically with everyone in the organization. And so what happened is, and we were about 25 people now. So we've grown a bit. 20,000 really cool. And what we found is it sort of like creates this parallel org chart where every single person in your org
just has their own plus one or their own claw that mirrors them. And then that sort of extends the work that they do. And that's really powerful. I can talk about this forever, but that's the basic idea. Isn't it crazy how much has changed since a dude dropped a tool? It's crazy. And it's like people were talking to me about this on Twitter today being like, oh, like OpenClaws, it's not a big deal. It just has a heartbeat. And it has gateways to all your messaging apps. And on the one hand, technically, that's true. And on the other hand, yeah, it actually does change everything. If you actually use it, there's a bunch of things in it that on their own don't really mean much, but altogether turn it into something totally different than what you mean. Is your personal one really buckled down from a security perspective? Or did you? I went pretty early. Yeah.
I didn't at first, but I've gotten increasingly yolo with time. I probably shouldn't say on a live stream, but I'm kind of like a yolo guy. But the way that we built plus ones is they have all the sort of default like good security practices. And essentially, you can only reach him through Slack. And you can only get in our Slack if you're a reasonably trustworthy person that we pay. Yeah. And he also doesn't listen to anyone except for me. So there are certain things you can do like that that make it a little bit safer. Give it its own accounts and some things as opposed to like, you know, have a buffer in there between like say your debit card. Yeah, he doesn't have access to credit card or debit card. So, but fair. People on my team do have their claws do that. And they just provision, you know, provision a mercury or ramp card. And it's, it works pretty well. It's really cool. Yeah. So as far as like using it as another coworker, do people like, because you know, you're the CEO of every do people ever come to your claw or your plus one
and ask for things like that they would have asked you for? Yes. And I think the biggest one is just proof like because I built this thing and their people are using it all the time internally. Anytime they have a question about it, how does this thing work? Or a bug usually it's a bug. They just, instead of tagging me, which is like, at a certain point, I was like tired of being tagged into bugs. Even though it was totally, it was totally my fault, but it's easier for them to tag my claw or my plus one. Yeah. And so yeah, it, and I can imagine right now it's not like a thing where if you asked him a strategy question like, what would Dan say about this? I haven't spent a lot of time doing that, but I think I could get that. I think I could get it to the point where he would be pretty good at that pretty quick. And I do feel, yeah, and I do feel, I do feel generally that, you know, like proof, I would not be able to do it even with like regular codex and clawed code
and whatever, I would not be able to do it at the level that I'm able to do it at and run the rest of the company and do all the other things that I do. If there wasn't this thing that's like hanging out on a server somewhere, or just like waiting to respond to requests and making me feel like, okay, that part of things is like more or less taken care of. Obviously, there's a lot I need to do myself, but he's at least the first line of defense, and that's really cool. Yeah. Yeah. Is your suspicion that this is how all companies will go, where we all have essentially like a digital like agent, a twin of us? I suspect so, but I'll say, I don't think there's any one size fits all thing. Different companies are going to be at and use this stuff in totally different ways that are fit for like how their organizations work. And like, you know, there are, there's like dry cleaner down the street from my house that like still doesn't take credit cards. So plenty of people will not have any of these for like a very long time. But I think there is, there's like a real debate around, especially internally,
but I think generally around, yeah, what does this look like when everyone has an agent? Are we all going to have one agent, or are we going to have agents that like specialize and if we are going to have them specialize, how, how do we get them to specialize? And what should they specialize in, you know, and I think there's what we have found so far is specialization is definitely a thing. Yeah. Even if you have this super smart, always on alien intelligence that could technically go across all different functions, somehow having, having it just focused on being a good, you know, marketer makes the whole thing better and it fits in our brains a little bit better. Yeah. Yeah. And the other really nice thing about the way that clause work is because your reputation is on the line for them because people see them as yours. It's a little bit like your kid, like you don't want your kid messing up because it reflects
on you. Yeah, totally. So people spend a lot of time making sure their clause are good. And I think that's an underappreciated benefit of this whole setup where it has a personality in a name and all that kind of stuff is like it activates all the, all the stuff in you that makes you want to care for it and make it good. Yeah. And by the same thing, it's literally the same thing and by caring for it and making it good, it's writing software to make itself better and you solve, you solve a lot of problems in AI with like around trust and all that kind of stuff through this weird mechanism that you wouldn't necessarily predict beforehand, but once you see it, like, oh, yeah, obviously, this is how it would work. Yeah. And I find that I use mine very much as like an actual in-person assistant, like I've got it trained on, hey, I need one of these things and it just, it pulls a skill knows what this thing is and really quickly delivers it back, can drop it into a Google Drive for me or whatever the case is.
And I still find myself leaning on codecs and codecodesome for like, I want to, I want to create something, like I feel like if I want to create something, I still sit down at the coding agent, but with like, it's almost admin almost, it's almost a companion in some ways, like, you know, you're, you're a bad guy. Yeah. Totally. I agree and I think that there's this thing, this other really interesting benefit of them is that they're connected to everything, so they know about everything in your whole life, in your whole work life. And that makes them fundamentally more useful in a lot of ways. And yeah, like if I'm doing something like serious coding or serious vibe coding, if that's a thing, I'm really going down to use codecs usually and like a little bit of cloud code. And it just feels like nothing's going to get lost. Like I think one of the problems with, with open clauses, like their memories, like kind of like shaky and like, you might all up an hour later being like, hey, did you fix that
bug and be like, what bug? And you're like, I'm literally going to kill you. And luckily, you're out of house and home with tokens if you're not careful. Yeah, yeah, yeah. So I, you know, that's why I'm, I'm really, I'm a more of like an agent maximalist, like we're going to have a lot of different agents doing a lot of different things. And the ergonomics of a codex right now for staying organized to make sure you actually like finish the work that you do and it's done well is, I think, quite helpful. But then for your claw, like I'm just constantly being like, okay, how's, how's the usage today or, you know, like file this bug or like the whole, the whole way that we do bug reporting is changing first and proof. But I'm hoping what we do that through the rest of the org where because proof is agent native, meaning agents can use it as first class users. It has a bug report function. And so agents just submit bug reports and the bug reports you get from agents are way better than the bug reports you get from humans, even if they're initiated by the user. Because they can be like, hey, like this is exactly what I did.
And here's the exact error message and here's like, like the line of code where I think it might be or, you know, depending on how much they know. And, and what happens is every morning, my agent, RTC2, then goes through all the issues that were submitted by agents and then clusters them and says, like, here are the main issues that we need to solve. And so that's just, it just like totally changes how you, how you work, even if you're using codex mostly for coding. Yeah. Okay. We got a question from the chat. And I think we can tie this into some of the other conversations we want to talk about. Rave Master 2000 says the only thing I don't know about AI is how to get the most out of agents. And luckily, we have Dan here who has quite a lot of ideas around this. I don't know if we want to jump right into agent native architecture or how, how would you address this, Dan? Well, I don't think that there's any one answer to that question. I agree. And there are, there are like some of these things that are like real, wow moments with agents.
In particular, let's, let's just narrow it down to open call because I think there's like, there's lots of different agents and they mean different things in whatever, but open claw or plus ones. One of the, one of the coolest things, for example, is once you have connected all your stuff to it, it can do like a really nice digest for you where in the morning, when you wake up, it's like, here's the weather, here's the stocks, here's all your, here's all your newsletters, here's like what's going on in your schedule. And people are like, kind of blown away by that. Yeah. That's the first thing I built. Oh, really? Yeah. Yeah. Yeah. And so plus ones come with that already built. So like what we try to do is figure out, okay, open claw is this like blank canvas. Like, what is, what is, what's the happy path to get you into it so that you don't have to think too much. You don't have to set up a Mac mini. You don't have to do any of that stuff. And it comes preloaded with things that like get you right to the like, wow moment. We also wrote a guide called open claw, the comprehensive beginner's guide, which I'll drop in the chat. Yes. Yes.
And that has, that has a lot of like our own ideas and experiences with what, what makes these things awesome. Like another, another like magic moment that happens a lot is if you work out using it to like help you plan and track your workouts and like find new workouts, it's like a huge like game. Oh, that's cool. For me, I think my magic moment was, um, was, uh, using it for reading. So I'll take a picture. I'll send the picture of the book to the, to my claw and then my claw keeps this like web page of all my reading notes on it and that's like really sick. Um, wait, how does that work? You like dictate your notes as you read or, yeah, I just, we just have a little conversation about this thing. And I'm like, Hey, like, okay, I highlighted this like throw it in the my book notes and it's like, great. Um, that's so cool. I think that the, the other big category is just seeing it do something you didn't expect.
So, um, Brandon, who's our COO, he was doing his email and had to run. So he was like, Hey, to his, to his claw, he was like, Hey, um, can you just like call me so we can do my email along while I'm walking? And it called him and he was like, that's crazy, you know, uh, how, how did it do that? I don't, I don't know. Exactly. I mean, it's just, I use like one of the, of it, like easily available APIs to, you know, probably use Twilio or something like that, um, and an art hit access to his email and, and it had it. It's kind of hard. Oh, it just like did it. Yeah. We had a weird one with my wife where one day she was in the living room. I was up here working and, uh, all of a sudden at like max volume, my computer started talking to her in the living room and, and she's texting me like, Hey, something's wrong with your computer and then she's like, it's talking to me. Next thing I just standing here at the corner of my desk, like, Hey, uh, so I get down there and it's got a brave browser open with like 75 tabs and, uh, one of them is some
university lecture on efficient tokenization that it's playing up on YouTube, yeah, it was, uh, I've had a few just little, little funny, little, little moments like that. Something I love that it unlocked for me is that I have spent years trying to find the right task app, like a good checklist and you know, I'll get one and I'll be like, this is great. And then like three days later, I never used again. I've paid for a year subscription and I forget about it and I forget about it. But like, I've got, I've got like cronjup setup to where, where my claw is able to be like, Hey, it's the end of the day. What's on for tomorrow? Do you want me to cut this? Did this happen? What's your, what have you missed? What's the big? You better get this tomorrow or you're unemployed task, you know, uh, and it's really been helpful because it comes to me and I think the fact that it talks to me in telegram makes it feel more like getting a message from somebody asking.
Yeah. And I seem to be more likely to go reply to it. That's really interesting. Where are you keeping the tasks or using an actual task manager? Is it just like keeping them in markdown or something? It's keeping them in markdown. It's keeping them in markdown rolling tally. I give it my, it doesn't have access to my work calendar, but I give it like a download once a week is like, Hey, here's the calendar. And so it'll keep track of those. It's doing most of it in markdown, uh, a little bit in notion, depending on, on which one we're doing. I've been trying out different areas to have it manage things as we go to. See, see which ones like most effective. Yeah. Yeah. Um, we got another question from the chat, FTLD says, are you making a ton of alternate accounts for these agents on the apps we use every day or how are you dealing with accounts or permissions? Good question. Um, well, a lot of the accounts, you actually just make a, um, an AP Aki or you sign it with Olaf. Um, so, and different people have different levels of comfort, like I know a lot of people
who are, um, like have a separate email address and all that kind of stuff. I'm a little more like, I give it access to my, my, not all my accounts, like it doesn't have access to my bank account or anything like that, but I give it access to, it has access to my work email, for example, um, and, uh, and my feeling is as long as you have really locked down the server that it's on and the channels that people can access it so that you, you're really the only one that can access it, then it's like probably fine. But yeah, different people have different strategies. Cool. Uh, and then another one, Sako Bambino says, how does your AI, agentic based product avoid or address model drifting issues that we often see with most of the AI models? So, I guess this is referring to plus one, right? I, I guess so and by model drift, are you talking about the, like, you know, uh, as models change, the harness is not as good or is, am I missing something about model drift?
I think that's what they're talking about, um, yeah, a lot of times I'll hear it referred to as talking about where, like, over time, the answers start to skew and are just, just kind of generally less good and it's, uh, oh, like, you just get used to the baseline. Yeah. Like, not just in a long context, but even over time where it just kind of, sometimes it'll get continually, continuously, I don't know how much of that is the actual model or if the human reaction to the output, I don't know, Dan, if you know about that, but, I mean, over time, it should not be over time because like each chat is like, basically new. Yeah. The over time thing would be as the, maybe as the models get updated, the, like, harness is not as good and luckily, you know, this is based on open cloth. So, but that, this plus ones are based on open cloth, so we don't necessarily have to worry about that so much on all of our products. I have this, like, philosophy of your job building products in AI is to surf the models.
And what that means is every time there is a new model update, you have to figure out how to use your product, how to, how to build your product and also modify your workflow to get the absolute most you can out of the, out of the model. And that's the way that you take advantage of model progress. And that's why you don't get your, your lunch eaten basically by like models getting good enough that you don't need an app. And what that requires though is you have to be willing to throw out your whole product or most of your product and, and a lot of your workflow every three to six months as the models change. And that, that kind of sucks, but also it's kind of awesome because you get to continually push the frontier and it's so much easier now to like rebuild products. And so I think that also kind of takes care of model drift. Like, I'm not, I'm not necessarily at this point trying to make something that like last, like make one piece of software that lasts for a long time. I'm trying to solve a, a, a task or a workflow for a certain kind of person.
And that will take many forms as the models get better. But it will still be a thing that people need to solve. It will be always evolving kind of thing, probably, yeah. Which it always was before it just, the rate of progress was slow enough that you could, you could like sort of kid yourself that you just do one thing and it's always good. You know, and that just wasn't ever the case. Well, those were the days, yeah. Yeah, when things were slow back in the .com boom or whatever, you know, it's like, yeah, it's not going to be no clarified. They said, trustworthiness and accuracy of the deliverable is what I'm referring to. How is the logic mechanism of the agent is safeguarded against possible drifting just to clarify myself? You know, this is, this is the, I assume maybe then what we're talking about is like hallucinations. And the, that's one of the real benefits of having a agent who is tied to a person is, you're using it all the time for yourself. So you're going to, and you're using it usually in places you're in experts.
So you're going to have a pretty good idea of like what is good at when it's not good at. You're going to have to fix things that are wrong and it's being used publicly by other people on your team and using a way that might reflect on your own reputation. And so there, there becomes this like major psychological incentive to make sure that it is working well. And I think that is sort of solves a lot of the trust problem. Yeah. That makes sense. Related to this, the blind dragon 13 just asked, have you found a great memory system that actually works for your claw? I have not. I think that the, I think it's out there though. And Willie, who's our head of platform is the guy that's building plus ones. And I know he's like experimenting with a lot of them. I don't have one off top of my head where I'm like, you should go check it out. But we'll definitely, yeah, we'll definitely put more memory stuff into the plus one. And my friend Nat Eliasin is also, I think, really good and Johnny Miller. If you haven't checked out their stuff, they're, you know, they're, you know, I feel like we're pretty head, but they're like miles ahead in terms of how, how claws, how claws
work. So if you're looking for memory systems, I check out what they, what they do. Cool. Let's, let's talk about proof because this is something that I think a lot of people can relate to. They have, we have these amazing coding tools now. People are, you know, if they're not building stuff, they want to build stuff. You actually built something and released it. Tell us, walk us through that whole process and what happened. Totally. So proof is an agent native document editor. And the underlying thought behind proof is most word processors are built for humans. And now that we're, I mean, all word processors really. And now that we have AI, we're kind of like bolting AI into it and trying to make it so that it can like write like you so that the stuff you put into the word processor is like mimicking what a human would do. I think that's, there's, there's a whole interesting line of work there. But there's this other thing that's happening, which is that I am actually reading a lot of AI writing. It's doing a lot of writing that I would prefer to read the AI's writing.
I don't want to read a human's writing. And that's in certain tasks, so like planning or like especially planning a feature in your coding app or, you know, a research report that uses a bunch of our like growth and stripe data, for example, I, if I asked a human to do that or a bug report, if I asked a human to do it, it's just going to be worse. And the, and the, the way that agents write documents right now is they write markdown files that are on your computer. And that's like just kind of clunky and, you know, if I try to open it, it opens Xcode and it's just not great. So what proof is is when an agent writes a markdown file, a plan, a research document, anything like that, it can just like put it in a web page that is collaborative. So you get a link. You can open it. You can write comments. You can type in it. You can, you know, do anything you would expect in in Google Docs. You can have your agent in there.
You can have other humans in there. They can have their agents in there. So it's a really good way to collaborate on documents between humans and agents with, but the idea that most of the writing is, is from AI and we also track and make it easy to see who wrote what. So you can be like, okay, I know most of this is, is AI written, but there's this little section that was written by human and I, I assume if they wrote it that they really wanted it in there for a reason and I'm going to pay attention to it. So that's, that's the idea. I, I built it, I, there were a couple different versions of it. I first built it as a Mac app and then I realized it should be a web app and so I like pivoted it to a web app like two weeks ago and it just kind of took off internally at every. Everyone was using it to share files and plans and all that kind of stuff and when we see that, it's usually a good sign like, hey, like we should release this and what I decided was it would be really cool to release it for free.
So anyone can do it without even logging in, log in unnecessary and open source and so we launched it and it went viral and people loved it and there was like, I don't know, four or five thousand documents created in the first day or two and so it was really cool and I've coded it and so there were a lot of, there were a lot of problems with it and say more, say more, like what are we, what are we talking about here? So collaborative documents are, they're effectively a solve problem. Like there's a couple of open source, like well known open source libraries that, that make, doing collaborative documents like fairly easy or like it can be fairly easy. And so I obviously like, I knew about those things and I asked Codex to use their, the stack I used is YJS and Hocus Pocus, YJS is this like underlying like library for collaborative
documents and Hocus Pocus is like a wrapper around it. And I asked it to use that and it did and it was working and what happened was as it was working, it hadn't really read all of the like YJS Hocus Pocus best practices and there are a couple like things you need to do at the very start of your project and a couple of ways of thinking about how data should flow and who gets to write data when for example because it gets very complicated when you have like, you know, someone typing over here and someone typing over here and an agent over here and you're trying to create a like unified always up to date version of the document. You have to be pretty careful about how you set that up so that no one gets confused right because you have to sink in between these different, the different exactly one. Yeah, yeah. And so there are a couple of specific ways that a specific best practices for how you set it up that make it, make it fairly simple and make it fairly unlikely that there are any
problems and Codex just actually doesn't know about those which was surprising to me because it's a fairly popular library and what I would normally do for any sort of production project is like when I'm in the plan mode, I'm like, hey, can you figure out how to, can you figure out all the best practices for this and I didn't do that. That's a really good tip for people, like when you're planning with your agent in the beginning stage, like make sure that it's like, okay, look up you know the best way to do this, because otherwise it will just riff, I guess. Exactly. And we have a plugin that we make at every called the Compound Engineering plugin. Yes. That has a plan mode that's like really good, really rigorous. Kirin, who's the GM of Cora, who made it, is like amazing and the workflow that he invented is I think incredible. So yeah, so I should have done that, but I didn't. And so what started to happen was I would start to have, we would start to have problems
and I'd be like, okay, here's the bug. And it would go off and research it and then fix it. But the fix was always like a sort of like duct tape thing, because it didn't want to go like solve the like really deep underlying thing, because like I said, it was going down and whatever. And it may be what a human would do. Exactly. And it depends on the human, but yes, it's only what I would do. And so basically like that just kept happening and so it kept duct taping and putting little guards and checks and like all this stuff here. And as a site started to go down, the complexity started to go up and each fix like would kind of fix it, but then kind of make it worse. And you know, I had a couple of experts look at this over the last couple weeks, because one of the fun things about doing this stuff so publicly is like when I, when you tweet and you're like, hey, like my thing is down, people who have a lot of experience with
YGS come out of the woodwork and be like, hey, like I could take a look. And what's really cool is like, because I like sent them the repo and then duct, you know, because I was just like, this is like, they're going to judge me so hard. Yeah. Yeah. I was like, I'm so sorry for this. What's really interesting is they're all like, yeah, this is actually very reasonable. It lacks some amount of coherence so you can, you can see that the agent was like solving local problems in a particular way, but then not zooming out and being like, well, I saw like this over here and like this over here and they should, those should match so I can understand the whole thing like it wasn't making like that. Which is what a good engineer would do. I actually don't think that that's a permanent thing. I imagine that that would get better over time. I actually want to get on the problem. Yeah. That's a context window limit thing like where it just can't possibly think about the whole project at the same time or it might be, it's also like a prompting thing.
I'm sure if I like prompted it a little bit better, it would be, it would be a bit better. And also just honestly, it makes it better to do these kinds of projects if it's a production app. If you hold in your own head, like the basic way that the architecture works, it's just right. Yeah. And they're like, you have to remember this thing is a super intelligent thing that pops out of a box every time you prompt it and it like hasn't, doesn't know anything for the last year and it's never seen your project before and it has to get up to speed every time. And like that just makes it, it's just hard, you know. And extra 30 seconds to really explain what you want clearly can probably save you a lot of grief on the back. Yeah. Exactly. Explain what you want and also part of that is knowing what you want and especially if you've coded it, like you may not fully know. And so basically like I had someone come in and help me just, okay, just be like, okay,
if we went back to first principles, how would we architect this? And I guarantee, like I already knew most of what he said, it just like wasn't fully there because I was like trying to transition from, I didn't even know that codex didn't know the best practices to, okay, I'm having codex see the best practices, but it's still like slightly doesn't want to do the full like rewrite and delete a lot of code and whatever. And the guy that I brought in who's super talented was just like, yeah, this is, this is exactly the thing that we need to do and like I'll just rip out a lot of the code and he used codex to do it. But wow. It's not, it doesn't necessarily come naturally to the AI models to do this yet. And now it's like fairly stable. I'm like happy with it. You know, some of the code is different, but it's not, it's not like a totally different app than it used to be and it happened very quick. Like he was able to essentially stabilize it in a couple days of work. So it's pretty crazy what you can do and it certainly at the scale that we're at and
then and the level of sleepless nights that I was having. I probably would be more careful next time, but I think generally this is fine. Just curious. Did you use codex's plan mode first? I was using codex's plan mode, yeah, I was like, I had really good luck with plan. But a lot of times I'll even go and use a different AI and be like, okay, I'm going to workshop this idea. You should know that I don't know what I'm doing in time. I have ideas and I can understand it if you tell me, but don't assume I know anything. And usually I can get like a good, here's the project I want to do prompt that has saved me some grief. But it's always a coin toss, you know, you just never know. And I think you also kind of hit on that element of this thing we've kind of seen since this all began is that in the hands of an expert who knows the right questions to ask,
it can do a lot more. Yeah, it really can. It contains all of the knowledge of all of humanity and you're only like kind of getting a little slice of what it knows, you know? Based on what you know to ask it, yeah, exactly. So it's a real skill to use these things and I don't think that's going away. No, I don't either. Do you think prompt engineering as like a skill is worth investing time in energy into still or is it more, is it more like if you talk to it long enough and you give it enough information, you'll get what you want? I mean, I just think a prompt engineer, I think everyone is sort of a prompt engineer but it's not like there are those little tricks that are, you know, like I'll pay you $2,000 or whatever and that you don't have to do anymore to get better results. They'll probably will probably always be like certain things you can do to like make it better but the big thing is knowing how to manage the model, knowing how to ask for what
you want and know if you're getting it back and that's like kind of prompt engineering but it's very specific to your workflow and what you want and the kind of thing that you're doing and so I think it's prompt engineering is like, maybe like writing, you can like you write in all these different circumstances and you can get better at writing but we're really talking about specific things that you're trying to get done, you know. So I don't expect people will study prompt engineering but I expect that they will know some basics about certain little tricks but mostly how to do it well for their specific use cases. It's almost as much about problem framing as it is about prompt engineering in itself and just understanding. I think prompt engineering has a role as like, here's step one for normal people. I think when average folks are like, I need to learn how to use AI, they can get a lot of unlock just from, you know, a quick prompt engineering course and course error or
something, you know, you really can, if you're new, really unlock a lot of capabilities you didn't have prior but it's very much the starting point and not the end zone anymore I would say. I agree. Well, I have one more question about proof. How well do you feel that you actually understand the code now and yeah, do you feel that? I still don't understand it. That's because I hired someone who's, he understands it. Okay. Cool. Somebody does though. That's all that matters, right? Yeah. Somebody knows. That's how I would modify this. Like we had a little bit of a retro at every about this whole thing and, you know, when we launch something, sometimes we label it as an experiment and it's okay for it to be like a little rougher on the edges but if we're really launching something, we want it to be good. And so, and also on the other hand, when we launch something, I don't want to have to be up all night like seven days in a row trying to fix it for my own health and for real. So what I've realized is we need a buddy system where if I'm doing this, I need someone else
who knows the code base and like knows a little bit about it so that when we launch it, if there are problems, like it's not just me and like my foxhole, you know, trying to like understand the code base while it's going down and everyone's looking at me because that's, you can switch off, you can switch every other night, exactly. And that led me to this articulation of how early product engineering teams should work, which is the pirate and the architect model where you want a pirate and that's, I'm a pirate, which is like, you're just going as fast as you can. You're just trying to find something that like works and people like. And then the architect is a little bit like I want to really understand how the whole system works and make the whole system work together well as a well oiled machine. And especially in early product work, you actually don't need a full time architect. I think you just need, you just need a pirate just like going hard and an architect coming in for a couple hours a week to be like, here's how all the things work and here's, here's
how we can tuck in the edges a little bit so that the core of it is stable, but you don't really want someone spending all their time making it perfect if you don't know if it's any good and you want to be able to explore if that's possible. And sometimes like some, some people on our team are kind of both, they can kind of flip in and out of pirate and architect mode. And I'm just like not that, like I just, I'm not care. Like something is very careful about this, I'm not. And so. You also have a lot going on, I mean, CEO, you're managing the machine company. Yeah. So I think that's a good, I think that's a good model. I would expect to see more of that and we definitely see that across, we've run five or six products internally and we definitely have, we have one person who's fully responsible for it all the time and then they often have one or two people who are spending some, some part of their day on it, helping them with, you know, some of the big, difficult more architecty tasks. Yeah. That makes sense. So I think this is a good transition into agent native architecture. The second you publish this, I fell in love with it.
I think it's an awesome way of thinking about building applications to ride the, or surf the models, as you said. And it was cool to see your interview with Mike Krieger from Anthropic yesterday and he said he actually uses it as a skill. Yeah. I don't know if you, I would have geeked out if I were you knowing that it was awesome. I was like, oh, yeah. Yeah. But yeah, could you just tell us kind of what that is and, you know, like if people are interested in building with agents, you know, how they could apply it themselves? Yeah. There's a new way of building software. I've been calling it building software that's agent native and it implies a new architecture for how your software works. And the way that you can think about it is normally in software, any piece of software is like a recipe. It has a set of steps for how it works and those steps are known beforehand by the programmers. And this new version of software, instead of like the whole thing being written out beforehand,
it's essentially cloud code and the trench code. It's like you take cloud code and you put some like nice buttons on top of it so that you're interacting with the UI that feels familiar. But when you press the button, it sends a prompt to the agent and the agent gets the work done as opposed to like running the recipe to get the work done. And there's a lot of really interesting effects of that. In particular, the interesting thing about agent native software is anything a user can do, the agent can do. So if you can push a button that prompts the agent, the agent's going to have to be able to like do anything in the app and that's super powerful. Another thing is that it creates this flexible way of working where the programmers don't necessarily know all the things it's going to do off like when they release it. It's going to do things they don't expect.
So I think cloud code is like the canonical agent native application. It's an agent. It sits on your computer. It has access to everything on your computer so anything on your computer you can do. And it works in this way, in this flexible way where it can run any bash command on your computer. So people are using it for code, which is the thing that it was intended for. But then they started using it for everything. It's like organize my files or like plan my schedule or whatever. And they're like, oh my god, this is so cool. And then they made co-work. And so it creates this much more flexible model of software that doesn't mean that traditional software doesn't work anymore. I think the whole sass is dead thing is like such bullshit. But there is this new class of software that I think is really powerful and proof is is an example of this kind of thing. There's different types of agent native. It can be agent native in the sense that it has an agent at its core internal to it or
it can be agent native in the sense that all agents can use it natively. So like Figma at this point is agent native because it has a CLI. So there's a whole new world of how software might work and also how software works with agents that it starts to open up. Does this apply, would you say that this also applies for people who are like building an agent to help them in their workflows, like not necessarily building software but like trying to, you know, whatever tool they use where they're having an agent deploy in production or is that a separate problem? I mean an example. So let's say you're someone who's not technical and you want to build an agent to help you whether it's open with OpenClaw or I guess actually OpenClaw would be a good example. Like you have an OpenClaw set up. You want to have your OpenClaw do something for you, say like help you manage your schedule. Would this apply in that circumstance to or is it more specific to your building a software that people are going to use? It definitely does apply but I think that OpenClaw they've already built it to be agent native.
And so you're kind of getting the benefit of writing along with that architecture. So an example of what makes the agent native is OpenClaw is built on pie, which is like agent harness. And pie is very, very basic. It doesn't really have much except an agent loop and the ability to modify itself. And OpenClaw puts a couple things on top of that. So it has a cron job. So it has a heartbeat so it like wakes up every 15 minutes or so. It connects natively to a couple of messaging apps but that's really it. And the core of OpenClaw is still this thing that can modify itself. And that means it's super flexible, right? Like Peter who built OpenClaw, he didn't, I use it for bug tracking and triage like he never built a bug tracking and triage feature into it. And the guy who made pie never built a bug tracking, it's not made for bug tracking
and triage. But it's just flexible enough and its tools are granular enough that it can be used for anything and that's the interesting part of it. That level of simplicity is what led to what is arguably one of the most transformative software we've seen is kind of awesome. Totally. It really is and it opens up a way of thinking that is quite different from how programmers normally think. And it's actually hard to get AI to think this way because it's trained to think like a programmer. And what programmers want is they want to be able to predict what's going to happen. Like they want to make a machine where they know how it all works. And I think that's one of the reasons why it took a long time to get something like a code code is because we were pretty afraid to unhobble the model. That's what Anthropic talks about internally unhobbling the model. We kind of like really locked it down and we're kind of like, oh, it's going to be in
this very specific type of workflow that it's going to work for. And the real answer is actually give it a basic set of general tools and let it run in a loop and people will figure out how to use it for whatever their specific use cases are. It's just a new way of thinking about software that's both scary and extremely useful. Yeah. You know, in the way most of this is built in big research labs with largely by software engineers, I think there are a lot of capabilities that get skipped over that maybe they don't even know where they're in some ways by just simply not taking that approach. I always tell people, ask it something you don't think it can do. Always ask the thing. Usually it will surprise you at how close it'll get. Yeah. I told the agree and I think that's why some of the stuff we're doing, the big model companies appreciate it and pay attention to it because they do have apps that they work on internally. But I also think of them a little bit like oven makers.
So they're making an oven. You can use an oven for a lot of things. And we make soufflé's. And so they make a new oven and they come to us and they're like, tell me about the soufflé you can make, you know, because that helps them figure out like, how do I make the oven better? But they're not going to make it into the temperature. Yeah. They didn't rise enough. Exactly. They're not going to make an oven just for soufflé's, but they, but it takes them seeing it being used in a particular context where someone's like pushing it as far as it can go for them to even realize, oh, there's an opening here. There's a vector along which I want to improve it. And I think people don't quite realize that and quite realize the difficulty and the difficulty and also promise in making tools that are so general that you can't fully predict how they'll be used and how to improve them. So the dominant startup metaphor for the last 10 or 15 years has been jobs to be done. You have to narrow in on a specific job that your customer needs your model for and or your product for and then make it like better.
And if you ask, okay, what problem does AI solve? It's like, well, it solves every problem. Theoretically, it could solve every problem. It's better and worse degrees. And that's really, it's really hard to figure out how to improve a product that's really meant to do everything. Yeah. Yeah. Well, first of all, I can't go any longer without addressing this before we move too far past the pirate thing. JD Burrow in the chat said, I'm late to the party, but is Dan's last name truly Shipper or is he just emoting pirate vibes? I didn't even really think about it because usually people talk about Shipper as, you know, shipping code. Yeah. Right. But I do love the Shipper as pirate thing. I didn't know a lot of that is it is truly, truly Shipper, but I have to make it make sense in every way possible. That's right. I love that. I saw that in the chat. I'd been eyeballing it too great. Had to sneak it in there. I want to get to a compound engineering as well. Dan, how long are you here for?
Are you here to 11? I can go a little longer. Okay. Cool. I just want to make sure we get to everything. So someone else in the chat asked Rave Master said, what's the first go to get started with agents? Is it usually VM where clawed instances or something else? Well, you really teed me up there. You should try plus one if you want to get started with agents. We have a new product on every every.io slash plus dash one and it is your very own hosted open claw instance. You can get it with one click. It has all of our every apps on it to help it do email to help it right well. You should really check it out. It lives in Slack. It has all the right presets. Other than that, yeah, I think the VMware one is pretty good. There's I think there's one for railway. It depends on your particular setup and your particular thing that you want to use it for. I like having a Mac mini. I think that's kind of cool. I think they're sold out now.
Definitely in Silicon Valley. Where are you guys based, actually? Or in New York? I'm in Brooklyn right now. Oh, New York. Okay. Yeah. I can also see that. I can't find anyone on that side of the country. I bet. Yeah. I bet. Cool. Cool. Cool. Cool. Cool. Cool. Cool. I wanted to ask one other question. Actually, let's just get into compound engineering because I think that would be kind of cool to talk about. Whether they're using it with their open laws or their coding agents. Can you tell us a little bit about this? And I'll add a little bit of context. Kieran was like the first, who's an engineer at every, was the first person I saw besides like faceless people on Twitter, who was like maximally going hard on multiple coding agents and sub agents, and like he, in my opinion, led the way on a lot of this stuff. And ever since, and he came up with this framework, but ever since I saw him, I was like, oh, this is like the new way to do things. Hold it. He's a true trailblazer.
And one of the few, I think, senior engineer types who are like willing to give up coding, like manual coding, even before it was, even before it was obvious it was going to work. Yeah. And I think basically what had like, I remember very clearly there was some model, some model came out. It was Claude Opus 3.7, and we were testing it before it came out. And when, when we test models being here and often are like on a video call together and just like chatting back and forth, and he was like, I don't, I think this is just work, like I don't think I have to look at the code anymore. And so we were like trying that and we were like, holy shit, this is crazy. And this is like maybe, it's probably almost a year ago now. And that then filtered into the rest of our company. And we started being like, I don't think we need to look at the code anymore. And that became a thing that we were doing, but everyone else is like, that's crazy.
And now it seems, you know, pretty much like that's the case. Pretty normal now. Yeah, it's pretty normal, it's pretty normal. And out of a lot of his experience with that came compound engineering, which is the idea that in normal engineering, each feature you build makes it harder to build the next feature. Why is that? Because the code base grows in complexity, all the complexities, interdependent usually, you know, you have all these tests, you have a bunch of stuff that like all depends on each other. Even if even in a modular code base, it still is like that. And in compound engineering, what you're trying to do is make each feature easier to build in the last, and that is by, as you do things, you compound the learnings from each feature into the next one. So like each bug that you find or each issue that you make, or if you're proof like each first principle of building YJS applications that you miss, you like compound that into a research, into a knowledge, into your knowledge base and your repo so that every engineer
has access to this. So. And Lincoln shared in the chat. Yeah. I'll just put it right in here. I don't know if you've got it. Yeah, go for it. Yeah. That should work. Yep. So I'm just making sure that was right. Okay. So yeah. So that's basically how it works is, it's a plug-in, it's a philosophy, but it's also implemented in a plug-in where there's four steps to it. The first step is planning where it takes into account like all the best practices, all the stuff you've learned in building your product, all that kind of stuff, it makes it really, really detailed plan. Then you kick off your agents to work on it and often you do that in parallel, so you have a bunch of agents all working in parallel. Then you review or assess what happens. So you have tests and you maybe test it manually. You maybe have a fleet of agents to testing and then you compound that learning. So you take everything that you learned and push it back into the first step of the process and that's what makes it compounding.
And I think this is something we were on early, this is something that Kieran in particular really noticed and was like, I'm doing this and I was like, that's sick. And I think has become, even if it's not call compound engineering, has become like a standard way that a lot of the model companies think about building engineering, like harnesses and doing programming. So it's, yeah, it's pretty cool. That's really, yeah, that's awesome. So we have how you and I test models, Grant. Me and you. Oh, how so. What do you mean? Sometimes it's hopping on a live stream and, oh, yeah, and throwing in stuff to see how good. Well, sometimes it's actually live stream. I was just meeting a Google call, a Google call or something. Yeah. I got to. Yeah. That's really cool. Why are we making that? There's a way to, there's a way to do compound engineering with like just being a regular anthropic user and it's with skills. Yeah. I don't know if you do this, Dan, but basically whenever I learn that the model can do something
that I'm often trying to do myself, I will turn it into a skill so that I don't have to ever prompt it to do that again. I can just say, you know, do this or in the off chance that it doesn't know to use the skill. I'll say use x, y, y, y, y, yeah, yeah, yeah, totally. And yeah, different people, like using, even if it's not a skill, it's just like, remember this for next time, you're effectively compounding. It's the same kind of idea, just out of bigger scale. Yeah. Yeah. Yeah. But the plugin is great for the engineering loop and I don't know, could you use it in like a regular workplace setting, like definitely a lot of people use a lot of non-typical people in co-work, for example, use it all the time and love it. And that's awesome. I'll try that. There's something about the models right now, which I think will probably always be the case where if you just get them to think more or like use more tokens on your problem and do more research and spend more time on it, you just get better results. And I think compound engineering, the plugin is a hack for that where all of the things
that the model comes back with and you're like, I don't know if this is like totally right or like it should have done a little bit more research here. It's just really good at getting the models to do the maximum possible amount of work. So for important stuff or stuff that requires a lot of thinking, it's a really good work flow to use. You know, I've gotten into using coding agents for everything now, like, I mean, even codex, I've now built out codex to do so many things that have nothing to do with like projects I'm working on. Like it's very much, you know, skills and automations and a little bit of everything. And like, I've gotten to where I love to write in it because it doesn't use like M-dash which is a lot of the normal AI tells are gone, which is an interesting thing because as Grant put it, code bases won't tolerate that. No, no, that's exactly why I'm like, no, it's not going to have M-dash us in your code base. What are you? That's really interesting. Wait, so when you write with it, what's your workflow? It's very simple. I usually have, I have a number of different skills that I use in there for like, I don't
know, this is my sub stack, this is a blog article and, you know, and I've essentially over time just managed to extract what a Corey blog article is from like my favorite pieces I've ever written over the years and had it analyze them and pull out the qualities that made them me, I guess. And then just build a skill where it's as simple as saying, hey, I want to write this use the skill. And you like that better than using 5.4 in Chagypt? Magnificantly. That's interesting. I really didn't think, I mean, I really thought, you know, the Chagypt app was something you'd proud of my cold dead fingers, but I've really grown to enjoy it in codec specifically. I think this is a really interesting thing. I mean, I use codec much more than I use Chagypt for things that I use use Chagypt for and the thing that happened is when around the time when me and Kirin were having this
like a whole realization about 3.7 and do you need to touch the code and whatever, that was a little bit after that was when Chagypt 5 came out. And GPD 5 was a very interesting model release because they continued, even though all this agenda coding stuff was happening, they continued to push forward this split between regular knowledge work and vibe coding happens in Chagypt and like professional pair programming happens in codecs. And they stuck with that split from GPD 5 really until like the last two or three months. Yeah, even maybe in December, even at like, exactly. And I think that really hampered them because what anthropocos able to do is a cloud code, like, people started to be like, there's a whole new engineering paradigm that I can't do with codecs because it's too hobbled and doesn't let me do this. And then people were like, but I could also do it for all this other work.
And so it started to explode. And I think Chagypt got left behind a little bit because Chagypt as a, as an app, it doesn't have access to your computer. It's like, it has desktop app, but like, it's, that's, has it always been sort of a site show. Yeah. And it's really like this chat thing. It like, yeah, you can use it. Yeah. It's just like the Chagypt website, just, just living in an app is what it was. Yeah. But you can double tap to open it. Exactly. And cloud code is just much more from the beginning and agent that you hand things off to and that has access to your whole computer and your whole life. And that's a much better, much more fertile ground to grow, more power and more work, it can do more work for you than, okay, it's really a website and now it's a mobile app. And now it has a desktop app, but like, it just doesn't, it doesn't, it didn't work for them. And I think that they, to their credit, even though it took a while, they realized this over the last like two or three months and like totally pivoted codecs and we're leaning into it super hard. I think one really good piece of evidence for that is the super bowl commercial they ran
was for codecs. It was not for touch. Yeah. Right. And I, and you hear all this like all these rumors about the internal kind of, they're reorganizing that cut Sora or whatever. I think that they're realizing this and they're like going full bore into it and I mean, so far, I really like codecs. I think cutting Sora was smart, like, you know, the fact is if you've got a big model around the corner and it needs a GPU, Sora was not cheap. Like, like, it needs to somehow be either making money or bringing in new users or something to justify not redirecting those GPUs over to, over to putting out something really cool. Yeah. The latest is that they're going to do basically Atlas, chat, GPT and codecs as one app. Yeah. I mean, that's interesting. I mean, codecs as an app on its own is just pretty great. Like, I don't really, I, I worry that if they try to combine it all together, it might not be as good as what, what they have now. I agree. I know that they're going, I mean, just not from any internal knowledge, but I do know like
just from looking at their tweets that they're really, I think they're really looking at re-architecting the codecs app, but hopefully in a way that's not like, we're making a chat GPT, but in a, we've realized how powerful this can be and we were like pushing it as far as we can and I think they're going to do a good job. They've been working on it in a way that I really like too, that is like reminds me of early perplexity where you'd see Arvind Srinivas on Twitter at night and he would be like, what do you want? Give me a feature. Yeah. And it's that way on Twitter right now, every single night with the codecs team and they're like, what do you want? Why do you want it? How would that work and like asking follow up questions? And then I think, I think that is such the way to build in 2026 is like, get that feedback right there. I mean, you know, the fact of the matter is if these are the way you want to please find out what they don't like and it's tough though because of course, you know, you've also got to separate, you know, signal from noise, which is tricky, but they're pretty good at that.
It's tricky even with a newsletter. You know? No, that's true. Where's like, when feedback comes, it's like, okay, is this, do we care? Do we not? Do we, is this an angry person? Is this person with a good idea? I think it behind being a jerk. You know, it comes all ways. Two more things because I know we're over the hour, Mark, that I want to touch on. I want to touch on, you know, your product suite. We mentioned agent native. Are all of your tools agent native now? How would you rank them and if you could tell us a little bit more about that, the other thing is, is this any advice for agent building and shipping in the agent area? What's really funny is if you talk to anyone on the every team and you ask them what my favorite word is or phrase, they'll say agent native because I just like, not, I can not stop saying it. Well, you coined the term as far as I know. I think so. I think so. I think you did. And it's just like a thing that I realized in December and then came, we came back from, you know, Christmas and years and I was like, agent native, agent native, is your
app agent native? Is it, is it? Is it? And we're really getting there. So Kora, which is our email agent, it has a CLI and that's really cool. And we're launching a new inbox for it. So you, you can just manage your entire inbox with Kora and it has an agent or any one of your agents can, can use it. Also, I think that's going to be amazing. Spiral is our, our ghost writer with taste and that is definitely agent native that has a CLI. So when I use my, use my plus one and I'm asking it to like write tweets, it'll just like go to spiral, talk to spiral about like what is, what is trying to write? Spiral has my voice and style and it just goes back and forth and gives me a few options in us. I think the writing is wonderful. So the agents are talking to each other, the two different agents and two different times are talking. That's, wow. And what's really cool is like, okay, so spiral has this interview mode where it, in order to do good writing, you have to download a lot of context from who you're writing with. And you can actually do a lot more context download agent to agent than human to agent
more quickly because my plus one has access to my whole life and spiral, we would normally have to like build integrations for all this stuff. But we don't have to do anymore because it just integrates into plus one and it just, it changes a lot. So spiral is agent native. Monologue is our speech to text app, it's like whisper flow. We're basically what we're doing, it's not really agent native, it's like, it's kind of in this like weird category where it might not necessarily even need to be. But what we're doing is creating a way where you can, whenever you activate it on your phone or your computer, whatever you say, you can direct it right to your plus one so that it doesn't go into the app that you're in, but it's sort of like having like a walkie talkie with your, with your plus one. That's cool. And I think that's going to be sick. And then sparkle is, it has, we're launching a version soon that has an internal agent. So I would say like, we'll just sparkle again, sparkle is our file organizer.
The file organizer, yeah. And then obviously proof our document editor is agent native. So I would say we're like 70% ish and we're really getting there and we are, as I said, like you have to kind of reset your product suite every three to six months if you really want to take advantage with the model models are capable of and we're definitely like in that process right now. Yeah. Do you worry about like the models becoming to the point where, you know, it's not worth it to build your own software? I mean, I know some people have that concern. I take it. You're not in that camp. I definitely am not in that camp because my, my answer is just surf the models. Every time there's a new better model, you can build a new better product right on top of it that the model can't do by itself. Yeah. And that's our job. Yeah. Awesome. Love it. And then I guess last thing is just building and shipping in the agent era, you've learned a lot with proof. We've kind of talked about a lot of things. But just what is, what is your advice, anyone who wants to build their own products with agents to use agents in this era?
So same, same advice is if you want to make sure, absolutely, that you have a job and you're thriving in AI, surf the models, your model comes out, use it, push it as far as you possibly can. And I guarantee you, there is no, just because of the way LLM's work, there is no way that that model is going to be better than you at using itself, like it, it's not trained on itself. So you're going to be finding all these different new places to push it. And if you just surf the model, that's going to be like a really, really valuable place to be. And that's what we try to do. And I think that's my number one piece of advice for anyone. I think for engineering or stuff specifically is like, we're in this world where you can build anything. And so if you are, if you're someone who like has ideas and wants to be shipping stuff and you're not, but you have to really question like, why? Because you could be. And so it's probably something, there's something there to work on. If you're someone who is building lots of things and not finishing, that's also a problem.
This stuff is, it can be actually kind of a dicking. And so that's true. That's true. That if you're using it. Browch. I can vouch for it. Yeah. Totally. If you're using it with a particular goal in mind and you're not hitting that goal, it's worth like really assessing that and evaluating whether or not it's doing the things that you want. I would not be too precious about it, just get stuff out there as much as you can and see what works. And it's just a really fun time to be building things. It is. Yeah. Do you have any particular way you like to ship stuff? Do you use a particular platform, what's any advice on the technical side? I mostly use Codex now. I use Clawed a bit. I think Clawed still has better empathy, so if I'm trying to design an API that agents have to use. So like ask Clawed, like what's the most ergonomic way to do this? I also think it's, I think Codex's design skills are a little rigid, and so if I'm trying
to do a good UI, Clawed is very helpful, especially I'm not doing it with a designer, like if it's just me and kind of riffing. That's my experience with Codex too, is that if you are a front end designer, you can probably get it to do exactly what you want, but if you're not, you're going to get a better result from Codex. Yes. Exactly. From Clawed. I also, I have to say, like check out Proof, use it with your agent, it's pretty sick, and if you want more stuff like this, check out every, every.to. We publish stuff like this all the time, and there's always new things coming out. Thank you so much, Dan. Yeah, we really appreciate it. This is awesome. Great to meet you and talk about all this stuff. Let's see here. All right. I'm just suddenly realizing we don't have a plan beyond that moment. It's totally fine. All of a sudden I was like, oh, now what?
So I guess what we can do is we can just kind of recap what we just talked about, and then a lot of wild moves this week too, where Earth may be tapping into if we wanted first. Yeah. We can do that. But first of all, I just want to say, maybe some people are familiar with Dan's work. They've already read some of the stuff we talked about, but obviously he has a ton of great resources. His Open Claw installation guide is the best, but it sounds like if you want just out of the box solution for that, that works with Slack, plus one is the tool there, and then, yeah, I just think his whole journey of proof, I think for me, the takeaway point there was definitely before you start building a project when you're in the planning stage, make sure that you have it go out in research, or you yourself research the best practices for using the tool. What I'll do is if CodeX or CloudCode comes to me and says, hey, here's the plan, we're
going to use XYZ Library, I then go and I ask it, why are you choosing this? What are the other alternatives? Why is this the best option? And if you don't know, go use web search and find out and tell me, yeah, I like the best practices. It's like a good, best practices is a good way to kind of condense all of that. You know, you might even want to provide those best practices, if there's a site you trust more than others, maybe go find what they have on those best practices and drop that in there. Could probably save you a lot of headache. Much like grant kind of my strategy is I don't want to ship it until I understand what it's doing. Yeah. You know, like I want to understand the code, at least at a rudimentary level, because, you know, I don't do this, I didn't go to, I don't have a degree in computer science or anything, I'm just kind of learning as I go. So I'll take it and drop it into chat GPT, or I'll drop it into even Cloud or Gemini. And I think it's really good to be like, just want me through this.
Tell me what we're looking at. What's the architecture here? Tell me the framework that allows this tool to work. And as I'm saying this, I'm thinking, I didn't do that with the one I just built. I'm sure really. Look, it happens. It's, you know, the coding, the coding part is trivial, right? So you can always rewrite the code. I'm always telling people to break better prompts, but the truth is when it's just me left to my own devices, I'm like, do that thing. The two things I want. That's damn it. I didn't find it opening for is if he, if and how he uses AI in which model in his writing process, because I'm very curious, you know, they're very, they're very, they're very high-taste over at every, like I think everything that they publish is really well done. They put a lot of thought into the ideas and the writing and they've slowly embraced using AI as part of their writing process, but I don't know where exactly they are with that these days. But everything that they publish is really great. And then, so that was one question I want to ask him is is like, where is he, where is
he using AI in the writing process? Sounds like he's doing a lot with, with his agents, yeah, and spiral and all that stuff. Another thing I wanted to ask him, actually, I forgot, there was another thing, if it comes up, I'll, I'll, I'll message him. Come up a little later, if it, if it returns from the memory lagoon, what was there this week? It's been crazy. We talked a little about Sora, yeah, that's really interesting. It's been hilarious to me watching the way these different companies are shipping, yeah, could be cause they're all shipping like insane right now, but they're each doing it in different ways. Claude is doing public releases that are a normal thing. It comes with a blog page. Here's, here's a release and they're going everywhere. With Chaggy B.T., it's more like a tacked on feedbacks. Yes, with Chaggy B.T., it's more of a tacked on feature.
With Gemini right now, I'm watching them like daily throw things out for Gemini CLI, for Google AI Studio, like they're all doing it. And the funny thing is, the part of this that I get such an incredible kick out of is how often it's a thing the other company already had and the people go crazy about it. And, and that goes every direction, like, like, what that tells me is that most people have a favorite and never leave it. Yeah, I think that's true. And then once yours gets the same capability, you're like, yes, this is amazing, this is a great big step. It's like, actually, the other one said that, it's kind of like when a show goes viral, when it hits a certain streaming app, like, not everyone has peacock, the show's always like, no one knew it existed yet, but then all of a sudden it hits Netflix and it's like a show from 12 years ago and it goes crazy viral just because everyone can watch it now. Yeah, I was left because like computer use, Claude did it first, Chaggy B.T. made it better with Chaggy B.T. agent.
It made a big loop forward with 5.4 and then Claude came back and made another giant loop forward this week. And every one of those four is like computer use never happened before it. It's absolutely like nobody ever heard the term before and I'm like, it's so funny. Yeah. No, but look at this. So everything Claude did in the last 52 days. This is a calendar. This is everything that they've published since what's the first day here February 2nd. And almost every day, a couple of Sundays they missed, a Thursday they missed, maybe because like, you know, some, there was like a red alert or something like opening an eye drop something. I don't know. You know, a Saturday they, or a Monday they missed, but it's just, it's just like something is going on over there where they can just build and ship things as quickly as they can. Some people think it's like they just have Claude 6 internally. I think we're watching them doing exactly what all three of them are doing because like
Gemini shipping like what 80 features a week or something, you know, small ones, but they're all features. They're all shipping so incredibly fast right now that to me, I always has it to say auto recursive, but I don't really have a better way to say it than they're really leaning on the models to build themselves. No, I think it's true. I think they have internal workflows where, you know, the agent is, is writing the code, the humans are reviewing it, the agents is reviewing it, and then they're shipping it. And you should assume every one of these companies is one to two models ahead of what is public right now, you know, Claude is, I mean, Anthropics not building with 4.6, which LGBT is not being built with 5.4, you know, I mean, they're absolutely working, or excuse me, it may be being built with it, but, you know, they're absolutely working with these models, a generation to even two generations ahead depending on the company.
And I'm sure Google is doing the same thing as you should be, you know, I'm trying to look at see what's what's happened since, because Thursday's a big day when they publish a lot of stuff, let's see what's happening here. Open Claude, the iPhone of tokens. Interesting, well, we can look at, we can look at what was published yesterday as well. The iPhone of tokens is such a weird way to say it, like, I get what he's meaning, he's meaning like, it's, it's, it's tokens having their iPhone moment kind of, yeah, exactly. Yeah, we have the Round the Horn diodesh that we published now where I, you know, look at Twitter every day, look at all my other sources that I checked before. You're doing the gaming now. I didn't even realize. Yeah, I'm doing it daily, because it's just easier for publishing based on. I'm not logging these, I need to be logging these, because they're not coming up in the right category. Yeah, but, yesterday you published a- Oh, that's not the same much I thought that was the one I published right before we went
live. Oh, no, this is Codex 101. Did you publish another one? Yeah, I published Codex 10 tips for non-coders. Oh, cool. Awesome. I wanted like normal people stuff you can do in those tools. Yeah, I think you have a good talk. I would call you Code version of that too. You published the one-on-one guide, right? It's okay. I'm just promoting you. All right. Thank you. Yeah. And, you know, I kind of distilled it down to like seven tips here. And then I clicked to the real, the, you know, the full guide. You have to plug the other one that you just wrote, which is really cool. But I really like this. I think this is a really great way of introducing people to Codex. I think the, you know, to the end point- It applies to all of those coding agents. I would say, oh, sorry. Yeah, no, I agree with you. I agree with you. And I also would say, let's switch off this for a second. I would also say that it applies to, like, using Codex for non-coding tasks.
Successful. Yeah. So. Yeah, that's fair. Like, like, I think, you know, just like Cloud Code was good for not just coders, but everyone. And then they made Codework. I think Codex is the same. I think it's a good app that's really good for coding. But it's also good for just using it with files on your computer. Yeah. I think the app is the least intimidating way to use it. I think they're, and I would say that with Clawed too. Because I just, I feel like less technical people are terrified of a command line interface. Because normally, if they interact with a command line interface, it means something is broken on their computer, and this thing is flashed up. So as opposed to, you know, it's, unless you're the, you know, the DOS generation. And I like that the app makes it a little more inviting, you know, there's still some coding lingo there, but it really doesn't matter. You could absolutely ignore much of that and just go in and use the automations and skills
and, and, and pick your models and so forth. Yeah. Really impressive. But yeah, this is a great, this is a great resource for people. It also pulls from the official Codex docs and, and other stuff. So it's like a great starting point if you've never used these tools before. And you want to do what Dan did, which is build an app and produce it and hopefully, you know, not, not, uh, skip the best practices and, you know, not three for seven days, but, yeah. But it's a nice of stress and sheer terror. Yeah. Yeah. Oh. Yeah. But yeah, it's a great resource that you published. Um, what else happened yesterday? Yesterday. Uh, the art prize. We were, oh, this was the other thing I wanted to ask Dan about, um, was benchmarks because you and I were having a discussion right before we got on here about the value of benchmarks and whether or not they're useful. And we can save everyone the back and forth between you and I, but I think your ultimate conclusion makes sense to me, which is benchmarks are toast, like they're cooked.
And the only benchmark that really matters is a list of tasks that the AI can accomplish. And you just list every human task you can think of and just, and mark off. And I don't care what you connected to to get there. I don't care if you're doing that in Cloud Code. I don't care if you're doing it through an API. But I think if there's a thing it can accomplish that is really good, we should see that. And honestly, I think a lot of research labs don't know what those tasks are in some cases until it gets out and it's in the hands of a really gigantic, diverse set of people. But like I, I would love to see, like, I need to see it. Like, you know, if you're going to test things like creativity, for example, I don't want to know a 46 versus a 51, I want to see what it wrote versus what it wrote. And a lot of these are very closed, you know, I think, because what's going to affect me and not just me, because like, you know, I tend to think of you and I as normal people. And then I remember that most people don't read as much of this as you and I do.
The fact is, you're probably a little buried. And I feel like I'm behind, but then I compare what I know to the average person. Oh, my goodness, people are not prepared for what it's getting bigger and it's getting big faster. And since, since I'm going to go back and say, Gemini 25 Pro from Gemini 25 Pro forward, like there have been these big hops from model to model, from tool to tool as they've come. And the hop keeps getting just a little bigger and a little bigger. And I just don't see a, I don't know, I don't know how you catch up today. Well, unless you're stuck by the neuron, not AI or newsletters, enjoy our fine podcast. If you do those things, you are sure to be well informed. Yeah. So for anyone who's watching, I don't even know how many people are still on the stream
at this point. If anyone is watching this, this is the Arc AGI 3, 10th. And basically, it's a series of video games that test your ability to reason and adapt to a new situation, figure out the rules of the situation, and then basically try to solve this puzzle. And I have failed this one three times because I missed a key thing there. Now I got it, and I'm going to try, this is kind of hard, but yeah, go ahead. We're only on LinkedIn and X. Apparently around 12 o'clock, we lost YouTube. Okay. Well, that's, it is what it is. It's why the chat is so quiet. Yeah, that makes sense. Yeah, because I was like, really? Nobody has any feedback? No, that makes sense. So we can wrap this up here in a minute. How about at the half hour mark?
I think so. I think that's good. I think it's good. This game looks fun though. Great. Yeah, this is Arc AGI test. So it's meant to test how well an agent can adapt to a new situation. And we wrote about this in today's neuron, but basically what happened was every frontier model tried this and was at like less than 1% ability to complete it. And not agents. Specifically models. Specifically models. Specifically models. And everything else. Yeah, because they're trying to test the actual underlying model and how good it is. And perhaps that's not really a fair assessment if the way that we're actually going to be using these things in real life is with the harness and as a part of a system. So I sort of agree with Cory's point there. If you want to expand on that, you can't. You know, part of my thing is that. I mean, I do think they're good metrics to know, but I don't know that any value comes from knowing this one's a three and this one's a 13 other than, yay, my team's winning.
Like, I feel like that's what we get out of it. And what I think is necessary is a much more practical approach. And I think instead of, you know, playing benchmark whack a mole, maybe we could move into more of a, you know, task specific stuff. And I don't mean your numbers go on tasks. I mean, show us examples. Show us. You know, maybe it's, you know, here's a 200 page, you know, research PDF that gets shared out, that's showing us like, all right, now here's what their last model did or here's how it works. I actually think we're so past. We're so past the point, like, like, I'm going to give OpenAI a bit of grief tomorrow about their ads and chat GPT. I'm going to give OpenAI grief tomorrow because their ads and chat GPT are so generic. I'm like, we have generative AI. Like, you could make generative UI at this point. And you're going to give us a little tiny image and add, like, I get it. Like, we don't want the ads to be obtrusive, but at the same time, like, we're past the point
where you should be putting out PDF documents, like, like, you should be able to build an entire website with videos embedded with all of the tasks, like, showing them. That's like, come on. We're way past the point of, like, PDF research reports. Like, you got a really good call, great. Entire websites for this stuff. Yeah. Yeah, I want living research papers that I can ask questions to when I'm unclear. Yes. I want all of the things. Yeah. Yeah. There was a great post I saw yesterday. I think I included it in around the horn digest. But it's someone who's saying there's two different types of websites that will exist in the future. And I agree with this. The first one is websites that make it really, really easy for agents to read them. Like, super easy. Like, like, basically, like, they're designed for agents. That's where you're at. Second kind of building is coming from. Yes. The second is designing for humans and making them as visually stimulating and as interesting and as complicated as possible. So in one world, you make the website as simple as possible.
And in another world, you make it as complicated and dynamic as possible. And I think that's where generative UI comes in. I think that's where you're going to have websites that feel dynamic and alive. And like, you're playing a video game, but you're on a website. Like, I just think that, like, that's where it needs to go. Because we have the tools to do that stuff so much easier now. So, like, now the level of complexity needs to go up. And really, just, like, meet people where they are. Like, yeah, if I'm going to read your website, you know, make it interesting. Make it cool. I can't stress enough that meet people where they are as important. And I think that's what's what's sourd me a little on benchmarks. You know, and I think it's important that we begin to try to make this stuff more accessible and explain to normal people what it means. I think it's important that, you know, more people than ever are, unfortunately, picking their sides in a battle as opposed to trying to figure out what this means. It's, you know, it's becoming very political, which I hate.
But I think it needs to happen, personally. The going political? Yeah, I do. Because I think, you know, at this point, there needs to be some sort of backlash to the progress. I think, like, the progress is happening so quickly that if there's not a pendulum swing the other way, then who knows what happens. And I think the natural order of things is for the pendulum to swing. Yeah. So I think that's healthy. Yeah. I think it's probably healthy. I'm not saying I agree with the form that that backlash will take. Let me put it that way. And what cost is my concern? Like, I mean, you know, and their argument would be at what cost are we going forward? And mine is, I just have a genuine feeling that whatever country leads the way in this, is going to be leading the world for a long time. That's fair. And that's the thing that has made me, I won't say anti-regulation. I do believe there are needs for regulation.
But I would rather see them around things like deepfakes, child safety, as opposed to pausing and preventing. I think coming around and kind of cleaning up the mess, perhaps you could legislate something like, I don't know, affordable training methods. Ways people can learn and reskill quickly. New ways to, you know, start revisiting income and what that's going to mean in a later world. I think, like, I'm absolutely for certain types of regulation. I just, I have this fear that the only two options are foot on the floor or slam the brakes. And because that's how we do everything in America now, it's everybody's just extreme. It's like everybody slow down, let's live in the gray a little here, let's find it. Yeah, find me. Well, that's why I want the pendulum to swing, so it comes back to the middle. Hopefully come back to the middle, yeah. Just, you know, maybe it'll look like, you know, a pendulum over a skyscraper,
or it takes out a couple, you know, takes out a couple rungs on a rope. If you think of it like a pendulum over a rope in a root goldberg machine, you know, where it's like slowly, it's like cutting away at the rope as it goes lower and lower. Yeah. Yeah. Oh, what was that? Who said? Oh, that, that takes me very Edgar Allan Poe. Yeah, exactly. I know exactly. Merges in the room. Maybe. The darkest hour. I forget what it is. But yeah, one of those. But I remember the guy laying on the floor. And the wait for him to just get closer and closer. Yeah. Yeah, I think also, I think it's just this capability unlock of where we're going next is going to necessitate we redesign our society in some ways. Yeah. I think it just has to. Yeah. And so it'll be interesting what form that takes. Hopefully the form is not centralizing power. Hopefully it's the centralizing power. As divided as we are right now as a country, as a human.
Maybe that fresh starts what we need. I don't know. But I would. I think it's going to necessitate it. Yeah. Yeah, I think it's going to lead to. And there's some interesting ideas. Like there was a great Axios article yesterday that I'm going to basically include in tomorrow's newsletter about Centraini Research published some ideas for, I think Gina Raimundo who's in the government, like the Department of Commerce or something, she had some ideas. And there's a lot of ideas for ways that you can build out the safety net to incentivize using humans and keeping people relevant to the workforce as these tools and capabilities increase. So there's some really cool ideas around that. That don't look like foot on the brakes do nothing. Yeah. Or Bernie Sanders is like, pause all day to center. I love that he goes and talks to Haiku out loud. Oh, yeah, that was funny. That was funny. The best meme of that is Old Man yells at Claude.
I was like, no, you send him to the one that's trained to be unsure about its consciousness. That's terrible. That is not the one you. Well, it's just funny because in the video you can see him totally getting Claude pills where his whole tone of voice changes. He's like, oh, it's somewhat reasonable. I feel like he was prepared to fight the AI or something. Yeah. He actually found it, you know. I actually like you better than most humans. Yeah. Yeah. Wow. I asked you a question. You answered me honestly. What? Without yelling? Nobody's screaming? No names? Yeah. Oh, goodness. Got anything else, Corey? I think that's it. I think it's it for today. Everyone, thank you so much for joining us. If you haven't yet, please remember popping the chat. Go find a link to the giveaway. You got your last shot to get in and get a chance at a DGX Spark. And you should because it's pretty sick. Also, make sure you check out video we dropped last night with Nick Heiner from Serge. It's really cool. We dropped three different ones last week.
Nicole Karte from Nicole Bear from Karta. We dropped Dr. Chichahu from SCS AI. We dropped Carrie Briske from NVIDIA who is an amazing person. Eman Boulire from Proton. Yeah. Yes. Eman Gouare from Proton. Yeah. You know, lots of these. You should go watch something. There's great stuff. And we appreciate you being here. We appreciate your continued patronage to our fine publication. And I don't know. Just be an nerd. I love it. Have a great week, everyone. And we'll see you back next time. Farewell for now, humans.
More episodes
More from The Neuron: AI Explained

BONUS: The AI Starter Kit: What to Try...and What to Ignore
The Neuron: AI Explained

Can AI Really Design New Drugs? Google DeepMind Spin-out Isomorphic Labs Explain...
The Neuron: AI Explained

BONUS: OpenAI Workspace Agents 101: Build, Run, and Scale AI Workflows
The Neuron: AI Explained

How Google's New AI Turns Anyone Into a Music Producer (Flow Music Demo)
The Neuron: AI Explained