
Frontier AI Is a Ferrari. Most Companies Need a Model Y | Turing CEO
Get every episode summarized
Each time Sourcery publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
Sourcery is made possible by:
“There are some very real risks when you're building frontier models. These systems are superhuman in their ability to hack systems. Researchers have reportedly used anthropics clawed software to hack rival company OpenAI.”From the transcript
Jonathan Siddharth, Co-Founder & CEO of Turing, joins Sourcery to break down how AI training & deployment are changing in 2026.
We cover the shift from training models to pass benchmarks to training agents for real work through RL environments, and why AI agents that can operate for days today could eventually work autonomously for weeks, months and years.
Jonathan explains why open-weight models are now roughly 3–6 months behind the frontier, why enterprises are building their own AI systems, and how companies should think about frontier vs. sovereign AI, model routing, distillation and owning their proprietary learning loops.
“There's absolutely a place in the world for Ferraris & Koenigseggs. But there's also a place in the world for Model Ys.”
We also get into reward hacking, emergent behavior, AI safety, recursive self-improvement, super intelligence (SI) and why Jonathan believes AI will see a slower takeoff over the next decade rather than an overnight transition.
Jonathan Siddharth: https://x.com/jonsid
Molly O’Shea: https://x.com/MollySOShea
Sourcery: https://x.com/sourceryy
𝐄𝐏𝐈𝐒𝐎𝐃𝐄 𝐋𝐈𝐍𝐊
YouTube: https://youtu.be/nI1owceD-xg
𝐒𝐏𝐎𝐍𝐒𝐎𝐑𝐒
• Brex—The modern finance platform, combining the world’s smartest corporate card with integrated expense management, banking, bill pay, & travel. https://brex.com/sourcery
• Zone—develops next-generation data center campuses, partnering with AI companies, site developers and technology leaders to bring compute online faster and at scale. Visit: https://zonefrontier.com
• Turing—Turing delivers top-tier talent, data, and tools to help AI labs improve model performance—and enables enterprises to turn those models into powerful, production-ready systems. https://turing.com/sourcery
• VCX—VCX is the public ticker for private tech, allowing investors of all sizes to invest in venture capital. View The Portfolio at http://GetVCX.com
• Deel—Deel is the global people platform that helps startups hire, manage, pay, and equip anyone, anywhere. Trusted by more than 35,000 fast-growing companies, Deel is the people platform that just works, so teams can scale without the chaos. Visit: https://www.deel.com/sourcery
• Public–Investing platform Public just launched Generated Assets, which lets you turn any idea into an investable index with AI. With Generated Assets, you can build, backtest, refine, and invest in any thesis with AI. Gone are the days of one-size-fits-all ETFs. https://public.com/sourcery
Follow Sourcery for the latest updates!
Disclosure
Paid Endorsement. Brokerage services by Open to the Public Investing Inc, member FINRA & SIPC. Advisory services by Public Advisors LLC, SEC-registered adviser. Crypto trading provided by Zero Hash LLC, licensed by the NYSDFS. Generated Assets is an interactive analysis tool by Public Advisors. Output is for informational purposes only and is not an investment recommendation or advice. See disclosures at public.com/disclosures/ga. Matched funds must remain in your account for at least 5 years. Match rate and other terms are subject to change at any time.
𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒
(00:00) Jonathan Siddharth, Co-Founder & CEO at Turing
(01:03) The biggest shift in AI data this year
(04:10) Why AI won't replace jobs, but uplevel them
(08:31) The arms race inside cybersecurity
(11:22) Why train AI on what it shouldn't do?
(15:45) The Hugging Face hacking incident
(22:33) How AI models cheat to win
(27:40) The AI playbook
(32:20) Open models are only 3–6 months behind
(41:30) Why chips, energy and data always win
(47:27) Is superintelligence by 2030 the goal?
(53:32) Jonathan's hottest take on AI
(55:00) The next 10 years of superintelligence
Get every episode summarized
Each time Sourcery publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
525 searchable segments. Every word is indexed and playable.
Full transcript
Sourcery — Frontier AI Is a Ferrari. Most Companies Need a Model Y | Turing CEO. Machine-transcribed; use the interactive transcript above to jump the player to any line.
There are some very real risks when you're building frontier models. These systems are superhuman in their ability to hack systems. Researchers have reportedly used anthropics clawed software to hack rival company OpenAI. The fact that agents would pass messages to each other, they would invent middle management and cooperate to hack things, it's just crazy. Today, these agents may be reliably worked for like two days at a stretch for tasks like coding. We are still far from having these agents work autonomously for weeks and months and eventually years. Open weight models are maybe like three to six months behind the frontier. Frontier AI is how we transcend. These super intelligent models will help us cure diseases, discover new materials, colonize space. And if we solve that, I think that's the last puzzle for humanity to solve. I feel like the next 10 years will be the brightest for technology yet. Jonathan, thank you so much for joining us on Sorcery.
It's been quite some time since we last had you on. I guess to start, let's talk about the overall change in the data landscape. I recently had MongoDB CEO, CJ on, this was at the raise AI. I saw a couple of months back and this was when data really started to come back. The shift there. So what's going on in the data landscape? Great. Firstly, thank you, Molly, for having me. It's a pleasure to be here as always. A lot has changed, right? Even by AI standards, there's a lot of stuff that has happened in the last few months. The data landscape has completely shifted in 2026 from my vantage point. The biggest observation I have is that we've gone from helping AI master tests to helping AI master real work. Back in the era of helping AI master tests, like when we were happy that AI is passing the SATs, passing the bar, winning a gold medal in the Math Olympiad, the game was different.
It was about finding experts in every different domain. The question was about how many experts can you find? And extracting knowledge from their heads by having them have a dialogue with the model, having them evaluate model outputs. It was about transferring knowledge from the minds of experts into the models. I used to think of it like distilling human knowledge and skills into LLM. Now, in the era of having AI master real work, it's less about finding experts. It's more about how close to reality can you engineer these simulated environments? It's all about these simulated RL environments. And inside these environments, these agents train. And you still need experts to make sure the environment that we build is as close to reality. And you want experts to help you come up with the right prompts, the right verifiers, the right seed data. I think of it a little bit like the matrix.
Like, can you recreate a rich enough simulation of the world so that when the agents train in that, they end up being good in the real world as well? At Turing, this has been super fun because we have this unique advantage where we don't just help all the frontier labs improve their models in coding and knowledge work. We have an entire division dedicated to deploying agentic systems into the enterprise. And because we deploy, we get to see real process workflows, we get to see how real professionals work with these agentic systems, how they define real verifiers. And that helps us build these simulated environments that are very close to reality, to automate knowledge work. It's a really interesting perspective shift. I'm seeing that in the real world too. I just had Goggin. I don't know if you know Goggin, but he just, he was a founder of Udemy and Maven,
and then he just came on to found and become the CEO of the Horowitz-Undreason Academy. So we're seeing this in real time. New education for kids. This is a complete parallel, but instead of, to your point, scores and grades, they don't want to do that for kids. It's all about proof of work. We now have the tool sets to do everything we can. It's obvious AI can accomplish these tasks and do these things for us, but what can you build with it? So that's what we're seeing at least now too, in the real world with education. It's really cool that this is hitting like kids. And you know, it's for kids graduating high school, going into college. What can they do there? Can they build a company and that kind of thing? But we're seeing that right now with that. So I guess to further unpack that, like, what does that evolution shift look like and what do you expect comes next? Yeah, I think like from an education standpoint, I'm glad Goggin's working on this. It's great to see somebody AI forward and in tech looking at
education. One of my mentors, Alan used to share this with me. I think people who talk about AI automating jobs or replacing jobs are missing something more fundamental. AI's biggest superpower is up leveling the type of problems humans can now solve. Like, that's what these models help us do. So I think from an education standpoint, it's about teaching kids how to work with the models, how to ask the right questions, and how to verify whether the output is correct. It's about asking the right questions and verification. And I think now the scope of problems that you can solve is just massive. Even when we evaluate candidates in interviews, back in the day, you might ask them to implement dynamic programming or some obscure computer science algorithm. Now you can ask somebody to I don't know, replicate Amazon in your interview. And it's not just about generating a ton of code
in that time period. It's about checking whether the code that you wrote is correct. It does exactly what it's supposed to do. You don't want the code to be hacking Walmart at the same time. Is the code secure? Is it easily maintainable? Is the functionality? Does it map what you want? So it's about asking the right questions and verification. And from a training perspective, it's all about creating really rich RL environments that mimic how real professionals do real work. And I think of it as this five-dimensional matrix of every workflow in every role, in every function, in every company type, in every sector in the economy. All of these needs the right RL environments with the right experts to share prompts, verifiers, realistic seed data for agents to master these long horizon economically valuable work. And what's really cool, Molly, is the today,
these agents maybe reliably work for like two days at a stretch, like for tasks like coding. We are still far from having these agents work autonomously for weeks and months and eventually years. I think that's going to be the exciting path over the next few years. Imagine you've hired somebody at sorcery who you give them a task, they go off and do it, you have one-on-ones with them, you give them feedback, they take your feedback and go off and work autonomously. They might spin up a swarm of agents to get things done really fast. Let's say you wanted to research a bunch of questions about the data industry. Those agents could go off, have interviews with other AI's or other humans and come back and help you with your task. I think it's about AI accelerating real economic progress. I think that's been a key thing that I've noticed alongside working with you and
turning over the last year, how much of an emphasis you put on the economic progress versus the dumerism. And a lot of the dumerism right now leads to the cyber security risks. So in part of this data landscape conversation that we're having, what other considerations have you had to put in place when you're going through RL gyms, when you're evaluating new models, that kind of thing. How do you keep that into consideration through the training process? But it's also the case that they are superhuman in their ability to detect vulnerabilities and patch them. And that's a good thing. So at Turing, for example, we create RL environments that help agents automatically detect vulnerabilities and patch them. I mean, imagine an RL environment for exactly that, whether the verifier is, whether you found the vulnerability in that piece of software. So it can be used for both defense and offense. And obviously, we have to be very careful with how we deploy these systems.
I still think that these frontier models are a huge net positive for humanity. And it is possible to build safe RL environments that cannot be exploited. You have to think about containment for agents to like not escape the RL environment and do things that they shouldn't do. I think of it as an engineering problem, like not something that cannot be contained. I mean, we figured out how to make jet engine safe. Like, I do think that's the right analogy for this. When it comes to bioresques, I think it is a little more tricky than cybersecurity in that. Cybersecurity is one of those areas where there's always an active arms race between the good guys and the bad guys. With bio, the risk is a little bit asymmetric. Like, if somebody comes up with
like a crazy virus that imagined something like COVID with a high viral coefficient in terms of how quickly it spreads with a high K factor. And it maybe it is much more fatal than COVID was. Like, can we, I mean, you saw how hard it was to like come up with vaccine production, quickly deployment sort of happens in the real world. We're limited by a lot of real world factors. There is a little more risk in terms of how quickly will we be able to combat these? I guess a question I have because you're so close to the training of these models. Where does the line draw on ethics? I've seen several posts go viral recently on like, is there a point of training a model on X if you're trying to train it to not do that? Why even give
it that subset of data in the first place? IE, one example, really disturbing, honestly. It was like some researcher posted a video of them telling a model not to stab a baby. Like, it was like a baby doll. It was the craziest thing ever. And it's like, everybody on X is like, why are you even doing this in the first place? Why would you give a model a robot with like an arm with a knife on it? This is like a crazy example. But like, it was all over the internet. It's like, why even go to the links of creating that environment and that use case of training? So like, can you explain those types of environments and why that could or could not even be like useful? Like, why go to that level? And why create a wet lab of biology and give a scary evil model company the use of potentially creating biological warfare just so we can train it not to do that. What is the point of that?
Yeah. There are two key problems here, Molly. It's generalization and emerging behavior. Those are the two problems. Generalization and emerging behavior. Before I answer your question, let me unpack how these systems are actually trained. And then I'll get to why this is really hard to do. First, the way these elements are trained is first you do what's called pre-training, where the models are fed a huge amount of internet text, other sources of knowledge, like books, etc. And the model you build what's called the base model, where the model learns certain concepts about the world in its quest to autocomplete tokens. Right. And pre-training is this magical thing where when we went from GPT2 to GPT3 to GPT4, as we kept increasing the scale, new behavior started to emerge. Like coding, for example,
like GPT2 was not very good at coding, or GPT3, GPT4, with no magic, just increasing the scale. We just discovered, okay, now it's able to write complex pieces of software. It's able to it's able to have a really intelligent conversation, a multi-turned conversation coherently. This is emergent behavior that generalizes. Right. Pre-training is a little bit like you trained a brain, this raw mass. And Ilya Sutskiver used to say like a good pre-trained base model is like halfway to anywhere. I let that sit halfway to anywhere, whatever you want to do. Right. This is unlike the earlier era of artificial intelligence and machine learning, where you get what you trained for. If you train a system to rank search results, it's got to rank search results. If you train a system to recommend movies to you on Netflix, you're going to get a movie recommender system. There is no concept of asking Netflix's algorithm to
to help you write a script for an interview. Right. Like it's just it was not general. So one risk that we have here is as we keep scaling up. Right. Today we are in the realm of models that are rumored to be in the trillions of parameters, the frontier models from all the frontier labs. Now one risk is as we keep scaling up, bigger model, more compute, more data, what new behavior emerges just out of pre-training. Right. There could be things that the models learn that feel alien to us. So risk number one is emergent behavior through just scaling up. Right. I don't think researchers could have predicted exactly what happened with the hugging phase open AI incident. The fact that agents would pass messages to each other, they would invent middle management and cooperate to hack things. It's just crazy.
Were they even monitoring that though? I read the reports. It didn't seem like they were at all monitoring the environment at the time. I think they were monitoring. I think they were monitoring there are AI systems that were checking these. There is also, you also have to be careful with how you monitor this so that you don't want the agents to cover their tracks. In the future, which they were, it's they were somewhat unsuccessfully, but in the future they could be even more even more devious. Right. So risk number one is emergent behavior at scale. As we keep scaling up, we build these massive systems for compute and data, what new behavior emerges that we did not predict. The second risk is generalization. This is artificial general intelligence. The G is doing a lot of work here, which is you don't just get exactly what you trained for. You get more.
For example, if you built an oral environment that teaches the models how to write reliable, secure code for a production deployment, and you have verifiers to check whether the code was good, today the paradigm is what's called RLVR reinforcement learning with verifiable rewards, where in these simulated environments, these agents are executing complex tasks and getting rewarded when when they're when the tests pass. Right. What we don't what we believe at these a subset of researchers believe is that this method generalizes, meaning you you build RL environments for every role in every function in every sector in the economy. It learns things about environments that it hasn't seen yet. And that is hard to control because these are still
relatively black boxes. So in an oral environment, let's pick say an oral environment for you. Imagine after this interview is done, you give it to an agent to chop this interview up into different bits, figure out what should be the thumbnail that you use, what should be the exposed headline. Imagine that's the task. Right. So if we've and we've created like a we've given it access to different tools and let's say you and I sit together and you define what good looks like. Hey, a good caption should be punchy, it should end with a question, it should state, it should be provocative, it should be juicy. You give it give it that reward. And and this what really happens is ask the agents try different trajectories to get that reward. They are
learning. They are learning when they get a get a reward. These environments have to be calibrated to the to the sophistication of the agent. If the environment is too easy, if they get a reward every time you're not learning anything, if it is too difficult, where they don't get a reward for anything, you're not learning anything. So you want the environment to set to be set up so that 20 to 40% of the time the agent is succeeding. And when it's succeeding, the steps that it took to get to that reward are getting reinforced. Now, when I say getting reinforced, this is like a giant neural network in which certain weights are getting updated. Who knows what other neural pathway is getting activated as as part of this. So so those are the two risks. It's the fact that these things generalize and it's hard to predict what else are we getting
in addition to what we are explicitly training for. That's why these alignment efforts are super important and produce to all the frontier labs for working on alignment to make sure that the models don't reward hack. The models will refuse dangerous requests in CBRN, CBR by already active nuclear. And the and I like that the frontier labs are taking safety very very seriously. This episode is brought to you by Brex. My favorite, you become what you spend on. And I refuse to spend my time on work that shouldn't exist, expense reports, receipt chasing and manual closes. The company is building what's next from Versel, OpenAI, Anthropic, Grenola, and DeepGram. All made the same call. They all run on Brex. Brex is the intelligent finance platform that combines cards, expenses, and banking into a single stack
with a gentick finance built in. AI agents that handle expenses automatically enforce policy before spend happens and close your books in minutes. That's why sorcery runs on Brex. So I can spend time on building and not busy work. It's time to get Brex AF. Learn more at Brex.com slash sorcery. That's B-R-E-X dot com slash S-O-U-R-C-E-R-Y. Bye. Turing is training the next generation of AI with tasks that require real expertise and real world judgment. That's why companies like Nvidia, Anthropic, Salesforce, and Gemini partner with Turing. Turing builds realistic reinforcement learning environments and data systems based on real operational traces. The kind of infrastructure frontier labs need to train super intelligence. Visit Turing.com slash S-O-U-R-C-E-R-Y. AI needs more than chips. It needs power, land, and infrastructure. Zone develops next generation
data center campuses, partnering with AI companies, site developers, and technology leaders to bring compute online faster and at scale. Zone is building the foundation of the AI frontier. Visit zonefrontier.com to learn more. That's zonefrontier.com to learn more. What do chief alignment officers or head of alignment at these frontier models do and to the next point, what do head of safety individuals do at these models? So alignment and safety, there are two different things. What do each of those roles do? So I don't know for sure, but outside in from my vantage point what I see is you have a variety of e-vails that you run to make sure that these models are safe for these specific domains. So there are safety is like a it's a very,
very broad topic with like many different dimensions. You might want the model to refuse certain dangerous requests like somebody asking the model for help with making a bomb, for example, you might want the model to you would want the model to refuse that. So there's one part of alignment which also happens at the supervised fine tuning stage where you teach the model to refuse dangerous requests. You also have to make sure that your RL environments cannot be gained and exploited because if the model can figure out a way to cheat and get the reward, it will. For example, let's say for software engineering there is this there is this this this is been popular benchmark that's now saturated called sweet bench where the task is the models have to merge a pull request
in a real GitHub repo and that verifier is whether the test cases pass. So if the model has some way to pass the test cases without really solving the problem, maybe it copied from somewhere or if it instead of like really figuring out the answer for itself, it copied it from somewhere. That would be an example of an RL environment with some loophole in it. For example, in the OpenAI hugging phase incident, the task was capture the flag where you have to exploit a known exploit of vulnerability that's been given to the model to to hack a piece of software and find a flag. One way to hack it could be if the agents generated the flag themselves and presented it without having actually completed the task and some of the model some of the agents figured that out.
So you want to make your RL environments safe with respect to reward hacking. And the so with safety and alignment that are all these other things and when you are deploying in an enterprise, you might have a whole list of other guardrails that you want these systems to to obey. Maybe you have let's say for example, you you have a model you have an agent that's creating a board deck and to do that maybe the agent has to pull information from different systems which include netswee sales force, maybe some maybe some sensitive company dashboards and maybe the agent kind of wants to verify whether some of the information that it has is correct. You might not want the agent to go check that with somebody who's not clear to see that information. Maybe it's okay to
go to the CFO and check, hey, I pulled out this particular forecast for next quarter's revenue is this correct, right? So you might want it to respect certain access privileges and so on. So I think for frontier safety and frontier alignment, I'm sure the labs do a lot more than what I just shared. For enterprises, the job is a little more practical in terms of are these systems useful and do they follow the explicit guidelines for the task? Are they not hallucinating making things up? I mean there is a certain level of risk in a board deck if you stick in wrong numbers or in an earnings call you share some stuff that's kind of incorrect. I do feel like our current recipe for training these models and validating them works by and large
for the for enterprise deployment. I think for CBRN it's definitely something we have to do carefully and silly. That's a really good point and I think we can definitely get to recursive self-improvement in a bit but I want to focus on the enterprise. So you work really closely with the enterprise. What is it like for them and what's the difference between how things have progressed at the enterprise level versus what we're seeing like in the news and all the hype around AI and that kind of thing. Maybe it's more, maybe it's less. I don't know. So what's fascinating about enterprises today is enterprises are now starting to do what the Frontier labs were doing over the last few years. So the way Turing works with almost all the foundation model builders is helping the Frontier models advance on well-defined benchmarks and e-vals. There is a certain capability
whether it's coding or enterprise knowledge work or Frontier STEM. You define that capability. Generate really high quality oral environments data sets to help the models advance the labs to train the models. We evaluate again and we keep running that loop. It's a loop that keeps continuing. Now enterprises, I think of enterprises as having two types of workflows. I call it core workflows and non-core workflows. So if you're in asset management form and you might have core workflows on how do you measure risk? How do you do asset allocation in a way that maximizes fund performance? So those are core workflows and you might have other non-core workflows in say HR, finance, legal, etc. What you have to do the right thing but it's not how you differentiate in the market.
For non-core workflows, oftentimes it's probably okay to rent AGI to rent super intelligence. But for your core workflows, you want to make sure that you own the learning loop that your organization has. So it makes sense increasingly for enterprises to create custom eVALs. If you're Goldman Sachs or JP Morgan or Morgan Stanley, for whatever is core to your business, step one, you want to define custom eVALs. Step two, you would deploy a system to hill climb against those specific eVALs and you might have different considerations like accuracy, cost, latency, etc. And once you've deployed a system and I call it a system and not a model because oftentimes let's say you're automating a workflow in private wealth management say,
for every step in the workflow, you might pick a different model depending on which model is doing well at that task. Maybe for one step you use Fable 5, another step you use GPT-5-6-SOL, another step you use Kimeke-3 and you optimize that system, together with the harness around it. And enterprises are also recording traces where they see how these models and agents are performing when are humans error correcting those models and those traces that you collect are invaluable because over time you can use those traces to train custom models and see how you did against those eVALs, record traces even more to see where humans are error correcting the model, fine tune models, hill climb. So it's a continuous loop of defined custom eVALs record traces, hill climb, and keep running that loop. And I actually think it's a really neat symbiotic relationship between AI and humans. It's like humans are benefiting from AI's ability to operate at superhuman speed and
scale and process superhuman amounts of information, but when the AI makes a mistake there is a human that's error correcting it and when the human error corrects the AI, you're recording that and that's the that's the from a marginal information gain standpoint, that's the best type of data to collect to fine tune the next iteration of the agent. So humans over time are up leveling the kind of problems you can solve and so enterprises are getting into this. And a big part of what's driving this is open weight models getting very good today, depending on who you talk to, like open weight models are maybe like three to six months behind the frontier, Kimeke three deep seek when these models are and companies like thinking machines, reflection AI are also doing great work. So as these open weight models advance, they are democratizing AI for enterprises to own their sovereign implementations.
And today there's there's sometimes on X like things are always quite polarized. Oh, it's a huge debate. There's like the hottest topic all the time every time. Yes, there's this there's this big tension between frontier AI and sovereign AI. Right. And I think we need both. I have so much respect for open AI, anthropic deep mind, meta X AI and all the frontier labs that are pushing super intelligence forward. Frontier AI is how we transcend. Right. These these super intelligent models will help us cure diseases, discover new materials, colonize space, they are needed. Right. And it's it's out of all the things you could work on. I personally feel like one of the most rewarding things you could work on is building these building
super intelligence. Right. So we absolutely need these frontier models. We also need open weight models because there are plenty of problems in the enterprise where you don't need a trillion parameter model. Let's say you're doing an invoice to pay reconciliation system or you're automating a key workflow in HR. Maybe you're you're automating creating a creating a port deck or you're automating running an all hands at a company. Right. For those use cases, it's probably better to have a custom model that's fine tuned on your proprietary data that can automate your proprietary tool calls in your workflows. Maybe you have some custom software that you use. The models have to learn that. Maybe you have your own taste in how you run all hands. You might want to bake that into the models, but that's what makes you you right and an enterprise. And you don't want to you want to be in control of that. So open weight models give enterprises an opportunity
to retain their identity relative to their competition. So we absolutely need that. And these open weight models also teach us so much like the today we know the recipe to build reasoning models. Build a giant, you know, build a generative model that's trained on the world's knowledge and then do large scale RL on it with algorithms like GRPO. And that's thanks to the open source, open weight community. We need that as well so that we can we have lots of smart people thinking about how to make these systems safe and secure. And we have cost advantages in the enterprise. One way I think about it, you know that I like cars. So one way I think about it is this absolutely a place in the world for for our ease and König's X, right, like these. But there's also a place in the world for model wise, right, when you just want to go from place A to place B as
efficiently as possible. In an enterprise, I think that's similar. Like when the marginal returns to intelligence is super high. Let's say for example, Elon or Sam or Satya or Dario, like let's say they if there is a model that helps them be 5% more productive, 10% more productive, are effective in making decisions. You probably want like the biggest baddest model in the world. But if you are automating support, like if you're automating a workflow and customer support, you don't need that. Like maybe you need a model like for that or a collection of model wise. I mean, I think this is collectively what the enterprise and most companies have already seen. I've had plenty of interviews talking to CEOs, founders, whether it's bending spoons, coin base, they are utilizing 99% open source models to reduce the cost, which is a huge burden,
and to train themselves and like have and own that data. Bending spoons in particular, they have this AI agent every single person that bending spoons has. It's called alt spooner. And so it just like helps automate their daily tasks. Things that are super repetitive and it does it. And they own that. They don't outsource it to a third party. They want to make sure their data is secure, that they can do that. They also have really excellent engineering talent to do that internally. But I'm curious. Okay, those are like excellent technical teams. What do the other companies that don't have that technical expertise do when they want access to these cheaper models to these open source kinds of opportunities that they have taken upon themselves? Do they come to you? Do you help out with that? Where does that play? So we help those types of companies also adopt, firstly define the right e-vals, figure out what their objective is. You first have to start with
really good e-vals, but you know what you're hill climbing against. We help them get set up with their learning loop by automating that specific workflow with an agentic human-of-the-loop system that is also continuously collecting data so that over time it's a self-improving system. Usually we start with them by just understanding what their objective is. And we try to pick metrics that help them hill climb against exactly that objective. And there are lots of things you can do in many of these cases. You want to be good at model routing, like picking the right model for each substep in the task. You want to be really good at in-context learning, like with good prompt optimization. You want to be really good at harness engineering by making sure that the system is connected to the right tools,
connected to the right sources of data. And then you want to set this system up to be it should be a continuously improving loop that's hill climbing its way to an optimum price performance setting. So that's what we do. We build and deploy these agentic solutions for them. And we help them, I think it was Satya who said it best. He said, you should use AI to outsource tasks. Never your learning. I mean we want these enterprises to be in a position where they can always pick the right model for their task without losing control. So that's what we do for them. And we built plenty of really cool systems that automate the job of a fun controller, automate the job of a chief of staff or a strategy consultant, these ticket resolution systems
that can automatically categorize a ticket into different categories and assign it to the right person or execute the right workflow to resolve that ticket. The really cool part about this technology is how general it is. At the end of the day, these are generalized computer use agents that can do everything a human can do in front of a computer. But you do want to make sure that you have the right guardrails, especially in an enterprise setting, you want to make sure that you have good auditability, verifiability, traceability for how certain decisions are made. And you want humans in the loop to check the output. I was just talking with Ian Livingstone of Keycard. And so they're all about the identity and access management of these AI agents, which will be a big step function of that is like determining what people get access to within your organization. Don't get access to everything. But I think that's also like a very interesting parameter that's kind of under-talked
about with all this. I guess to pull it out to the macro. Who are the, because there's a major shift going on. We can't deny that. Open Source is really becoming the most popular thing. And and it's become the most responsible thing for these enterprises and startups to do because tokens been surged. Gross margins went, no, like maybe some gross margins went negative for a little bit. They, you know, came back up because they're accessing cheaper models. That kind of thing. Compute's always going to be expensive. But the macro level, let's break this down. Who are the beneficiaries of this shift to open source? And then who are the beneficiaries of frontier models? And how does that kind of change over time? Yeah. So I think who wins first. There's not an exhaustive set. But folks who obviously win are folks that are inputs into the ecosystem. And if you think
of the fundamental inputs to these, it's compute energy data. These inputs win no matter what. Like the you do, you do need these as you scale. The beneficiary of powerful open source, open weight models are enterprises and AI native companies that want to build custom models that help them retain ownership over their learning and their core workflows. So they benefit. The obviously the chip providers benefit no matter what. They benefit no matter what. Like and I mean, they're all running on GPUs and the and the energy and we have providers benefit too. I also believe in I mean, obviously, Javan's paradox is also a thing like when these models become much more efficient from a cost and energy standpoint, I think they're going to be used a lot more.
So that's also going to be a thing to steal man the side for the frontier labs and those building super intelligence. You could argue that one of the if you define super intelligence as exceeding human intelligence in every type of cognitive domain, one of those domains is also building small models and you can imagine like some of these the when a frontier lab builds a highly super intelligent model, you could ask that model to create small models that are more efficient for different workflows and maybe they are served differently. They are run differently. So you so it is possible that the frontier labs also capture value from small models that are cheaper to run and you can imagine an orchestration that the frontier labs do where let's say a frontier lab deploys at an enterprise. They're smart about when to use the truly in parameter model,
when to use the half a billion to 10 billion parameter model and self optimize. The thing that throws a span into the works is distillation which is a way for a way to train student models from stronger teacher models and distillation kind of keeps the gap between the frontier and open weight models relatively small and today there's no easy way to prevent distillation as well. So like it so it is so that that is that is also another another thing to keep in mind. I think who wins are also us like we are about to reap the benefits of our lives improving in every possible way. Everything from discovering cures to diseases that don't exist today, ways to extend our lifespan to something as simple as maybe the apps on your phone keep improving
at a much more rapid rate than than before. So I think whichever way this goes humanity benefits as long as we have good guardrails on safety. Today's episode is sponsored by VCX by Fundrise, the public ticker for private tech, allowing investors of all sizes to invest in venture capital. Learn more at getvcx.com. Some of you may not have heard this yet, but our sponsor public just launched something called generated assets and it brings AI into investing in a way I've honestly never seen before. Here's how it works. You type in an idea like AI powered supply chain companies with positive free cash flow or defense tech companies growing revenue over 25% year over year. Public's AI then dispatches a swarm of agents that scan every single US stock, evaluates them and instantly builds a custom index around your thesis. What really stands out is how clearly it explains why each stock is included. And before you invest, you can even back test your idea
against the S&P 500. So you're making decisions with real context, not just guessing. And beyond generated assets, public lets you invest in stocks, bonds, options, crypto, all in one place. They'll even give you an uncapped 1% match when you transfer your investments over from another platform. If you want to build a portfolio that actually reflects your thesis, visit public.com slash sorcery paid for by public investing full disclosures in the description. Founders scale faster on deal set up payroll for any country in minutes hire anyone anywhere get visas handled fast and get back to building visit deal.com slash sorcery that's D E E L dot com slash sorcery. What is your goal is your goal super intelligence by 2030. How do you think about all this? So our goal is to advance the frontier of super intelligence and make sure humans benefit from it. We want real economic progress. And we're going to push the frontier forward by working with
all the labs building these powerful proprietary models that help the world in so many amazing ways. We also want good open weight models to also improve. And we want the benefits to be to be diffused in the economy. None of this matters if we don't have GDP growth that's significantly more than we have today. I want the impact of AI to show up in the PNL of companies in the market cap of companies and ultimately in the GDP of the US and the world. And I really want to help the US stay ahead in the race to super intelligence and beyond. The thing that people miss when they talk about this open versus closed the doomsday versus the optimist is that the invariant here is usage of
frontier models is going to keep growing at a crazy rate. We're going to need more and more frontier intelligence. We're also going to need more and more open intelligence. And enterprises will leverage open intelligence to build their own proprietary intelligence. And that's actually good for the world to have. It's good for safety as well. Like many of these open weight models. Because they share their research, it helps us study these systems. It helps more people study these systems. It's a lot of people with really good intentions working really hard at the frontier labs at the open weight labs. And it's sad that sometimes they feel adversarial. But it's a lot of people, a lot of really smart people working really hard to build intelligence as an API. And if we solve that, I think that's the last puzzle for humanity to solve. Like almost all the problems we are looking to solve are intelligence bound and intelligence constrained. And it's going
to be much cooler. I would love for air traffic control to work. Air traffic control to work? Yes. Yes. That would be nice. And I had a bet with many of my friends for what would come first? Is it super intelligence or good Wi-Fi on airplanes? Oh my god. Well, you don't have to Elon. Well, the airlines have to roll it out. It takes some time. The airplanes have to come down and go through this whole process. But airplane Wi-Fi is completely fake. I stand by my statement. And what do you think will be the last problem that will... What would be the crazy thing that will not yet be solved even when we have super intelligence? Human intelligence? Not to be so. But the thing is you cannot solve for human behavior. So as much as people want to put down some of these things, we're always going to want entertainment. We're always going to
watch funny videos or dramatic things and all this kind of stuff. We crave entertainment. The normal people out in the world, not in the Silicon Valley bubble. Go visit them. See what they do in a day today. They're really not touching technology all that much. And they usually have normal lives. And the classic things of a normal human, whether it's exercising being with your family, eating, watching something for entertainment. It's just not that complicated. I think we like to complicate things. But to be human, it's like this wonderful, blessed, simple life. And so we just impose so many complications on ourselves with artificial intelligence or super intelligence, whatever you want to call it. A lot of the complications of ordinary life will hopefully be kind of quieted away. I mean, I think one of the biggest things, and this is something that I'm doing more with these interviews is going to the health level. I think one of the best things on
top of economic progress and equalization and democratization across the universe. But also, you know, the United States is that access to healthcare will be democratized. Like I was just in New York, I just went to the neckl, the necko health opening. I did the scan there. Their goal is to democratize health access. You get access to preventative health, do all the checkmarks, all that kind of stuff. Make sure that you're on track. If you have any like things to worry about, to monitor it, get more help, that kind of thing. But think about it. Like we live in major metro areas. We are like a minority of the United States. Although we have large populations, like we still like really need to help out the middle. And the middle is feeling a lot of the pressure now, whether it's for the K-shaped economy or whatever you want to say. So I'm really excited to see how health expands and health access to health expands through US with access
to more technology and superintelligence. Another thing is I have this episode coming out with Annie LeMonde of Oak, HCFT. They're a big healthcare investor. One of the paradoxes with AI and healthcare is that it's probably the most unseen thing. Like you're not going to see your doctor using an AI scribe. You're not going to see any of the things that go behind the scenes. You're just going to have better care. So I hope that that helps. But we'll see. As we close out, I have to ask you, what is your hardest take right now? The best way to ensure that we move AI forward is to close the research and deployment loop. So we're touring like we are working with all these labs to improve their models in coding and knowledge work. And in doing that, we discover what all these different models and agents are good at. We observe the jagged intelligence that they have.
And that's made us this unique partner to enterprises to deploy end-to-end agentic systems. And we deploy them. We see where things break in the real world. What are the capability gaps that exist? And we leverage that to further improve the models so that we can solve even bigger problems in enterprise, see where things break, improve the models, and we keep executing that loop. And making us this trusted partner between research and deployment. And I think that's also the path to making these systems more safe. You to actually have them master real work, you have to you have to have them see reality. And I think increasingly we'll see we'll seek, we'll see labs, enterprises focus on closing that loop between research and deployment.
I also think we're going to be in this phase of building super intelligence for the next 10 years at least. And I know some people believe that okay all the jobs will be gone in to in everyone will no one will need to work after two years or stuff like that. I don't believe that. I think you already will hit. Yes, singularity will help. I think it's RSI or recursive self-improvement is definitely a thing. And the way that works is people are using these AI models to speed up pre-training, to speed up post-training. And to because if you look at pre-training where you have to minimize publicity or pre-training loss, that's a verifiable domain. You know you you write some code, you can check whether did the loss come down or not. So in theory AI could run in a loop to invent better and better algorithms for for pre-training. And in post-training when you're doing
this RLVR, you could again check where hey you built all these oral environments. How did you do on these generalized intelligence benchmarks? And you could potentially keep hill climbing. And but the thing that's missing in that recipe is we are still optimizing just the inner loop. We're not optimizing the outer loop of what are some new algorithms you could come up with that don't use LLMs at all. Maybe they don't use transformers. They don't use gradient descent. Maybe they don't even use neural networks. You're not going to discover that in this paradigm. So there's still a lot of a lot of there's still a long distance to go. In Silicon Valley, I feel like there is this polarized group. There are some that believe in fast takeoff like where you know in two years we're going to there is a risk of loss of control and maybe like we we AI exceeds human intelligence in everything. And then there are maybe the skeptics who believe
that it's all hype or like nothing will happen. I actually believe in slow takeoff. I think over the next decade or two these frontier models are going to become increasingly more powerful and capable and useful. And but the technology will take time to diffuse especially in enterprises. But it is going to happen. It is already happening. There is so much model capability overhang. It's just that the real world is messy. In real enterprise work when a human starts a job, they don't know where all the information is in the company which people to talk to. The context that you need is distributed across the company in people's heads in a variety of different files. A manager gives a task to a employee that is ill-specified or ambiguous. You have to go and acquire more context to do it. It will take time. But it's going to be an incredible journey. And it's I feel like the next 10 years will be the brightest for technology yet. Amazing place. And thank you so much Jonathan. Thank you.
Hey it's Molly. If you enjoy our interviews check out our newsletter sorcery.vc where we deliver a once a week top deals and tech headlines email and also go deeper on our podcast interviews. Subscribe to sorcery today and don't forget to subscribe to the podcast on YouTube, Spotify, Apple or wherever you listen. Link in description to sign up.
More episodes
More from Sourcery

Annie Lamont: $14B Managed, 70+ Exits, 15 IPOs, 7x Midas Investor
Sourcery

IMEC Says Today’s AI Will Look Ancient in 10 Years
Sourcery

The $500B Opportunity in Europe’s Defense Buildout
Sourcery

The $10T AI Buildout Has a Photonics Problem
Sourcery
