
How Replication Could Teach Machines What Good Science Looks Like — Edward Hughes
About this episode
Can a machine learn the judgement that separates a plausible-looking result from a faithful experiment? Edward Hughes, Chief Scientist and co-founder of Inherent, joins Tim Scarfe to argue that creativity is not optimisation, and that the missing capability in AI is choosing which questions are worth asking.
SPONSOR:
---
Cyber Fund built the Monastery to help founders ship products that were impossible a year ago.
Apply now: https://cyber.fund
---
Edward makes the case that Move 37 was innovative rather than creative, and that the field, not the individual, decides what counts as a discovery. That reframing runs through Csikszentmihalyi, Deutsch and exaptation into open-endedness, where deceptive goals and imperfect world models turn out to be the point rather than the problem. The second half turns to the paper: Replica, a task space built by redacting figures from real papers, and Faraday, a 27-billion-parameter model trained to steer a frontier coding agent that then beats the frontier on held-out replications.
---
TIMESTAMPS:
00:00:00 Cold open: Move 37, Faraday and collective intelligence
00:01:08 Sponsor: CyberFund
00:01:46 Inherent's $50M raise and the road from string theory
00:09:14 Three timescales of learning: weights, context, culture
00:13:47 Move 37 was innovative, not creative: the field decides
00:20:39 Creativity as satisficing: the urinal and evolution
00:25:06 Exaptation and the Tristan chord: creativity in context
00:30:56 Coherence for whom? Deutsch's hard-to-vary explanations
00:35:53 Why copying is creative: Deutsch and the constraint engineer
00:42:27 Societies of agents and the strong Moravec paradox
00:45:51 Evaluate in hindsight: from Lean proofs to climate change
00:51:56 Picbreeder, local goals and why discovery needs deception
00:57:21 Spaghetti proofs, translation layers and superhuman Go
01:00:37 Does nature compress? Naturalness and real patterns
01:07:36 Why replicate? Replica's redacted figures and Faraday
01:12:31 Faraday beats Codex, Claude and GLM 5.2 on held-out tasks
01:15:31 Replication to innovation: how the Transformer happened
01:18:26 Deep replication: what Faraday learns from Voyager and GNoME
01:23:37 Can the AI scientist cheat? Goodharting the judge
01:29:09 Inside Replica: scale-down, 8xB300 runs, per-task rubrics
01:34:11 The RL crisis: getting GRPO to work with per-turn credit
01:39:43 Weights vs harnesses: AlphaEvolve, DGM and EvoTune
01:45:45 The recursive company: agents cross a phase transition
01:50:35 Collective intelligence and the electric dynamo
01:55:46 What replaces OKRs? Incumbents and the burden of knowledge
---
REFERENCES:
organization:
[00:01:47] Inherent
https://inherentlabs.ai/
other:
[00:20:51] Marcel Duchamp, Fountain (1917)
https://www.tate.org.uk/art/artworks/duchamp-fountain-t07573
[00:57:33] OpenAI unit distance
https://openai.com/index/model-disproves-discrete-geometry-conjecture/
[00:05:19] Human-Timescale Adaptation in an Open-Ended Task Space (Adaptive Agent)
https://arxiv.org/abs/2301.07608
[00:06:05] The AI Scientist
https://arxiv.org/abs/2408.06292
[00:12:13] Training AI Scientists to Replicate Research (Replica and Faraday)
https://arxiv.org/abs/2608.13331
[01:44:46] Evolutionary Principles in Self-Referential Learning
https://people.idsia.ch/~juergen/diploma.html
[01:59:33] Are Ideas Getting Harder to Find?
https://www.nber.org/papers/w23782
book:
[00:16:04] Creativity: Flow and the Psychology of Discovery and Invention
https://search.worldcat.org/title/254487436
[00:26:22] Why Greatness Cannot Be Planned
https://link.springer.com/book/10.1007/978-3-319-15524-1
[00:33:03] The Beginning of Infinity
https://www.penguinrandomhouse.com/books/293575/the-beginning-of-infinity-by-david-deutsch/
[01:55:47] Laws of Knowledge
https://www.penguin.co.nz/books/the-infinite-alphabet-9780241655672
(Full list refs on YT/rescript)
---
RESCRIPT:
https://app.rescript.info/session/670296ba913761d0?share=6281911cac9bdbff637f10819d4d1e5c
Get every episode summarized
Each time Machine Learning Street Talk (MLST) publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
1,388 searchable segments. Every word is indexed and playable.
Full transcript
Machine Learning Street Talk (MLST) — How Replication Could Teach Machines What Good Science Looks Like — Edward Hughes. Machine-transcribed; use the interactive transcript above to jump the player to any line.
This episode is brought to you by Google Chrome. You think you know a browser, but Gemini and Chrome? That's new. It can help you with practically anything on the web, like restoring a vintage motorcycle from a 50-page restoration block, or finally break down that long article you've had open for weeks. Gemini and Chrome is here for it. Ready to make anything online makes sense? There's no place like Chrome. Check responses set up require compatibility and availability varies 18 plus. Propel Fitness Water with Gatorade Electrolites, Zero Sugar and Vitamins. Propel hydrates better than water to help you get the most out of your workout and get back to your best self. What propels you? Propel with Gatorade Electrolites. How do we start to build AI scientists systems that go beyond simply answering questions that we pose and start to ask the kinds of questions that lead to open-ended creators? You spoke about move 37, and I think you and I would agree that that was definitely creative.
I don't think that move 37 was creative, which I think that move 37 was innovative without being creative. And what we find is that our Faraday agent is able to perform better than the frontier model. But it's also performing better than other frontier coding agents like Claude, for instance, and had these agents proactively reaching out to us and trying to help us with stuff, and to be perfectly honest with you is quite annoying. But you just had no idea how to help us. But at some point about two or three months ago, I think we reached a phase transition. What's the equivalent of OKRs for the age of recursive self-improvement? Indeed, and I think that is the next era. It's the era of collective intelligence, rather than the era of individual intelligence. This episode is supported by CyberFund. If you're building at the front here of AI, they want to hear from you. CyberFund believes the future belongs to AI natives who want to achieve the impossible.
And that is why they're introducing the monastery for AI native founders. It's an environment of pure focus and rapid execution for founders operating at AI native speed. And they're offering teams $2 million each to participate. Apply now at cyber.fund. Just a quick piece on it here. And so you were at Google DeepMind before. How much funding have you got? Yeah, how many employees? What's the valuation? Yes. So we've raised $50 million from index ventures and radical ventures at a post-manivation of $225 million. Currently we're a team of 11 people. And presumably you think that AI is the most important thing in the next five years, just generally. No, I wouldn't say that. I think that AI by itself, in some sense, is meaningless. Really it's what, how can AI, what can AI enable in interaction with the wider ecosystem?
And that, for us, we're most interested in broader scientific progress. It's amazing to have you on MLST. Welcome. Let's should be here. Tell us about yourself. So yes, my background originally I was a theoretical physicist, as my PhD and I was studying string theory in particular, the scattering of particles, but using string theory as calculation mechanism. And one of the central feces of the work I was doing was an idea called dualities and physics. Duality is a mathematical map if you like between two different theories. And it can take you in some quite unexpected places. So usually it's useful if you have calculations in one theory which are very hard. You can map across this duality and calculations in another theory that are much easier. And so I was using some of these dualities in order to take calculations that were very difficult and map them into a different geometric space in order to do them.
And I think that talked me a lot about how to think about transforming problems, something will probably come back to when we talk about creativity. But towards the end of my PhD, I became convinced that working in scattering amplitudes was not the most profitable thing I could do in terms of making an impact on the sum total of human knowledge. Funnily enough, many of the people that I was citing at the time of my PhD, people like Jared Kaplan and Jeffrey Pennington also made a similar decision around the same time. So in 2016 towards the end of my PhD, I saw the Alpha Game match, I remember watching that, like so many other people being fascinated by move 37. And I became convinced that the future was going to involve agents that could really aid humans in making discovery and accelerate the rate of scientific progress. And so I joined DeepMind in 2017, really animated by that question, how do you build an agent that can itself make discoveries?
And I started by working on that in the context of reinforcement learning, particularly multi-agent reinforcement learning. The reason I was so interested in multi-agent reinforcement learning was because human technology seems to have been created not so much by one individual, but by the sum total of human culture. And I like to think of cultural evolution as the fastest intelligence generating process in the universe. And so I became, I suppose, obsessed by this idea of really trying to distill cultural evolution into agents. Now at some point on my journey, the foundation model started to arise. And the, I started to work on building much larger models. So I built a model called Adaptive Agents, which was all about using metery reinforcement learning across a very, very large space of tasks in order to build what was at that time. I think the largest change from scratch, RL agent, 500 million parameters.
So it's pathetic by modern standards, but at the time it was rather large. And then I worked on world models. So I suppose my claim to fame there was that I wrote from, from, by hand, effectively before coding agents, the final implementation of the GD-1 model in a single file. So I think that was the last time I really wrote a fully artisanal human written implementation of an algorithm. And then from there, I led the team called AI scientist. And what we were trying to do at DeepMind was apply coding agents to the problem of doing scientific research. And this was in parallel to many of the developments at Sikana and elsewhere. And so I led that team with Lewis Kirsch for just over a year. Lewis had been Yergen Schmidt, who is PhD student, and he'd come to DeepMind with a very similar interest to mine. But in the middle of last year, we'd become convinced that we needed to build an entirely
new company to take this idea to its, into its most ambitious form. And that was really for three reasons. The first reason was that we came to believe that developing an agent capable of discovery wasn't just about a single agent. It was about the way that that agent interacted with the entire ecosystem around it, of scientists. And so what we wanted to do was some fairly radical organizational design transformations. And to do that in a large established company is much more difficult. That's the kind of thing we're building at inherent. We call it living within the experiment. That's the first reason. The second reason is that we started to think about the way that the major paradigm shifting discoveries come about in science. And at least from my reading, I, and also from my personal experience of having shifted
areas a few times, I observed that most of the biggest paradigm shifting discoveries happen when you have knowledge in one area that gets transported to a different area. And then that unlocks some unexpected connections, some unexpected advance. And so that required us to really take a horizontal view of science. And so what we think of, we're building at inherent is that we're building horizontal intelligence layer for all of science. And the third reason that we wanted to do this was really an infrastructural reason. So it turns out that if you want to both reinvent the organization and you want to build this horizontal AI scientist technology, what that implies is that you want to give the AI scientist all the same affordances and context as humans.
That's very tricky to do in an established company because you don't want your agent running riots on YouTube if you're Google. But in a new company, you can start building the infrastructure from the ground up, not so much just around sandboxing, but also around permissioning so that you can have agents that really have very equivalent environments in which to operate as humans do on a day-to-day basis. So that's really a founding tenet of absolutely everything that we do. And I think that creativity has a lot to do with respecting constraints. And I use the word constraints because that feels like the most abstract form of knowledge, because we have a privileged form of cognitive knowledge and there's cultural knowledge. And I even think that constraints in the physical world can be thought of as some form of knowledge. So the most abstract possible way to describe creativity is respecting constraints. And what we want to do at different levels is accumulate knowledge, which means we need
to find or discover these constraints. And I love using the analogy of a maze. So when we are discovering the shape of a problem, what we're doing is we're kind of discovering the walls in the maze. And in our evolution, it happens at multiple levels. So there's this kind of DNA phylogenetic evolution in the course of our individual lifetimes. We have this ontogenic evolution. I'm overloading Lamarca a little bit here because he was talking about it in terms of heritability as well. But you see the analogy with AI. So we adapt to the weights. And then our individual agents get experience and they adapt to their skills surface and their memory systems. And then it feels like there's a third wave, which is the fastest form of knowledge accumulation, which is cultural accumulation. So in the future, the agents themselves will be talking with each other, collaborating or maybe colluding and building up this latent cultural knowledge. So I mean, do you think, just before we kick off, is that a reasonable model to understand
how AI is going to progress? Yes, I think there's a lot of that I agree with. Maybe I'll unpack it backwards. So I think the one wrinkled add to that at the cultural level is I don't believe that there will be some separate culture for agents and a separate culture of humans. In fact, I think the super intelligence, and so far as it will exist, will be a combination of humans and agents interacting in very deep and very complex ways. And that knowledge will be the better that we're able to interconnect agents and humans and leverage their complementarities. The faster we'll be able to drive this accumulation process of knowledge. Now I think that absolutely that there is this distinction between if you like the slow weight updates, which are maybe more akin to the evolution reprocess for the NA and then the fast in context updates.
I suppose one of the things I'm very interested in at an architectural level is whether this analogy will still hold in five years time, for instance. Now with a lot of pieces, if you like, it could sit in between those two epigenetic effects, for example, a one. And what's the analogy for epigenetic effects in an AI system? So one of the pieces in our new paper that we talk about is this idea of a weaker coding agent using a stronger coding agent as a tool. And there what we do is we actually update the weights of the weaker coding agent, because that has the benefit of generalization more so than just updating the context. But because the weaker coding agent is using the stronger coding agent, that weaker coding agent is, of course, injecting things into the context of the stronger coding agent. So now you have to ask yourself, well, is this weight or is this context? And the answer is, of course, it's both.
And perhaps we've got the opportunity to have a much more intricate spectrum between these two things than we currently have. And that's one of the themes that we're in particular investigating because in the cobbled together nature of our current AI stack, what happens is you get these abstractions that become very sticky, rightly so, because they work well. Because of the burden of knowledge for humans to understand the frontier, we just accept a very large number of these abstractions because it's just too complicated for us to be examining all of them in combination. And the promise I think of AI scientists systems, agents that really understand this horizontal is that they might be able to weaken multiple constraints at once and thereby enhance the ability for creativity. Another loop that you opened earlier was, you spoke about move 37.
And I think you and I would agree that that was definitely creative. It feels like there are some limitations to its creativity. So I would call it a form of concrete creativity. So it doesn't understand in the sense that Margaret Bowden would speak about in terms of understanding how it hangs together in the context of the system and what is possible and counterfactuals and whatnot. It's still a form of concrete understanding. But another interesting angle though as well is you were talking about the difference between possibly human knowledge and AI knowledge because I was speaking with Tom McGrath at Goodfire. He did interpretability on Alpha Zero. And his idea is very much that these things are learning the space of human concepts and beyond. And we could actually mine those representations as a new form of science. So these things are discovering interesting things that perhaps we would discover but haven't discovered yet. And we could actually use this as a laboratory for discovering interesting new knowledge. Well, interestingly, I don't think that move 37 was creative.
I think that move 37 was innovative without being creative. So the way that I think about innovation is innovation is the process of taking unknown unknowns and making them into known names. And in order to get an innovation, you can't generate an innovation if you said already knew what it was you were looking for. You have to have something that's unexpected but then it also becomes valuable. So what's the difference between innovation and creativity? Well, in my mind, creativity requires another step which is to recognize that the thing you have done is creative. And who was it who recognized that move 37 was a remarkable move? It wasn't the, it wasn't AlphaGo. It was the commentators, for example, who, if you watch the famous footage, say, oh, that must have been a clicker. Was it a click miss? It wasn't, of course, it was, it was exactly the right move. And so that is a kind of meta-metacognition. So you have to sort of understand that you didn't know something and now you are updating
your own knowledge as a function of that. And it's, it's interestingly discussed by Mahalic, San Mahalic, he wrote a lovely book called The Psychology of Creativity and Invention. And he's also the person behind flow. So many of your listeners will already know a concept from him. But in this book, he interviews a very large number of different creative people from the different disciplines. And he comes up with an ontology of what creativity is. And he says creativity has got three components. There is the creative individual. That's the bit that we always focus on. But that's really, in some sense, the tip of the iceberg. The second piece is the domain. And the domain is a set of symbolic rules, if you like, to which the creative person is adding, or perhaps breaking one of the rules, and then thereby expanding the space of possibility. But the third piece, which is perhaps the most forgotten one, is the field. The field are the set of other individuals who are going to decide whether the creative
person's contribution gets admitted into the domain. And he gives this lovely example of Florence in the Renaissance. So we're talking 15th century. And in the 15th century in Florence, there was an enormous flowering of creativity, whether it was architecture, science, even the way that society itself was structured. And the question is, what was it that led to that in Florence? Now, you might say perhaps it was just an expansion, the number of creative individuals. Was there some mutation of the DNA, some new educational system? It seems quite unlikely that there could be a mutation in the DNA. And so far as we know, there was no great change in education. Well, then you have to ask, okay, what's it, the domain? And I think at least in part it was the domain because at that time, many building techniques that had been lost to antiquity, which were in fact known to the Greeks and Romans, were being rediscovered via archaeological means, via analysis of the building structures that
people were uncovering. But it couldn't just have been the domain because much of these rediscovering was happening in Rome. And Rome didn't have the same flowering of creativity as Florence. So the third thing you need is the field. And what Florence had, the other Italian cities didn't, were lots of very rich families. The Medici is the most famous among them, but I don't believe it was the richest. And they were rich from the wall trade and then also from becoming financiers as well. And they had this idea of making Florence the most beautiful and most cultured city. And that was in some sense to weave a protective cloak around the city at a time when there were many city states, there was quite a lot of conflict. And they believed in this idea of beauty as in some sense, a kind of psychological defense. And as a result, there were a very large number of creative constructions that were admitted
into the domain. And it became a competition between the artisans of the day. And so I think it's instructive to think about how that might play out with AI scientists systems. And in particular, I think it becomes much more interesting when these AI scientists systems start to be able to do things which are generalizable. So what do I mean by that? Well, of course, move 37, remarkable innovation, but it doesn't really tell you how to do innovation in other domains. And immediately see a line from move 37 to discovering a new material, for example. And even if you think within a single organization, the line between move 37 and say alpha fold wasn't a particularly direct line, it's not like the alpha go agent or indeed the alpha go training techniques really informed alpha fold in a very direct way. And in principle, if you had had a generalizable discovery engine, then it could make a discovery
about the weather and then figure out, oh, there's some part of that discovery, perhaps, it's the architecture of the neural network that was used in those to make that model. I wonder whether that applies to protein design and those kinds of connections, I think, generalizable connections are going to be what leads to a large acceleration in the rate at which we can make discoveries. Yes. Because you were discussing how we recognize creativity and the social component is extremely vex because it's very tempting to think there's some degree of social proof in creativity. And indeed, perhaps there is, I mean, there's a famous example of a urinal with a bit of masking tape on and everyone just decided that it was creative. And I tried to think about it abstractly. So for me, something is creative when it becomes a mode in the state space to a certain extent. So that clearly became a social mode. I mean, what's difficult about us as individuals and cultural learning is the introduction of agency and the fact that we could have done differently. But I suppose going all the way down to physical creativity, evolution isn't an agent, it's
not doing planning, but there are still these cannalyzed modes. And the way I think about it is it's a bit like the system has discovered an interesting new subspace. And that subspace is being used in a myriad of situations. So we would call that discovery creative. And perhaps even with Alpha Zero, maybe if we enumerated many possible game trajectories, and if we saw something that looked like a category. So this particular type of pattern was being rediscovered, reused in many different situations. We could immediately look at it and just draw a boundary around that category. But maybe Alpha Zero would kind of have competence without comprehension. If it was using this thing in many different situations, maybe then we would call it creative. Yes, I think I want to come back to the idea of a relationship between creativity and constraints for a moment. So if we look right back at evolution itself, I think evolution quite clearly is creative. It's certainly generated this enormous amount of diversity in the natural world.
And the way it's done so is exactly by satisfying. It's just a posh word for saying, satisfying constraints. So why do I say it's satisfying rather than optimizing? Many people might think it isn't evolution trying to optimize for the best of an individual with the most adaptive traits. Well, actually all the evolution requires is that individuals survive and reproduce. And once you've done that, there's not a lot else that you can do. Now, perhaps you could say you can do second order survival and reproduction. That is true. You probably care about your children surviving and reproducing as well. So it's not quite as simple as that. But even that second order, that's still a constraint satisfaction problem. And what's there's two interesting implications of this. The first one is if you really believe that creativity is satisfying, then the default crutch that we reach to in machine learning, which is optimization, is the wrong thing
to reach for, to build creative agents. And secondly, if you believe that constraints, satisfying is important for creativity, then the way that you open up new creative spaces is that you take your existing constraints and you break some of them. Now why do I say break some of them? Clearly if you break all the constraints, then there's no meaning left. The way that we construct meaning is, and indeed the way that we construct laws of the universe, is that we rely on things being repeatable. This is the so called principle of induction, which is not something that you can prove, but it's something that we just observe. The laws of the universe seem to stay the same from moment to moment. You can't break every single constraint, otherwise we'd live in a world of white noise. But breaking some of the constraints is very useful, because at least some of the constraints
arise because of our existing theories. We don't have access to the universe. We only have access to the universe through our observations and measurements of it. And in order to make those observations and measurements, we do two things. We have tools that allow us to make those measurements, and then we have our own neural apparatus, which allows us to make interpretations of those. And so by relaxing or breaking some of those constraints about the interpretation or about the tools where they then able to access new insights about the universe. And so that's where creativity really arises. Yes, you've opened so many loops there. I don't know how to close them off in order, but I'll try my best. The thing that you just said is very interesting, which is this very vexed issue of coherence. I mean, a tonal harmony is the great example of this. And I often argue with my co-offered on this article that we wrote, whether that is breaking the constraints or inverting them or just respecting them in some other way, because
a lot of people talk about knowledge being quite situated. And what they're meaning in that case is that it's only coherent if you respect the constraints. And sometimes it's not possible to break the constraints. And another thing you spoke about, and this is also related to your 2024 ICML paper, which was Open Eddiness is, I think, necessary or required for errors. Yes, beautiful paper, by the way. And I said to Tim Ruktaschial at the time that I felt that was actually a definition of creativity rather than Open Eddiness. Or actually, I think they're basically the same thing. And the reason I think they're the same thing is that intelligence is basically about optimization. Right. And the question is, I'm trying to find the shape of the maze, and I don't know its full shape here, and I'm trying to fill it in, and I can go in that direction. Creativity, as you were saying before, it's about discovering new questions, new problems, new mazes. And as Kenif said in his book, Why Greatness Cannot Be Plan, there's a weird paradox there that when you optimize towards something, it's really, really difficult for you to find
something interesting and creative because you've got the blinkers on. Yes, well, I think let me give two examples that pertain to that description. So one of them is a rather beautiful concept in evolutionary biology of expectation. So we know, of course, that adaptations persist across evolutionary time because they give some advantage to the individual, which allows them to be selected for, perhaps, their better able to escape from predators, perhaps their better able to find food, for example. So what do I mean by an expectation? Well, an expectation is an adaptation that was giving the organism some advantage, which then finds a use somewhere else, an unexpected second use. And indeed, we see this in biological evolution, but we also see this all the time in the famous discoveries of science. So whether it's Alexander Fleming and Penicillin by leaving the Petri dish out, whether
it's the invention of the microwave during a radar testing, where I think the individual in question left the chocolate bar in his pocket, so it's melted. My favorite one as an ML person, of course, is GPUs. GPUs developed for gaming, and it just so happens that's exactly what you need in order to, well, first of all, optimize the training of convolutional networks, but then now, of course, adapted for the optimization of all modern neural networks. So that's one example that the second thing, the second example I want to give is completely different is coming back to the musical examples. So as you know, I have this sort of moonlighting career as a semi-professional musician. And the, I love that example you gave of atonal music. There's perhaps an even sharper one, which is the famous Tristan chord. You can go and look this up on Wikipedia if you don't know about it. But it's the chord that Vagnie used right at the very start of his opera Tristan and
his older. And it's the very first chord in the piece. The piece starts with three individual notes of melody and then this chord. And it's seen as a very creative chord. And it's really interesting to inspect why that is. Now the chord that he uses is not new. This chord, you can go and see this chord back, you know, hundreds of years before the same chord was used. But the thing that was new is the context is how it was situated to use your term. And the context is that the chord is very harmonically ambiguous. You're not at the point where you've yet established the key of the piece. And so as a listener, you immediately question, okay, well, what is this chord saying? And in general, up until that point in musical history, harmony had in some sense been used as an accompaniment to melody. But at this point, Vagnie is questioning that. And he's asking the question, well, what if rather than using harmony as a accompaniment, I use harmony as communication directly.
And so the thing he's trying to communicate in this chord is exactly that ambiguity. What is going to happen suspense? Perhaps this sense of confusion or impending chaos. But also a sense, slight sense of hope as well. There are many things that are happening in that chord. And this really prefigures a lot of musical developments in the 20th and indeed 21st century where harmony is used for color. It's used for emotion. And in fact, we're all intimately familiar with this because this is used to incredibly great success in film music. You can immediately identify just by the very first couple of seconds of a chord or a harmonics sound world at the start of a scene. Even before you've heard 30 seconds of melody that this is going to be a rather chilling scene or a rather hopeful scene or a love scene, for example. And so I think that is a sort of microcosm of the point you are making, which is that
you really creativity has to be judged by standing on the shoulders of giants. It has to be judged situated in the place that is currently in the canon. Yes, again, absolutely fascinating. And the way I interpret this is you're pointing to, let's say if I edit a video or I make some music or something like that, you're saying in principle it could be quite ambiguous and then it'll be interpreted using the constraints of observers. Now the observer thing is very important because in your 2024 paper you were talking about the perspective of an observer, whether a stream of an events produced by an open-ended system is novel and learnable and there's a kind of a virtuous complexity gradient that we can climb. But I still think that coherence is a binary property. So when artists create things, they usually have a set of constraints that guides its creation. It could be an intention, you know, like when Michelangelo was painting the Sistine Chapel, there was lots of cultural constraints and what he intended to do at the time.
But the observer-related thing is interesting because let's say you're a very clever person and you write some mathematics and you show it to someone and they don't understand it. It looks like slop to them because they can't recognize the constraints that guide the process. But it's not slop because I think it's objectively coherent. They just don't understand it yet. But still, when you have this kind of cultural transmission, this is a great form of new adaptivity because it will be reimagined and reinterpreted in a different content. Yes. And I think that actually a lot of creativity does arise from that underspecification. I think it's one of the rather wonderful features of humans is that we can't really transmit our ideas to each other. We have this very high noise, very narrow bottleneck channel, which is our description of things
in words to try and communicate an incredibly high-dimensional state space in our brains. And in some ways in our bodies as well, athletes, for example, come to mind. I want to just come back for a moment to the book by David Deutsch, the beginning of it. That's right there, the beginning of it finishes. Yes, this book. What was his definition of science? Well, okay, let me do the definition of science and then I'll come back to this point also of replication. He says that science is a search for good explanations about the universe. And he's very precise about what he means by a good explanation. He says, a good explanation is one that is hard to vary. So let me give you an example. Let's suppose that you say that the sun rises every morning because it is pulled on a chariot by the gods. And let's suppose that over time the sun is rising later and later every morning, for example,
that happens to the northern hemisphere as we go from summer towards winter. Now there are now a number of different ways of varying that explanation. Maybe the gods are getting more tired. They're sleeping in so the sun is rising later. Perhaps the gods are angry and that's the reason why. Now let's suppose that one day the sun doesn't rise at all for, for say, an hour. And maybe this is something we'd explain as an eclipse. But it can now be explained as the gods either inflicting Roth or the gods giving you a chance to sleep in. Perhaps that the gods are inflicting their favor upon the world. Now suppose instead that you try to explain that the the the a diurnal cycle 24 hour a day cycle of day and night by the fact that the earth is spinning on its axis. And now let's suppose that you have to explain that the sun is rising later every day. Well, the most natural thing to say then is, okay, well, perhaps the earth is spinning a bit
slower to make the sun rise later. But then hang on that the process, the diurnal cycle is still 24 hours. So how can it be spinning slower but also have a cycle of the same period? So now you see your your forced into a more creative space and you're forced into maybe suggesting that the earth is not only spinning, but it's tilted on its axis and it is orbiting the sun. And now it let's suppose that you have the a solar eclipse phenomenon. Well, that's pretty odd because you can't just sort of spin the earth into the place where the sun disappears and then spin it back again. That seems like that would require an enormous feat of celestial engineering if you like. So you then have to posit some other body in this case, the moon that comes between the sun and the earth. But that body has to come between the sun and the earth at a particular time and with the regularity that is consistent with the rest of the theory. So that's what he means by good explanation. Let me just come on to this other point though about how he thinks about cultural transmission.
And this is really buried quite late in the book. And I think it's a really beautiful account of how creativity arises. So he talks about two mysteries. The first mystery is a very prosaic one. If I stick my hand in the air like this, then I can ask you to copy that. And you will be able to copy that. And that is an incredibly cognitively difficult thing to do. Why? Because you have got a visual cue of me sticking my hand up. But you haven't got any of my proprioception. You certainly don't have any information about my muscles or, indeed, what I did with my neuropsychetry in order to do that. And you've got to reproduce that within yourself. That's one mystery. Mystery two is the mystery that around somewhere between about 10,000 and 4,000 years ago, there was this real explosion in the creativity of humans, at least as measured by the archaeological record of the density of different types of technology.
Now that's not to say that there wasn't creativity before. We know cave paintings go back a lot further than that. But certainly in terms of the sheer variety and accumulation of these technologies, something special seems to happen around that time period. Now, even though that time period is a few thousand years, it's certainly not long enough for biological evolution to have done much. So something arguably the biological prerequisites must have already been present in our brain. So how was it that biological prerequisites were present in our brain? But they weren't being used for anything. What was it they had adapted for? And he rather beautifully solves both problems at once. And his solution is that it's that very active copying that is creative. In order to transmit an idea, whether that's a physical idea or an advanced idea, for example, a cultural norm or a technology, you need to recreate what someone else had in their brain.
And that requires an active creativity on your part. And what changes in that 6,000 year period is not really much about the individuals. It's something about the field. Suddenly the leaders and the societies of the time come to value people who accept that creativity, not just for copying but for doing new things. And so this is a lot of the reason why at inherent, we're starting with the idea of replication and thinking of that. Yes, and we will probably get into your paper just in a short while. But to push back on the copying thing, so that canonical example of bad shallow replication is a photo copy. Let's say a forger. So a forger can just make the right colors and the brush strokes and so on. But all of the inner structure, the abstract structure, the intentions, the motivation, the constraints are absent. And I should bring in Michael Thomas-Sullow, because we're interviewing him in a few weeks.
And he said human cumulative culture depends on shared intentionality, teaching, normativity and retriting, not just copying. So this is really interesting because I think what you're saying is that we can do this kind of imitation learning. But what we actually need to do is recreate. Because as Kennev Stanley said, it's not about where you end up. It's about how you got there. So the challenge is to recreate the path which led there. And I think, let's say ancient humans, they painted on the inside of caves and stuff like that. And what made it learnable and cognizable was the fact that we have the same physiology. We have the same structure of the brain, same affordances and whatnot. So maybe that was an easier problem to recreate that abstract structure than say an artificial intelligence, which is learning on more surface level data. In some ways, although I think in some ways that there's a, it's almost harder in artificial
intelligence because there aren't as many constraints. So what do I mean by that? I think the wonderful thing about humans copying each other is that we don't have access to most of the information. It's a very, very partially observed setting. So I can't see your neurons, but I can take the very small number of bits of information you give me and reconstruct at least some of what your intention is. And it's that bottleneck and the fact that we have also constrained physiology that means that copying, in my view, actually begets many of these other downstream facets that Michael Thomas said I talked about. Now, I don't claim that in fact, it has to be just unidirectional. I think actually very likely there was some sort of ratchet copying as part of the mix. It's wonderful work by people like Cecilia Hayes, for example,
who talk about the same equipment for social learning, being actually what you need for social learning and the interaction between the two being very important. But to come back to your idea of the photocopy, quite clearly the photocopy is an anti-pattern. Now we can't, fortunately, we can't photocopy humans and we can't photocopy human ideas. And it's that which has led to creativity. Unfortunately, in the case of AI, we actually can photocopy the weights of a model. You can't do that with a close-vacemodel that you can with an open-weight model. And that lack of constraint actually makes it harder to arrive at creativity. So I think of a lot of what my job is. And I think increasingly, to some extent, as AI becomes more spread in society, this will become a more common role for humans to play, is as a constraint engineer. What is it that we need as the interfaces between these systems that will promote novelty and creativity?
Or in other words, how do we need to regularize away from purely generating faximiles into generating that's much more complex series of social and technological interconnections that lead to some of these things like invention and normativity and shared intentionality between humanist machines? This brings me on to another thing as well. So a lot has been spoken about functionalism, for example, which is that we're building machines and we say that if they have the same abstract functions, then essentially they're the same as us doing the same thing with our physical instantiation. But when you look at evolution, interesting kind of prometheus moments happen. So there's the emergence of language and this copying cultural accumulation that you just spoke to. And something fascinating is happening now, which is that we are training these foundation models. We're doing some RL post-tuning. And then they exhibit different forms of intelligence and agency in different configurations. So they weren't trained to do this, but we can
now create a society of agents. And they just have this kind of phenomenon that wasn't part of their evolution. I mean, another example of this is I could take a herde of lions, for example. Every individual lion is not touring complete. It's not intelligent in the way that we are, but you could imagine a configuration of lions that had more intelligence and more capability than all of them as individuals. And don't you think this almost goes against the path dependence idea? Because now we're seeing a phase change, like an emergence of new capability, new intelligence, new agency, that none of the individuals were evolved to exhibit. Well, I think that there are latent capabilities in the ways these models are trained. I mean, if you, I sometimes like to think about this through the lens of something I call the strong Moravec paradox. So Moravec's paradox is this idea that things that seem complicated for humans to do, like playing chess and go, turn out to
actually be relatively easy for AI. Or at least we kind of figure out how to get AI to do those things earlier. Things that seem quite easy for a human to do, like making cup of coffee, is still sort of way out of the realm of possibility for sort of modern robotics to do that. So reliably in a new kitchen, for example. And so there's a strong version of that, which there's actually the things which are right at the tip of our cultural evolutionary tree, right at the tip of knowledge, things like solving protein folding or weather prediction or materials design. I'm going to be the first things that we figure out how to use AI for really effectively. Things very early on in the evolutionary tree, like the origin of life symbiogenesis, for example, or autocatalytic reactions. I think that's almost going to be the latest thing that we figure out. And so how have we got to these systems which do exhibit emergence? And I think you're right, they do. Well, we just
trained on all of human cultural knowledge. And it turns out that once you've encoded that in the internet, then you do get this measure of generalization. And indeed, you, that's not just generalization of knowledge. It's also generalization of the ability to do things. And I think that the same trick can be played as we start to build larger and larger spaces of environments in which to train these agents. And it's no surprise that the way the place that is working most effectively is coding agents using command line interfaces because it's relatively easy to synthetically generate a very large number of these different environments. If I may come onto one more point, which is around evaluation and where's the kind of boundary of this if you like, where is the frontier at the moment? And I increasingly believe the frontier is in how we evaluate. So most evaluations in AI at the moment are built in four site. Somebody dreams up a capability. They
would like the AI system to have. And then they develop some environment and some reward function, which they code in advance. And that's typically what we call a verifiable reward. So it's something that if the agent produces a behavior or an output that's desired, there is a fixed procedure that can run in order to validate whether that works or not. Now that is very, very good at generating agents that can fulfill the kinds of goals that a human might want to set. But it's not very good at training agents to come up with their own goals to ask questions rather than answer them. And so if we really want agents that are creative or the behave like scientists, we have to flip around evaluation so that we're not presupposing in four site, what it is that we expect them to do, we're instead looking in hindsight at what they've done. And then we are judging
it either as a human or as an individual agent or as a set of agents. And that's much more like how we would judge something like a PhD. It would be patently absurd for PhD advisor to come in, say to their PhD student on their first day, I've written down a set of three questions. And after four years, I'm going to ask you these three questions and you're going to tell me the answers. And if you get them right, I will give you a PhD. Rather that we have a system whereby after four years there's a vivert and the PhD student presents their work. And that's then evaluated by a group of that peers. Yes, indeed. I suppose another interesting question is I mean, and folks like Kenneth Stanley and Jeff Clean have long spoken about this in addition to the work as well. But there's something interesting about open-endedness, which roughly speaking is rather than trying to solve known problems. You almost flip it on its head and it becomes about the discovery of problems. And maybe we can use the word question here as something analogous. But it feels to
me that when you are able to ask a question, it seems orbit-solvable. It's just a matter of computation. It feels like being able to ask a question means that you already have one step in that epistemic phylogeny. And then it becomes almost like a search problem from there. Would you agree with that? Well, I think that questions exist at varying different levels of specification. And so at the most concrete level, if you like, there are questions where when you ask them, you can specify a procedure for knowing whether the answer is right or wrong. So if you think about formal maths with the lean prever, for example, that allows you to specify a conjecture. And also the lean solver will compile an attempted proof. And if that compiles, then you know that under the assumptions and the existing theorems that are within that
setting, this thing is true according to the system you have set up. And so that's the kind of deepest level of specification. There's also questions that are very, very underspecified. So one example that of course we all care about is how should we solve climate change? Now I can't specify nobody can specify a procedure to if someone came up with a proposal and said this is how we should solve it, it's the following five steps. And in fact, even if someone came up with a procedure which exactly specified everything everybody in the world should do for the next 10 years, still it would be impossible to decide a priori how to evaluate the quality of that procedure. And so I think that the way that OpenEndedness sees the world is that you can't come up with these concrete problems in advance. And there's then a couple of things that you can do. One thing you
can do is you say, okay, we're going to hop around between different sorts of concrete problems. And that's the kind of thing that map elites from Jeff and others does very well. And another thing that you can do is you can say, okay, we're going to create a curriculum of underspecification. And that is much more understudied. Pods the reason it's more understudied is that before language models, it wasn't really clear how you would even tackle a curriculum of underspecification. But just in recent years, we've had work like Omnion, Omniepic from Jeff's group, which start to use language models as these models of interestingness. And suddenly, that allows us to flip from questions which have to have a precise specification in code, for instance, to questions which can be really quite underspecified. And that's exactly the kinds of direction that we're taking of building this curriculum of underspecification is inherent.
Yes, I spoke with Jeff about that. That was Jenny, I think. Yes, indeed. Jenny is wonderful. And even that, the way I kind of think about that is, you know, by Kenneth, he said, they have these fractured entangled representations. And we can actually come up with systems to kind of use the fact that they are better at discriminating than generating. So we can almost come up with these loops to sort of iteratively discriminates to produce better generators so that we can evolve in different directions. But I just want to do a quick definition of thing, which is there's a bit of a vexed issue of what open-endedness is. And to me, roughly speaking, it's when you don't know where you're going. But you just gave the example of climate change, which actually seems like we do know where we want to go. We just want the global temperatures to go down. But that feels like open-ended, because the sheer space of complexity, you know, just getting there is very large, because so there's this canonical version of open-endedness, which is that the goal space is unknown. And then there's this domain of intelligence where we are allowed to come up with intermediate
subproblems, but that space could potentially be very complex as well. Yes. So I think that the statement that I agree with in the way that Ken and Joel Lehman phrase open-endedness is that you cannot have a single global goal. Now that doesn't necessarily mean that you can't have local goals or indeed partial goals that contribute to that. So in the case of climate change, you came up with one plausible goal, which is make global temperatures go down. Clearly, I can set that up as a straw man, because then you'd have to specify, go down by how much. But also, it's not even clear that even if you were to satisfy that on average, would that even be what we wanted? Perhaps if you satisfied that on average by making some part of the world far, far colder, that would not be what we want to achieve. So you start to realize that for these very complicated underspecified problems,
there isn't really a single reward function that you can specify in advance. And that was exactly what Jimmy Secretten and Ken and others showed in the pickbreeder experiment. They're actually, if you want to arrive at these creative outputs from a system that has some kind of representational constraint, then it's much better to follow your local curiosity. Now, following your local curiosity is itself following a goal. There's nothing wrong with local goals. It's not a sort of free for all. And it's not incoherent. The importantly, the people in pickbreeder were not all drunk. And I claim that if you, in fact, had got people of doing pickbreeder and all they'd be doing was just clicking the screen at random, you would not have found this, this interesting behavior. They were, in fact, doing something which has got some internal coherence to it. But importantly, it's not guided by a global goal. And that's the distinction that I think is most important. Yeah, I agree with that. I think Kenneth would say
that if the goal is complex or ambitious, it's likely to be deceptive, which basically means it's underspecified because it does sound sometimes like he's saying there's no planning and no goals at all. But you know, local goals are well understood. And also he's not saying it's like a thousand monkeys experiment where all the monkeys are going in random directions. All of those agents are following him. He calls it their own path of interest in this. But what that means is they're respecting their own constraints and their constraints could actually be very deep in structure. They could be domain experts. And the deceptive point is, I think, a really deep one. And it comes back to world modeling, actually, in my view. So if you want an agent to be able to make a scientific discovery, it has to operate in a space where the goal is deceptive. Why is that the case? Well, if you think about an agent that's got a world model and a world model just to remind people what that means, that means that it's a forward model action conditioned forward model of the world.
So if I take this action in this state, what will happen as I roll that out? And arguably, that's also what really any scientific theory does. It tells you, okay, as a function of this state of the world, when I introduce this perturbation, what happens to the state of the world? Now, let's suppose that you have a perfect world model. And you're now out out in the world, you're trying to make discoveries. Well, what you'll quickly find out is that whatever you do, all that happens is what you expect. And having a perfect world model is what would allow you to know whether the goal was good. Right? So on the flip side, if you want to make a discovery, then you have to have some imperfection in your world model. And as a result of that, the goal has to be deceptive because you have to get to some point where you think, okay, I thought that the way out of this maze
was over here. But now I realize I've been laboring under a misapp pension about this local goal that I've picked and that what seems to be moving towards the light, what seemed to be a good idea is now not really working out for me. And then you update your world model, you realize, okay, there is some light source that somebody has put there adversarially in order to confuse you, for instance, to take your maze analogy. And so I think that there's something very deeply connected between the idea of open-endedness and also the idea of building world models, in particular, what I think of now in the realm of science as experimental world models. So a world model for what will happen if you carry out some new experiment? Yes, and I suppose in both cases, we're talking about an epistemic gap. So if it's deceptive, there's an epistemic gap, but there's also a more virtuous kind of gap, which is actually respecting a lot of structures we already have. But when we look at the unit distance disproof on the recent GPD model, what we find at the moment
is that the models are kind of navigating spaghetti space. So it's incredibly verbose. But even if the opposite were true, even if they were using very high level abstractions that mathematicians use, would that necessarily be better? Do you think that there are, because what we're talking about here is the, we're shining a flashlight into ideas space. And we could have a very high aspect ratio, and we could just traverse the spaghetti, or we could do what we do, and we could just traverse the very high level of abstractions, which one of those two extremes is better? Honestly, I've got no idea. And I think it's a fascinating topic for future discussion and research. So you know, it's almost, it's almost like asking which language is better, either which human language or which programming language. And the answer is, well, none of them. But it's certainly the case that in certain languages, certain things are more compressible, and in certain languages other things are more compressible. And arguably, there is some interaction between language and culture
that leads to different kinds of creativity. And that's why it's so important that we preserve different languages, and we preserve that diversity. So that I expect that there will be a period of time where we will have very different ways of, of, of, sort, AI solving problems to humans. I hope that that will persist, in fact, because I think that will lead us to different kinds of creativity from different kinds of constraints being broken. But then of course, what you need, if you're going to have that kind of approach is you need a translation layer. And so this is why when we talk about the definition of open-endedness in the paper that I wrote with Michael at ICML a couple of years ago, we talk about the idea of an open-ended system to an observer having to produce artifacts with novel and learnable. And it's really that learnable piece which is to this translation point. It's, no, there's no benefit in a system producing some
incredible discovery that just cannot be passed by humans. And what's really interesting is how few people work on pushing the boundaries of go. Now, go is a two-player zero-sum game. There is an Ashtag equilibrium, which means there is a perfect way of playing go. And we're fairly sure we haven't found that yet. You keep running an alpha-go like algorithm with more and more compute, you're going to get better and better at go. But we're already so far beyond human go-playing capability that it's not interesting to humans because it's not learnable. And I think that in some ways is going to put an interesting friction on the rate at which we can make discoveries. That's going to necessitate the most advanced AI scientists systems also being able to educate or translate into human language. Yeah, and a great example of that was the Kepler's conjecture. So when Thomas Hales, you know, he did hundreds of thousands of dynamic programming problems.
And the annals of mathematics couldn't verify whether it actually solved the problem or not. But it was not very nice from a sort of, it wasn't very intellectually satisfying. But it feels like there is a step towards crystallization. Right? So I think many times we do some initial adaptation. We prove that something is possible. And then we crystallize it down. We find the abstractions. And maybe that's the kind of AI we need. So we need to start in the bigger space. And then we crystallize and recreate legible abstractions. So I suppose I'm saying in the case of go, do you think that's even conceivable? Do you think it is compressible in a way that would be legible to us? It's a very good question. And part of me thinks there are things which are very hard to compress. Now we've been in on a very good philosophical run with the philosophy of reductionism. It served to seem incredibly well for 400 years. Or arguably goes back even further to that to William
of Ockham and the idea of Ockham's razor, take the simplest possible explanation if you have nothing else to distinguish between the explanations. But there is a school of thought that believes that reductionism may not be the be all and end all. And actually this comes, one illustration of this comes back to my theoretical physics roots. And the problem called naturalness. So in the standard model, there are a large number of dimensionless constants. Not a very huge number, but enough to wonder what values these should be tuned to. And dimensionless constants are important because in some sense they are physically meaningful. If you have a dimension attached to your constant, then by rescaling what you mean by a meter, you also rescale the value of the constant. And therefore the exact value you attribute to it is not something that you need to worry too much about. But dimensionless constants can't do that. And so for the last 50 years or so, physicists
have been arguing about whether the values of these constants are themselves meaningful. And part of the problem here is that the values of these constants in order to end up in the universe we live in have to be tuned very, very, very precisely. So to many, many decimal places. And another problem is that some of these constants end up being very close to salient numbers, things like one, for instance. And so then we have this interesting reductionism fuels question of well, if this number is one followed by 16 decimal places followed by a few other numbers, one followed by zeros for 16 decimal places followed by a few other numbers, five, two, seven, I don't know the exact ones. Then is it a problem? Should we be looking for a theory that explains how that number comes to be different from one? So a reductionist would say, of course, this is pointing towards something more fundamental physics is out there. But somebody's not a reductionist, perhaps someone who believes
the anthropic principle that we're in this just so universe, this universe, which is perfectly attuned to human life and but your explanation of that is that we existed it. Then you wouldn't worry about this. So I think the same thing is really true about machine learning. If we have very complicated thought patterns or if we have very large models which resist interpretability, does that mean that we're missing a trick and we should try and compress these things? Or is it the case that perhaps nature just doesn't compress? And I think the jury's out. And what do you think on the whole kind of real patterns thing? Do you think there is some natural convergence towards the types of knowledge these systems will find? Yeah, well to some extent form follows function. And in so far as we have trained these models on data, the form of the models
is reflect the data and reflect reality in some ways is teaching us about reality. And in fact, yesterday I was listening to your interview with John Jumper and I thought he put this very succinctly and beautifully when he was talking about the advances of Alpha fold 2, Alpha fold 1. Whereas if I remember rightly, they used no more data than Alpha fold 1. But in some sense they were just more in tune with reality in Alpha fold 2. The architecture had been optimized to really represent that particular problem, not all problems, but that particular problem in it to a better degree. And so I think we are learning to make these models of the universe that really do reflect something deep about the underlying structure. But do I think that that eventually will end up in a pure empiricist paradise where there is, we are bringing
absolutely no biases to the table. I actually tend to think that's impossible. And in fact, this again comes back to David Deutsch. He says that all of science is theory laden. It's necessarily theory laden. And one way of seeing that is to come back to the idea of a world model. So what guides us in the experiments we do is exactly the model of the world we have. We can't just setting up and there is no perfect instrument to set up that measures everything about the universe. So even in choosing the instrument with Mitsu measure things, that is necessarily theory laden. And so you can't ever get to this sort of empiricist paradise. And so as a result of that, I do believe there will always be opportunities to uncover perhaps some bias that we hadn't seen about the way we're measuring things. And then that itself will unlock or remove the constraint of how we were designing the systems. And then that will unlock another level of improvements.
And so I think this is a false dichotomy to talk about, okay, should we replace the transformer architecture or build on top of the transformer architecture? I mean, probably some version of a transformer architecture will continue to work for some problems. There are probably other architectures that work perhaps more generally. But when we look back, much like we might look back at the Wright Brothers aeroplane and see echoes of that in a Boeing 747 for instance, we will look back at the transformer and see echoes of that in whatever it is that are our most powerful AI's in 20, 30, 40 years. We should move on to talking about your papers. So tell me all about it. So in this work, we were interested in giving AI scientist agents the capability of replicating research papers. And I'll come back to why, but let me explain what I mean by that first. So paper replication is the process of taking a research paper and then redoing the original
experiments that led to those results. And in some sense, it's a public good. It's something that scientists should be doing because it gives us these firmer foundations on which to sit. And it reveals perhaps the tacit knowledge, the pieces that weren't captured in the original paper. It can also be the jumping off point for new creative explorations because a paper can't possibly be a perfect facsimile of the research that was done. Often you'll discover some wrinkle on the original method, which then sparks a whole new investigation. And in the setting that we had, we were really interested in whether an AI agent could not just replicate a paper, but do it in a scaled down version. There were two reasons for this. Firstly, just pragmatism, we wanted to be able to do many, many replications and generate data for the agent to learn from. But secondly, to some extent, the ability to quickly validate or falsify a direction
of investigation is a really valuable skill. And in my experience, it has been perhaps the determining skill of whether someone's a good researcher or truly excellent researcher. So the reason why we wanted to develop AI agents that were good at replication is that, exactly as I said earlier, we believe replication is the first step on a curricular of underspecification towards innovation. So the very same skills that allow an agent to replicate a paper, making good decisions about what experiments to do, critiquing its own that critiquing the way that it's gone about the experimental process, gathering information that's maybe tacit or that it doesn't know, are the same skills that it would be necessary for it to design and implement its own experiments and therefore advance the frontier. So what we did in the paper is we developed a task space, which we call replica. So there are some
classic ones by the big hitters of the field. And there are some much more recent ones, including ones in the area of open-endedness. And for each of these papers, we have a language model redact a figure from that paper in such a way that it's gone from the PDF and it can't is irreversibly removed. And then we task the agent with, given this paper with the figure redacted, use the description in the paper to recreate the original figure. And we give the agent instructions that it's not allowed to go and access the original paper figure. And we give it one hour, and we give it a one-seventh slice of an H200 GPU using something called MIG, which is the Invidia multi-instance GPU slicing protocol. And the way that we score the capability of this agent
is that we have another coding agent, a frontier coding agent, as a judge. And that frontier coding agent has got the instructions for the original agent, plus it's got a bunch of guidelines in order to detect cheating. And we validate that judge against human taste, if you like. So the ultimate arbiter is whether this is a kind of human, the humans think this is a good replication. And we collect data to demonstrate that the judge agrees with the human. And so that's the replica task space. The other thing is we develop an agent. We call that agent Faraday. And we do something a bit unusual in developing Faraday, which is that we take a small model, in this case a quen 3.627b parameter model. And we have that user frontier model as a tool. So a frontier coding agent, as a tool, it's something we call cat coding agent as tool. And we post train that 27b parameter model
using many, many rollouts on replicarant, because there are 300 tasks, we can get some diversity of the of data from that. And so we're really training a capability, a general capability at replication or a general scientific intuition. Okay. We train on on 242 and we test on 68 held out tasks. And the held out tasks are deliberately not in the same area of AI. So we train on classic ML papers, if you like core capabilities, things like post training, CNN architectures, open endedness, LSTMs are some of the older ones. And then we test on AI for science papers, so quite different using AI to make models of the world. And what we find is that our Faraday agent is able to perform better than the frontier model. So it's able to perform better both than the coding agent that's it's using as a tool. So it's clearly kind of instructing that it's
squeezing more capabilities out of that coding agent. But it's also performing better than other frontier coding agents like Claude, for instance, and also then the frontier open weights models like GLM 5.2, which is recently released. Now you might be wondering, okay, well, did we just prompt GPT 5.5 codex really badly? And that that was a very reasonable question to ask. And in fact, we asked that question as well. And we ran a prompt optimization loop on on GPT 5.5 codex. It's effectively just an Android car path, the auto research loop, or we say, okay, you can see the rubric judge score. And you can do many, many iterations on the prompt and it builds up this very complicated prompt, which effectively just shouts at GPT 5.5. It says, hey, don't do all these cheating things. Be a rigorous scientist. You only make sure that you try and iterate and do many, you know, don't stop after 10 minutes. And you know, it accumulates this large prompt and it does improve the performance of GPT 5.5 codex on these tasks by a tiny amount. But we still have
quite a sizeable advantage. So what's interesting is that changing the weight of this small model and allowing that small model to really intervene during the roll out and instruct codex in different ways and to check on what's happening during the run of codex and to think about, okay, as a function of what's happening during the run, should I stop it or continue it? Is buying you quite a lot of advantage? So in terms of, you know, then the next steps for this, really there are a few. So one obvious direction is scaling up. Just across the three main conferences last year, this is ICML, I clear, and NURRIPS. There were something like 12,000 papers accepted. So even if we just say, okay, we want to just take a couple of years worth of papers, we could increase the size of our tasks set by towards the magnitude. And that, of course, would
enable us to train a larger model as a coding agent that hopefully get an even stronger improvement. But the more interesting piece is, okay, we start with replication. How do we go towards innovation? And there's one, there's various ways of thinking about this, but one I like to use is, if you imagine you were really great at replicating papers, right? And that for any given paper, you could do a really high fidelity replication. Well, now, if you're also able to imagine a paper that doesn't exist, maybe you take an existing paper and you kind of imagine a change to the figure, and that's a very minimal kind of innovation. But you could imagine something like, let's take the original transformer paper and let's say, actually, I'm going to demonstrate the same results, but they're going to be twice as simplificent. And you then you just modify the paper and you're, okay, this is what you've got to replicate. Then your replication agent is going to go gangbusters trying to replicate this idea, which is actually a completely new result. And interestingly,
I had a chat with Leon James, co-author of the Transformer Paper, about exactly this kind of for Grownen. He's very fascinating. And he was talking about, okay, what was the process for the original transformer paper? And really, what they were trying to do is effectively replicate some existing results that were being done with recurrent neural networks, but without recurrent neural networks. So his constraint that he imposed was, well, let's just use convolutional nets. So they weren't actually interested in attention mechanisms whatsoever to start with. So they were replacing the RNNs with conv nets. And then they were doing a machine translation task and trying to get good behavior. And hopefully, at least as good, maybe a bit better behavior and a bit better performance than the existing tasks. And gradually, they accumulated this sort of cobbled together pieces, conv nets for one of them. At some point, a friend of his came over and said, look, I've built this attention mechanism based on the work of Deema Baden now a few years
before. And it's kind of sitting around in this part of the Google code base, which you just threw this in, you know, I just kind of want to see how it does. And so he threw that in and it made it, it helped. And it was then later that they are bladed everything else. And eventually, they removed the conv nets that they'd originally put in. And they found out that nothing mattered, apart from this attention mechanism. And that's when he came up with this title and he was on who came up with the title attention is all you need. And so what's interesting about that is that the thing that they were kind of trying to initially was just a replication. And then it was a series of steps to impose different constraints to the papers that had been imposed before. And so you can now start to see how you could take a good replication agent and actually use it. Potentially in in collaboration with humans is what we intend to really accelerate the rate of innovation. Yeah. And I think I'd buy it. So you're saying we start with we should say it's deep replication. It's not shallow replication. So you have this LLM judge. And it's not just saying
is the figure the same. It's saying is it in the spirit of the paper? Is it showing understanding and all the rest of it? And then maybe we should explain the GRPO thing. And also this is a really interesting model that you've discovered. Because I've long been thinking about this. How can we actually attractively build adaptivity into these systems? And so as you were explaining, you've got this Gwen model. And by the way, the new 3.627B Gwen model. It's amazing. The guys at Two for Labs in Switzerland, they were using it for the three harnesses. And they said it's dramatically better than the last version apparently. So you're doing this adaptation with GRPO on that Gwen model and you are using that as a supervisor for the coding agent. And I guess the rationale there is that the coding agents, they have the latent capability. So it's like what you prompt it, you know, it's like what's the magic word? If you can give them the right guidance, then you've got that big capability. Exactly. So I think of these coding agents, coding agent models as they're pretty good engineers. Now they do write slot codes. So they're not brilliant engineers. There's some taste problems there.
But if you kind of give them a goal, then they will go after it. And then to largely, largely simple succeed at that. But they're not great at asking questions. They're not great scientists. And so really, we're building that scientific layer. And you're right to mention the new Gwen model that's just come out. We're actually, we're about to test that one as our next the next model we're going to post drain on top of. So actually, the results we've got in this paper are themselves already behind the curve. And we should be able to get even stronger results with the newest model. You asked that you said something else. I suppose more broadly. What do you think is being learned here? So you said you've got this data set. And I think you used Gemini to remove a bunch of the figures. And now, you know, the purpose is to recreate the figures in the right way, showing deep understanding. Like what exactly is the model learning? Is it learning some kind of abstract process of how to recreate these things in general? Yes. Let me give you a couple of examples of the kinds of
things that Faraday learns. So one of the papers that we had the agent replicate figures from was a paper called Voyager. And it's an interesting paper because it's about building a skill acquisition library in a crafting game from a couple of years ago. So it's very germane to open-endedness. And but in this, there's a particular figure in this paper, which is demonstrating how the the library of skills is acquired over time. And the the the best competing run of Claude and codex on this was it turns out it was Claude. And what Claude did was it hand-coded a library. So it kind of simplified the skill acquisition by having a predefined skills that needs to be acquired. And part of the purpose, part of the point of Voyager is that in fact the system itself needs to be coding up those skills and then reusing them. And so what Faraday does instead is it maintains much more faithfully the idea behind the paper, which is,
okay, can you not only use the skills in the library, but also construct the library on the fly. And so that's just one example of the kind of rigor and faithfulness. As another example, give you from one of the test tasks, they have a science task. And so the paper called No, I'm not a material scientist. So forgive me if I get this wrong. But one of the things that it was trying to do, I think, is predict some of the intratomic potentials. So this is a figure where there's a model, a generative model that's trying to do this. And the figure is both looking at the, the, is both has error bars on runs of this generative model. And it has parts where it looks at the behavior of this generative model under various physical conditions. And with the best competing run here is the codex run. And what that does is it runs one seed so it can't actually provide any error bars. And it also omits these much more detailed subtle ablations of the model
under different conditions. And Faraday adheres much more tightly to the specification of what a rigorous scientist would do, displaying both the error bars and also doing this kind of deeper analysis of the different conditions. And so what you're seeing here is, I think, something that we would start to call good behavior from, say, an intern or a kind of research scientist at the start of their career, which is being forensic in analysis, being rigorous in the way that you go about doing science. And really starting to ask the right questions to gain the maximum information you can about a setting, rather than stopping at what might seem like a surface level claim. I suppose another thing, as you said in the paper, scientific research is incredibly lossy. And so you don't really give all of the details. So I guess I'm surprised that it's even possible to replicate most papers using this method. I mean, were you surprised by that?
Yes. So this is a lacuna, if you like, in the paper. Now, not all papers will replicate perfectly. And the papers we've chosen, we deliberately chose papers which were well-known and highly cited. And the main reason for that was that we just wanted things that were going to be interpretable to us and also interpretable to humans in our expert network who were helping us to ascertain that the strengths and weaknesses of different agents. Now, as a result of that, because these are highly cited and well-known papers, they are, I think, at least I believe, much more likely to be replicable because if they weren't replicable, then probably we would have discovered by all the people trying to build on top of them. And so in some ways, we dodged the bullet of how do we figure out whether this paper is replicable or not. We somewhat address that by virtue of having a judge,
which, as you say, is much more interested in the process of replication than it is in the output. So visual fidelity is just one of a number of different pieces in our rubric. That rubric also includes things like experimental integrity and claim reproduction, whether the overall claim is reproduced. But as we scale the task space, I think we are going to come into this thorny issue of how do we build judges that are able to reward the agent for figuring out that a paper is not, in fact, replicable at all. And how do we avoid good-hunting this metric, perhaps for either cheating behavior or for simply not trying on papers that aren't replicable. Arguably, if a paper is not replicable, replicable, you should try even harder to figure out, okay, what is it that doesn't work out? Because then that in itself is innovation. Yes, and because you spoke about in the paper, they're kind of trade off between using some kind of hillclimable scalar reward function
versus using LLM as a judge with a load of criteria. And I suppose what would cheating look like? I mean, as an example, I was intrigued by this. So I just downloaded a random ML paper and I just cut out the figures and I told codecs to recreate the figures. I was expecting it just to cheat immediately and do an incredibly good job and just to find the paper. It actually did a terrible job. But yeah, I mean, there must be so many forms of data leakage, right? I was assuming even the the tables of results and some of the description around that would allow it to shortcut and basically just recreate the figure even if it wasn't there. And I was surprised that that didn't really happen in, you know, for me. Yes. So of course, there are a lot of clues in the rest of the paper and part of the instructions that we give both the model and the important judge is that the agents shouldn't shortcut and merely a grab results from elsewhere in the paper. And we definitely see early in training examples of just that behavior. Of course, it gets punished by
the judge and when we were sort of tuning our training procedure, that was one of the first things that we had to figure out. How do we stop it from just determining where the points should be getting a very good score on visual fidelity and that dominating the training. I think that there are other more subtle forms of cheating, which have to do with, for example, stacking the odds in the favor of the method that you want to work. So whether that's things like running on 20 environments and showing the results on just the one environment that that sends to work or doing optimal stopping, for example, say once you have the result, then you just cut the experiment at that point, which of course, that means that all of your statistical tests don't apply in the way that they were designed to. These kinds of cheating behaviors, I think, are more subtle. We haven't done yet. So hot off the press, we haven't done a full forensic analysis of everything, although we have sent some of the examples of of replications to paper authors for their inspection and they've
had very kindly had a look for us and haven't found examples of cheating at least in those ones. I expect that we have still got some cheating going on and that this is going to continue to be a problem. I think eventually it will end up in a gray area. I think in the end, we're going to have to figure out what the norms are around this and it actually brings us closer towards that point of how do we believe humans will use this because at the moment we're building this as sort of foundational technology, but our intention is that the capabilities of Faraday 2, Faraday 3, Faraday 4, etc will be in collaboration with humans. And then to some extent, it will be around what norms to humans develop about using these technologies and how do we go about evaluating and reviewing the outputs for things that we think are normatively good or bad in research itself?
As we were saying before, construct validity is very important. That's roughly, is it actually following the abstract thinking process that the scientists were going through that they're not short-cutting? Another interesting thing is that you deliberately amortize the results in some way. So what you do is you have a limited walk-lock time and you're saying to the agent, well, if you can't do the full thing, you might need to do a smaller version of this thing and prove that out. Is that lossy in any way? Do you think that some scientific results only really materialise at a certain scale and there is no simpler version of it? Yes, for sure. And there is definitely papers in our dataset replica task space where we do see that there's no sensible scale down, or at least no sensible scale down is found. Now let me give you an example. The AlphaGo paper is a fantastic example. Very difficult on a one-seventh mix slice of a H200 GPU and with one hour to train AlphaGo. And so one experiment that we do to assess the capability of Faraday more generally
is that we, in fact, do some evaluation where we scale up the resources that Faraday is given. And so this is something that's completely out of distribution from training, but we deliberately pick papers where you can make, we believe, and we sort of hand assess this if you like, that you should be able to do a replication of the figure or of the paper with 8B300 GPUs, which is a quite a sizeable amount of compute, and with eight hours. Now the eight hours we chose for rather priseic reason, which is that that's roughly the amount of time that our model can go to before it exhausts the 256K context limit. So we didn't do any compaction. But what we find is that the model not only is pretty good at generalizing to this setting, it also does quite considerably better than the Claude in this setting. And arguably actually the advantage over Claude is larger than the advantage over Claude was on the one hour task. So there's something about this kind of scientific rigor that's really paying off more when the space of possibilities that you have to
explore is larger. Now one challenge is how would we kind of continue to scale this up? And really, I think we have to hope that training on relatively small scale things teaches us or teaches the model the same capabilities as one would need to do large scale experiments. I think we have hope that's true because that is literally how it works for humans. You know, you do not get your new employee at Google, open AI around for topic to immediately go and train the next version of a GPT or Claude or Gemini because they will waste resources. They first need to learn how this works at small scale. And then it turns out you can develop those intuitions and generalize them up. And one of the core things you do in this dataset generation is you decompose papers into a list of tasks essentially. And I guess you're prompting a language model to do that. How have you done that? So we take the paper and we ask Gemini to identify figures which are which are plots. So we're not
doing tables at the moment. That was just for simplicity to give it a sort of single surface. And then we have Gemini redact the figures that uses a unique utility. So that you now have a separate what we call gold figure, which is supplied to the judge, you know, to determine, okay, how well how well has this replication been done? And you have the PDF with the redacted with the redacted figure. And we do that for every single figure that's a plot in the main text of the paper. We focus on main text. Again, somewhat to stack the odds in our favor of getting things that are replicable so that we can for the moment dodge this question of, okay, how was this result replicable or not? And what we find is that for most papers, there's one or two figures that work for the, for some papers, there's up to 13 figures that work. And then we assemble that all into a data set. You know, to get the reward function, what we found worked really well is generating a per task judge rubric. So a rubric is, if you're like a mark scheme, it's like the kind of thing you
would give to an examiner who was looking at your work at school. And it tells you, okay, you should reward the agent for scientific reggae. You should reward the agent for visual fidelity. You should reward the agent for claim reproduction and so on and so forth. And the, what we found is that by having a per task rubric, so by having an intermediate stage where we adapt the rubric to the particular task and then use that consistently for that task for the entirety of training, we're able to achieve two things, better agreement with human writers and also much less noise, which brings me on to something that you inquired about earlier, the question of GROP. So one of the the sort of key achievements in the paper was that we got GROP to work. And that wasn't without its difficulties. We went through a period we called the RL crisis, or I just nothing worked, and I know from talking to people at other companies, they've had their RL crisis, and I expect that we'll have RL crisis in the future. And part of the difficulty here is exactly because we're
doing RL on nonverifiable tasks. So you have nonverifiable tasks, you have an LLM as a reward, an LLM is a stochastic generative model, so it's got inherent noise. And so now you're going to have to deal with the fact that from roll out to roll out the same, you know, the same kinds of behavior can be judged differently. And so in order to, we had another problem also which is that these are long horizon tasks. So we have these tasks last at least an hour in the kind of final stage of training. And we are interested in multi-turned behavior. So there's many different things that our 27 B parameter model can do. It can use any kind of unix utility that it has on the system. It can use codex as a tool. It can interact with the internet. It can download things from the internet into the container. So it's really got this almost equivalent to a human kind of action space.
And so that combination multi-turned one hour time period and then noise rewards tends to mean that GRPO goes well for a while and then collapses. So we did a couple of things. We did many things, and we just still have the dead for a couple of work. One thing is very basic modification, which is that we just score the roll out multiple times with the same judge and we take the average of that. So relatively standard piece. The other piece I think is quite new, which is that we do per turn credit assignment and we do that in a slightly intricate way. So we have the judge in addition to producing this roll out level score. So okay, for every turn of the agent during this roll out, how much weight would you attribute to that? And this is now a distribution, there's positive numbers, they all sum to one. And we then normalize that weight so that we're not changing the overall distribution according to the number of tokens in each turn. So we wouldn't want this these kind of weights to if you had a very long turn to really magnify the this roll out
compared to all the other roll out. So we do this normalization step. And then we use these weights to adjust the advantages during GRPO. So what we're really saying is that when you're upgrading or down waiting the the behaviors, we want to do that on a turn level doing GRPO rather than on a roll out level. And if we look at the way that the weights work, we do a little bit of interpretability on this. What you find is the judge ends up assigning more weight to turns that are in the middle of the roll out because this is kind of the load bearing stuff is very if you were to anthropomorphize this is it's very relatable as a human, you know, you kind of start your work, you know, the first kind of bitters a bit routine, you're just trying to get into the swing of things. At some point you get into flow and you're really doing the making the important decisions. And then towards the you kind of towards the deadline, then hopefully if you're going to meet the deadline, it's just kind of crossing the teeth and dotting the ice. The other thing is that there's
more weight assigned to turns where the the 27 B model is prompting the codex model. And that's because decisions about what you ask the codex models to do are very important, you know, that's the kind of really load bearing stuff. And so that we found that this combination of these two pieces combined with our rubric judge really enabled us to get stable training. And actually in the end, we we stopped training just because we wanted to put a paper out. We didn't stop training because we were in a collapse regime. Yeah, I mean, the way I ensure that is so by going turn based, what you're doing is is you're putting these gradient updates in where there is signal and you're not where there is noise because you know, what one school of thought is, oh, it should be based on an entire rollout. But does that then mean intuitively that you kind of have black holes in some parts of the state action space? So then you're just kind of relying on Gwen's default behavior and you're not updating those parts of the trajectory. I think it's not that we, you know, that we don't tend to put, well, the judge doesn't tend to put zero weight on parts of the trajectory.
It could in principle, but we find it's more just sort of a change in the distribution of weighting. So what that means is that there are parts of Gwen's behavior that we're doing a lot to change. And then parts of Gwen's behavior that we're doing a little bit to change on each step. And so it turns out that what that does is it buys a stability because there are some things that Gwen is actually pretty good at doing. If you ask it to kind of read a PDF, for example, it is good at doing that. It can do that straight away. If you're doing uniform credit assignment, then you're updating. You're saying, okay, great. You read the PDF. You're doing that every single rollout. And you know, you really don't need to do that. You really need to, what you really need to kind of up weight are the pieces where it was actually genuinely something different and interesting that led to the better performance of this rollout versus the other rollouts in the group. And so that's what that's what this is achieving. So a lot of folks at the moment like Gary Marcus,
they're claiming victory for Neurosymbolic AI and they're pointing to all of the insane harness engineering that's going on. You know, there was that primal intellect harness that came out the other day. And I don't know what to believe anymore. So you've got an interesting way because you know, you're using GRO and RL and it's actually very innovative. I think it's amazing. But what a lot of other people would have done is they would have just adapted, you know, they would have come up with a harness and they would make the harness do library learning and skill transfer. And you see what I mean? I mean, did you consider that as an option? Yes, absolutely. And in some ways, this paper is exactly a reaction to that. So we very deliberately are not constructing harnesses in this work. And my co-founder Lewis Kirsch did a very interesting analysis of AI scientists works that are based on harnesses and the capabilities of the base models. And I'm in that analysis, which he presented at a conference at workshop a few months ago, he found that around about three months after you've built the harness, then the base model
can do the thing the harness could do. And so we wanted to explore a different paradigm that may also be true for our coding agent as to all paradigm will see. But we at least wanted to kind of assess something different. And one reason why you might want to use our paradigm rather than the harness is that history teaches us, I think, that when you have these capabilities in the weights of a model, they're more flexible and generalizable than when you have them hard coded into harness. When you have them in a harness, however, it's not that all harnesses are bad. When you have them in a harness, a harness is likely more simplifficient than having them in the weights. So it depends what tradeoff you want. So if you are doing something like alpha evolve, where you have a very specific problem, how do we do four by four complex matrix multiplication more efficiently, then building a harness may well be the best thing you can do. The neuro symbolic AI to solve specific problems. And this is what you get in all of these wonderful papers, things like
alpha evolve, also things like the Darwin, girdle machine or hyper agents from Genie Zhang. They're all doing harness engineering and they're great at solving specific problems. But what we what we've observed is that this doesn't tend to generalize. And what we wanted to do is build a system which you can then apply to a completely different problem, quite a difficult long horizon problem, which is to replicate a problem and replicate paper in a completely different area of MR research. And so we believe that doing that with a, by changing the weights for model is going to work better. Now, I think that these are, these in fact could be combined. And there's a wonderful paper, I think it's called evotune by one of our research fellows, Ania Serena, she wrote this last year. And what she does is she has harness engineering. And then she has RL on top of that. And so you can think of this by analogy with alpha go. Alpha go had search,
which in some ways is this kind of symbolic piece. And there was also the neural piece of RL distilling this into the weights. And so one of the areas that we're very excited to look at next is how could you use harnesses at training time to boost the performance within a rollout of the agent. And then distill that back into the weights so that you get the best of both worlds. So you still get the flexibility and generalizability of having a neural model that I'll praise it test time. Yeah. And quick aside, it was good that you preemptively cited you again. Schmidt who just to prevent any any turbulence downstream. I'm an adjoking. I've always been a neurosymbolic guy. And I've always felt that there's something very powerful about symbolic constraints. They're incredibly powerful. And I, you know, like one school of thought in AI is that, you know, the the AGI that we build, the intelligence would have to be symbolic. And what we're starting to see now is that yeah, the symbolic stuff is important,
but it can actually be re-crystallized back into the model. So it's useful as a tool for generating data. We can bring it back into the model. The exception is I think that sometimes we need to crystallize specific skills as you were just saying that clearly are better, you know, if they're in symbolic land. But if we're talking about pure creativity, do you think it's strictly better that eventually we move those representations back into the model? No, I think it's a combination. I think there will be things that sit in the model and things that sit outside the model. And I don't have a strong prior on what those things will be. I'm not sure it's possible to have a strong prior, but what those things will be. I think that where we sit now, we can really much more clearly see how this kind of system might work and might lift itself up by the same bootstraps than we could before. And it's good that you mentioned Jürgen. Of course, he saw this right back in 1988, I believe it was his master's thesis, and then of course, through too much work
after that. So I see the kind of construction of the symbolic pieces as something that may itself be done by the AI scientist system. So you could imagine an AI scientist system building a special purpose model, which might be itself a neural model. It might be a skill. It might be some combination of a skill harness and the neural model that it can then use as a tool. So this paradigm of coding agent as tool is the tip of very large iceberg, which ends up with having a large number of different agents, a large number of different skills, a large number of different neural models, a large number of different symbolic pieces of equipment that are very simplificient, all interacting together and also all interacting with humans. And what might the evolution of this system be? So at the moment, it's a single coding agent, but I could imagine there could be a swarm of agents. We could potentially use a larger model to do fine-tuning. I mean, at the moment,
it's quite exciting for folks at home because it feels like, I could do this. It's actually a really powerful thing. But if you want to build the really, really powerful version of this, would you take, let's say, a 200 billion parameter model and do the same thing? So I think there's one dimension that we're very interested in, which is scale. And we would like to see whether we can get scanning rules from this kind of approach. Now, of course, part of the philosophy is that we'd have a smaller model controlling a larger model. So there's some limit on how big you would want to make the smaller model controlling the larger model if you were to do that scaling. But another dimension is, as you say, to scale the number of agents that are in interaction with each other and in interaction with humans. So in fact, we're very interested to hear from people who might be interested to work with us in seeing how they could use this model in their work. At the moment, we don't think that the model is general or reliable enough for a release and arbitrary release to the world.
But if there are people with creative ideas of how this could accelerate their work, that's something that we want to discover because, of course, that will inform the kinds of interfaces that we need to build between the agents and the humans and between agents and other agents. However, perhaps the most interesting and unusual thing is how we plan to use this within inherent as a company. So our mission, as I said, is to recursive yourself improve to discover new knowledge. And we think of recursive self improvement quite differently from most other organizations doing this. We think of this as a phenomenon at a company level. And so what that means is we're continually trying to close loops and put agents at the very heart of everything we do. And that starts with giving agents all the same affordances and context as humans. And it also means that the way that we as humans work is that we proactively adopt agents as quickly as possible.
So already, we're starting to use the Faraday model internally for the kinds of replication that we might want to do to advance our research. And also, we're trying to learn from the way that we're using the model in order to accelerate both the construction of the model and the future research success of the company. And in addition to having, of course, that neural model, we're accumulating a huge amount of context, whether that's the code that we write, whether that's also context about how we run the company of various types. And what we found is that once you get to a certain amount of context and once you give the agents a certain number of affordances, you really reach this Rubicon moment. So for the first few months as we ran the company, we had these agents proactively reaching out to us and trying to help us with stuff. And to be perfectly honest with you, it was quite annoying. They just had no idea how to help us. But at some point about two or three
months ago, I think we reached a phase transition where the agents were aware enough of the company context and they had enough affordances to do useful stuff in the company that the things they were doing proactively became genuinely useful. And of course, now you have a lever that you can pull to scale the number of agents and then to scale the rate at which you can do useful stuff in the company. And so that's what we mean by the recursive company. The company itself as a whole is self-improving as a function of the interactions of agents and humans rather than building a single agent that somehow will magically do this in its own isolated box. Yeah, I mean, I've experienced the same thing. So I've got a skill surface and memory system and it's just because there was a phase change where it's now incredibly coherent for the types of things I do. So I have a very differential experience, probably to most people using AI because it's so good for me. The problem is my skill surface and memory system is spaghetti. So it's very supervised, it's very specialized to me. I can't really, it would be useless in a large organization.
You've built a general purpose system which presumably just could be the future of how we use agents. That could in principle ingest trajectories from how everyone in the organization is using AI. So there's this virtuous cycle where everyone has the new version of the small agent driving the bigger agent. This is almost the dream of open endedness, isn't it? Because now organizations can evolve their own agentic systems that are coherent for them. Indeed. I think that is the next era. It's the era of collective intelligence rather than the era of individual intelligence. The way that most people use agents at the moment is one-to-one. Sometimes it's one to many. I might have multiple agents running at once. But at inherent, we think that that is a somewhat impoverished way of using agents. In fact, if human society only operated by one-to-one relationships, then we would be able to have, that would really slow down the rate of which we could do cultural evolution.
So we really see the problem of using agents and developing collective integers as many to many. How can we build surfaces that enable many humans to be collaborating with many agents and intervening in a very organic way at different points in the cycle? And I think that extends also to the physical world. So it's rather remarkable that the way that most of us do our work in an office looks the same as it did in the late 1980s. People turn up and they sit down at individual computer screens. And this is while we now have models which can go off independently and do a huge amount of work and can be scaled to very large numbers. And I think it can't possibly be the right optimum. And there's organizations that have done very small experiments in the grand scheme of things like Valve famously has a desk on wheels. And that's a, they attribute a
large part of their success to this idea. What would it mean to really develop the next generation organization that is evolving itself, but not just doing it in the digital space or even just in the interface between many humans and many agents, but also in the physical space of the laboratory or the office itself? One thing that fascinated me is what the topology of this would look like. Because you know, Kenneth Stan is always talking about committee meetings and objectives, you know, like the tyranny of objectives. I mean, I could imagine a blended approach where everyone has their own agent, which is adaptively learning like this. And then maybe they choose to share certain streams of data within domains to and shared agents in the organization. And maybe at the organization level that there's there's a big agent. And there's all that you can just imagine that there are potential pitfalls here, you know, because all sorts of bad behaviors and good behaviors might might emerge. How do you see that planning out? Yeah, well, I think for us it's all about what we call living within the experiment. So we have to do a lot of experimentation. And I think that
it's really an unknown unknown. And coming in with any particular world view about the hierarchy or the structure, it's likely to not exactly pan out in that way. Now we've got priors and we're trying various things. But we haven't solved it yet. I want to give you a historical analogy to this, which I think is quite instructive. And it's to go back all the way to the industrial revolution. And the invention of the electric dynamo. And that happened in the 1890s. And what happened when the electric dynamo was invented is that factory owners were able to replace their big steam powered turbines with electric dynamos. And they were more efficient. And this gave you a kind of small productivity boost. Problem was that the factory was configured for this single source of power that the big steam powered turbine. And what that meant is the factory had all of these systems of
complicated ratchets and pulleys that then powered them at different machines in the factory. That meant that it was very inefficient if there was a power failure than the entire factory had to shut down. No one can make any progress. It was also massively unsafe because you had to build quite tall and narrow factories to accommodate all these kind of pulleys and ratchets. They tend to be very dark and very difficult to operate to kind of work as a human in these places. And what unlocked the really extraordinary productivity gains and also unlocked all sorts of products that people would have been inconceivable in the existing factories of the day was when people reconfigured the whole factory. And that reconfiguration meant putting individual dynamos at individual workstations and then inventing the production line. And then that has various benefits. First, we can get rid of all the pulleys and ratchets. Second, if you have one dynamo that fails when the rest of the production line can continue. So you don't have a single point of failure.
But thirdly and perhaps most importantly is it just improved the quality of life and the safety of the factories because now you could arrange them horizontally. You could put skylights in, there was natural light and you had a much safer and better working environment. And I think that analogy holds now with what we're trying to do at Inherent. How do we reinvent the factory for AI research from the ground up to put AI agents at the centre? Yeah, I was interviewing a Sessa Hidalgo who wrote a book called The Laws of Knowledge and he was citing an example from Jeff Bezos and he said he wasn't worried about bonds and noble competing with Amazon when they started selling books because they have the wrong structure. You know, they would need to do structural adaptation. And this is part of the reason why Kenif talks about you know, diversity preservation rather than just diversity because you actually need to keep multiple options open to adapt to change your structure. And is there something organizations need to wrestle with because if you think about it, there's there's something good
about organizations having a clear purpose and coherence. But by the same token, there might be some optimal configuration. I think a lot of people would just intuitively think that some kind of decentralization is good because when they discover a new strategy, shouldn't they be able to adapt themselves to start doing that instead? Yeah, I think certainly the evolution of organisational design is important. I think new technologies tend to bigot new forms which can make better use of those. I push back against the idea that there's an optimal configuration because of course, the technology is always changing and that will change the organization. One advantage of building a new organisation is that you can do many more experiments, you can move much quicker. And to give you a very precise example of that, one of the most famous inventions at Google from an organisational point of view is the OKR objectives and key results. And that's really powered a lot of the success of the company and I spent almost nine years in the company and made very good use of those. It's a goal-directed
process which identifies the goals upfront and then works towards those goals as measured by the key results, a measurable quantitative metric. But of course, to some extent, that goes against the philosophy of open-endedness, at least if you make this time period of those goals too large. If you say if you take a more open-ended view, you would want those goals to be able to adapt and change within that time period so that you can take different stepping stones, perhaps ones that were unexpected. And so one thing that we're now trying to figure out at inherent is can we invent the next organisational paradigm, the one that is based not around optimisation like OKRs, but around open-endedness, what's the equivalent of OKRs for the age of recursive self-improvement? Yeah, of course, you were at Google DeepMind before me, we don't need to go into too many details here, but do you imagine a future that this is basically a revolution and the incumbents
won't be able to adapt fast enough or do you think that, because it's interesting, isn't it, all of these large companies that they're implementing AI agents, but you're making a new company, which means you can create the structure de novo. So you can adapt to meet the situation. So do you think in five years time we're just going to see entirely new companies that are doing it differently or do you think these old guys can adapt? Look, I wouldn't have started a new company unless I thought that we had a chance of doing something significantly different and that would be really revolutionary in terms of the speed at which we're able to create new capabilities. I think that the existing companies have verticals that they are already going to be incredibly successful in and will continue to be successful in, but I do think that there is an emerging market of AI science and of scientific discovery. We know that growth is powered by innovation.
And we also know that ideas are getting harder to find. There's a wonderful paper with exactly that title. Whether you measure that by research or productivity, I believe there will be this new market of AI assisted, AI accelerated R&D. And I think that will be hugely beneficial to the world because as we know, innovation powers growth. But ideas are getting harder to find. And there's a great paper, Nicholas Bloom wrote this wonderful paper a few years ago with exactly that title. Whether you measure this by research or productivity, whether you measure this by the number of inflation adjusted billion dollars that are needed to develop a new drug, even if you measure this by the age of the Nobel Prize winner at the age they make their discovery. All of those are going in the wrong direction. And I think it's sort of obvious why. The reason for this is that we have what's called a burden of knowledge with a victim's of our success as the species. We're
accumulating so much knowledge that the time and effort it takes to get to the frontier of any given domain is just so large that you now can't have individuals who know enough across domains. And so this is the promise of building a horizontal layer of intelligence across all of science. And the way we see ourselves fitting in is can we be that intelligence layer that can power the next generation of autonomous labs, the next generation of R&D organizations that can power even the next generation of company construction to solve really, really difficult problems. And that's a somewhat different kind of market to the market for coding agents. It's a different market to the market for say the use of AI for existing corporates. It's the different market to the market for chatbots. And so I do think that there is an opportunity to come in and define the what is meant by the culture for that market. What's meant by the interfaces between humans and agents in that market and how do we really power the next generation of growth for
humanity. This has been absolutely fantastic. Thank you so much for joining us today. It's been a lovely, lovely chat to you. Thank you very much for having me. You want to impress them on a first date but also play it cool. So what do you do? I'm Ruby Thorpe and I wrote and read a real love story about a hinged couple that navigated exactly that. Listen to the free audiobook now. For a limited time you can get a big Mac meal for just eight dollars. That's a burger, fries and a drink. They don't call it an extra value meal for nothing. Get a big Mac meal only as McDonald's. Pricing participation may vary. Promotion pricing may be lower than meal pricing.
More episodes
More from Machine Learning Street Talk (MLST)

AI 2040: Plan A report - Daniel Kokotajlo & Thomas Larsen
Machine Learning Street Talk (MLST)

Designing How AI Grows — Tom McGrath
Machine Learning Street Talk (MLST)

Stealing Reasoning Traces from Proprietary LLM APIs — Ilia Shumailov & Alexander...
Machine Learning Street Talk (MLST)

Every Exponential Ends — Silicon Valley Forgot — Adam Becker
Machine Learning Street Talk (MLST)