
AI Just Democratized Filmmaking (w/ LTX Co-Founder)
About this episode
In this episode, we sit down with Yaron Inger, co-founder of Lightricks and LTX, to explore the future of open-source AI video.
LTX-2 is currently the #1 ranked open-source audio & video model on Hugging Face — with over 4.5 million downloads in just two months.
But what makes it different?
It runs locally.
It can be fine-tuned on your own IP.
It integrates into real video workflows.
And it might change how filmmaking, education, and creative work evolve in the AI era.
We talk about:
• Why open models are catching up to Big Tech
• How smaller models are getting better through distillation
• Running AI video on consumer GPUs
• Infinite, autoregressive video generation
• AI teachers that change environments in real time
• Whether AI will replace filmmakers — or empower them
If you care about the future of creativity, open AI, or the economics of filmmaking… this one is worth your time.
Check out LTX: https://ltx.io
LTX-2 on Hugging Face: https://huggingface.co/Lightricks/LTX-2.3
LTX Desktop Repo: https://github.com/Lightricks/LTX-Desk
For more practical, grounded conversations on AI systems that actually work, subscribe to The Neuron newsletter at https://theneuron.ai.
Get every episode summarized
Each time The Neuron: AI Explained publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
602 searchable segments. Every word is indexed and playable.
Full transcript
The Neuron: AI Explained — AI Just Democratized Filmmaking (w/ LTX Co-Founder). Machine-transcribed; use the interactive transcript above to jump the player to any line.
We have to build a model ourselves because we can't rely on anybody else. And like this model has to be like a foundational model and it has to be big and it has to do basically everything in video generation. The fact that everyone can create now means that, you know, the content that will shine is the actual content that, you know, when you look at like the basic, you know, storytelling, the ability to convey, like feeling, we realize that, hey, this is going to make itself into production. And moreover, it's really going to change the way that content is going to be produced. Having this flexibility, having this ability to do, like hybrid approach of running locally, maybe upscaling the videos on the back end if you want to find tuning, like keeping the IP safe. Like these are things that you can only do with open weights models. We want to democratize creativity. And this was our mission from day one. Welcome humans and interested lobster butts to the neuron AI explained podcast.
I'm Corey Knowles and joined as always by the man who is part human and part RSS feed Grant Harvey. How are you, Grant? I wish I was part RSS feed. It would be easier to write the neuron every day. It would be so much easier, right? It is downloaded. I also like that part that you mentioned about lobster bots. Do you think our audience is getting a clawed bots in it now? I regularly end my day to a brave browser on my computer with videos opened. And I assume that's part of a task I have running. But it always creeps me out. Just a little. It's all right. Well, before we get started, please take just a quick moment to like and subscribe to the channel. So you can keep up with all of the latest conversations with AI movers and shakers that we bring on just like today's guest, who is the CTO and co-founder of light tricks makers of the LTX and LTX two open source video models and a little more to come later. Yarninger. Yarn. Welcome to the neuron. It's great to have you.
Thank you guys for hosting me. So if you could, could you just give us a quick, like overview of yourself and how you started at the company, you know, when you co-founded light tricks and how that led to LTX? Sure. So I was a PhD at the Hebrew University of Jerusalem together with four other co-founders. We were doing, you know, image processing, computer vision, machine learning, but you know, the old style, the things that, you know, now foundational models like do all in one. And basically, it was the, you know, end of 2012. Instagram got acquired by Facebook, the ones who remember for the outrageous sum of $1 billion, which is now, you know, which now sounds like paintings. And, you know, we were saying, okay, so we're doing this state of the art,
you know, computer, computer vision, computer graphics, research. And hey, we can also like create these apps that, you know, back at the day, it was mobile, the mobile platform was very, very popular. And we said, okay, we can just like, you know, leave the academy and just, you know, go through the garage and start building stuff until we figure things out. And then we released Facetune, which become a massive hit. Got, you know, highlighted by Apple and we managed to bootstrap the company for two and a half years until we got our first fund raising. That's awesome. At 25, we were 30 employees, basically completely bootstrapped. And, you know, from back then, we were building a lot of content creation apps, video, photo, got plenty of awards, both on Android, iOS, from Apple, one Apple design award,
Apple app of the year twice. So, you know, we were very deep in, you know, building consumer apps, understanding how creators want to create digital content on this platform. And mainly focused on bridging the gap between, you know, this is our vision. Like bridging the gap between imagination and creation. So, understanding how to basically manifest your, what you think into, you know, the digital world, which is something that, you know, we worked very hard on and worked very hard on how to build these interfaces that, you know, you can use very easily, but produce the results that you actually want. So, have the very tight control over the results and also worked a lot on the, you know, more the back-end side of things, like the tech side of things on how to make things very, very fast, especially when they run on mobile phone that are disconnected from, you know,
from power outlet and we were very much focused on running all the things on mobile. So, on the mobile device itself. So, it runs very, very quickly. But then, you know, the years passed and about three years ago, we, you know, we were tracking everything. We have a pretty big research team in the company and we were tracking everything that's happening with GNI, but at some point we really saw, like, especially with Dali too and stable diffusion, right? It was around, like, mid-22. We saw, like, this big leap in quality of images specifically that are produced and we realized that, hey, like, this is made, like, this is going to make itself into production and moreover, it's really going to change the way that content is going to be produced, right? Because now, you know, we worked and, you know, as PhD students, right? We worked very, very hard on very, you know, fundamental and basic
problems in, you know, computational photography and computer vision and now you realize that, hey, I can just, like, prompt these models and they create these images and later on, these videos, you know, just from scratch, right? So, basically, we started, first of all, to deploy all these models to our apps, to our users. But eventually, we also realized that two things are that are bigger than that are happening. So, one, we were building all of our software on top of models that were open source. But then, you know, everything with stability happened and they basically completely, like, pivoted the business and they might, like, left the CEO position back then and he told us, like, listen, guys, we're not going to train a video model that we thought that is coming and going to be open source. And on the other end, all the companies started, you know, OpenAI weren't open anymore, right? And all, like, the big tech as well,
like, started, like, closing the doors and stop publishing papers, right? And so, we realized that, hey, like, if we want to compete, like, we have to be there. And basically, that led us to, and so this is like one thing. The second is that initially, we thought, okay, let's build a model, but let's build a model that's very, you know, specific. So, it's like, let's take a niche, like, face to you, for example, which was, like, most popular app, right? And let's build a model that's very specific just to retouching and not, like, doing anything else in this world. And we thought that, hey, like, this is going to be an easier task to build a model that's very specific, but very quickly, we realized that the opposite is actually true. So, and this is maybe a bit counterintuitive, but we also see this today, like, with the, you know, the large language models and other models. And the fact is that as you train on more data and you train bigger models,
then they actually perform specific tasks better. And so, eventually, this led us to say, like, okay, we, first of all, like, have to build a model ourselves because we can't rely on anybody else. And second, like, this model has to be like a foundational model, and it has to be big, and it has to do basically everything in video generation. So, this is making a foundation model. Yeah. So, so this is where the world, you know, like, took us. And this is the effort that we're focused on right now with LTX. So, we released LTX 2, which is the number one ranked open source audio and video model two months ago. We've more than four and a half. Yeah. It's really good. So, did you, did you play with it? I played with LTX 1 sum and just, and was very impressed.
And I've done a limited amount of tinkering with LTX 2, but it is really good. I was, I was very impressed. Like, the first one I did was I took, I, I play a lot with AI images and try to make these wild kind of surrealist artists sort of things. And, and the first thing I did was took one of those. And there's something really cool about an image that was just a thing in your mind that you put into words. And then you go and you generate an image and you take that image and you throw it into this tool. And suddenly, it's a living breathing thing in front of you. And I was very impressed. I'm happy to hear that. And I think that, you know, we got actually, it's really exceeded our expectations like this release. And you, we, we now have more than four and a half million downloads of the model two months after, from Hugging Face, got a ton of traction in the community.
You see like the, the hunger and the need of, you know, everybody in the community to have an open source model that they can, you know, tinker with that they can run on their own hardware. And I think this is the, you know, the big deal about this. And this is why it's very, very important for us to release this model as an open source one. So, first of all, we really want the community. We want to democratize creativity. And this is, this was our mission from day one, right? This is why we started creating, you know, apps for consumers and try to bring like these capabilities to everybody. And we don't want like these models to be limited via APIs, right? And limited by, by the big corpse, we want you to have the ways of the model and be able to control it. We want researchers to be able to perform research on the models themselves and not just on outputs of the model. And for that, you actually need to have the ways of the model,
you need to have the code that runs it. And that's super important for us. And we really care about, you know, seeing like papers written on top of the model and helping people from the academia and from the community who fine tune the model for their own needs. So, by fine tuning, it means that you can take the model, right? And you can take your own data and you can make the model learn new concepts that the model didn't see in training. So, it can be like motion style, it can be visual style, because the model is always also doing audio, right? So, you can train the audio, you can train characters with their audio. So, these things are highly important for the community, but also for professionals who really care about training the model on their IP. And because the model runs locally, it means that you can also run it on your own network without exposing your IP to third parties. So, you can run it on your own, you know,
even consumer hardware GPUs, like 50, 90, 40, 90, and even people in the community manage to run it on like 30, 70 or even, you know, lower-end GPUs. These guys have a lot of time to try, like, and tinker with things, and run it on YouTube. But you can also run it, you know, on your data center, local GPUs that are disconnected from the internet if you want to. So, having these flexibility, having this ability to do like hybrid approach of running locally, running on maybe upscaling the videos on the back end, if you want to find tuning, like keeping the IP safe, like these are things that you can only do with open weights models. You just can't do them with with close source, and we believe that the future of AI in general is hybrid. So, you're going to have models like the big brains, big models that are going to be on the back end,
right? And you'll access them through the internet via APIs. But you are going to have, you know, smaller models, or, you know, just to generate like lower resolution images and videos locally, and then you'll send it to the back end, which will also really improve the, you know, energy consumption, because we already have like these mobile devices, laptops, desks, apps, and so on. But also really improve the operational expenses. So, you don't have to pay every time you press generate, and I think this is going to be a very big deal. Yeah. You said so many things in there that I'd love to wrap a hole on. But the first one that I think the people listening are really going to want to know is when we say we can run these locally, what mechanism are we using to run it? Like, like you mentioned a couple different options. Like, is there a tool that you use? Like, is there a specific setup that you recommend? And what can they expect when they try to set this up?
Cool. So, currently, I think there are two main options to run the model. Obviously, when you release things, you know, to the community, there are a million, you know, options, and people build a lot of stuff on top of that. We officially support two options. One is running the model through ConfUI, which is a very popular open source project, which is like a graph based editor that you can use to run the model. And second is that we have our own codebase, LTX2 codebase, that is currently mostly like built for researchers and for developers who want to build on top of the model. So, one of the demo applications, the applications that we're going to release very soon, and when this podcast will be released, this will already be available, right? It's called LTX Desktop, which is a non-linear editor that we vibe coded very, very quickly that
builds on top of the model. So, thank you about having some. Oh my gosh. Now, everything is vibe coded, you know, so we polished it a bit, right? So, actually, it's a funny story because we've had one of our marketing creative guys. He created a lot of content with LTX2, and then he said, okay, but here's the tool I want to build for myself, and then he basically built that, right? Which is, it took him like a day to build like a non-linear editor that looks like, you know, premiere or avid, but all the generations can happen locally. So, you can use the backend, but the focus of the of this application is running the LTX2 locally. So, you can basically, you have a timeline and you can generate videos, you can generate, you have like multiple modes to run the model. Maybe you can talk about it a bit later because I think it's way more than
LTX2 video, you can do like audio to video, you can do first and last frame, like key frame interpolation, you can do many things. And it also has some nice interface to work with the model. So, for example, if you, you want to create like a clip from given like a first and last frame, it also knows how to generate a prompt given this first and last frame. So, you don't even have to write something yourself if you want to. And then you just press generate and it creates the video and everything runs locally. So, you basically don't have to pay for that at all if you have the right GPUs. So, this is the, this is the, basically the code base that we provide. And we want as many people as many developers and, you know, now everyone is a developer again because everyone can vibe code, right? So, this is a, yeah, so this is a friendly repo that you can just, you know, you know, start code code or any other, you know, such a application and just like start building on top of that and use the model and running locally. You should know
with my first reaction when I saw the pictures of LTX desktop was they've got a heck of a team building this. Yeah, no, Corey was saying he's like, oh, this has got to cost money, right? They can't support this. So, it's like two, let's say two and a half people or two people who polished this for like a week and a half or something like that. It's crazy, like what you can build. It's insane. It's insane. And, and I think the, you know, we now realize, you know, and you see everything that's happening with SaaS companies these days, right? And like the stocks and how how the, you know, the market reacts to this, what's really, really important for us is to build the foundational layer that, you know, can build applications on top of that because the options are quite endless and they, and they're related to video editing, but also to many, many other use cases that you can use for video models. For example, you know, dubbing, you can do like,
you can do a lot of things related to ad tech and marketing. We have a really cool flow called retake in the model that you can use to basically replace parts of an existing video. So, you can take, you know, part of a clip where you say something and then you can just replace that part and the model knows how to render this part. So, you can replace what you're saying and the model knows how to take your voice signature automatically from other parts of the clip, or you can, you know, prompt to say, okay, I want to look left and not look to the right. So, you can also like do like process production video editing with it. So, so there are many, many use cases, but one, for example, is an ad tech where you can just personalize ads for different audiences after the fact, even after you shot a real video. It doesn't have to be generated. That's a really good workflow, especially for people who work in film, but maybe skeptical of like ad generative AI, you know, for all the reasons that they might be.
That's really cool. Cory, if you have another question, go, but I have like 30. So, I'll let you jump in. Same. First off, the model is really impressive. The desktop version looks, looks amazing. I'd like to know about kind of the decision to make an AI tool feel like what video editors are used to, like like the timeline vibe is there. Everything is looks like exactly what you would expect from a tool like that. And I'm really impressed with the layout and can't wait to get my hands on it. Yeah. I have a lot of ideas. Yeah, but that's the point. I mean, this is one, this is not, I wouldn't call it like how we imagine the manifestation of the model, right? This is one, only one option and what we're trying to show here is how it feels like I think it's extremely usable, right?
But we want to show how it feels to have this local editing approach that is packaged in a way that is specific, you know, to a specific task at hand, which is like a non-linear video editor, right? But because you know, config, for example, is a great application and it's extremely flexible, but on the other hand, it's not very opinionated in a way because you can mix and match like different kind of building blocks together, which is why it's so powerful. It has a learning curve. And again, like you don't have a timeline, for example, there are plugins that provide you of various things, but I think, you know, after a decade of building consumer applications, I think that there is a lot of power of being opinionated in what you're building. You optimize for a specific flow, but you're trying to really nail that specific flow. So this is one example like let's take of like how to reimagine like a classic non-linear editor given that you're
building it specifically for AI use cases and you can run it also locally, which is very important for us also for the experience, but the point is that the options like are endless, right? Maybe you want to have, I don't know, a chrome extension that uses the model for something that's very specific. And I think the cool thing about releasing these two open source and letting builders build on top of that is that I want to see things that I didn't even imagine that people can build on top of the model. That's the, to me, like that's the, that's the fuel, right? That's the really like interesting part of what people can build with it. And the point is that with the model, like I think again, when we, when most of people think about like video models, you think about, okay, let's prompt let's prompt it and let's get, you know, a result. So text to video or image to video, right? You're like, the input is like the first frame of the clip and then we have your input text and we generate
the video. The point with, I think video models is that the, you can integrate with them in like a million different ways. So I gave an example of like retake, right? Where you can just give the video and just like replace part of it, right? And it's the same although it's the same thing. It's just like a different, a bit of a different way of how to integrate with it. We also have flows like audio to video where you can just take an existing audio and we have a lot of people from the community creating new music videos, right? So they, they take a song, right? And then they generate a new music video given that audio track. And the model knows how to do lip syncing extremely well. I think this is one of like the best, like the, you know, really got like critically acclaimed for it, like doing an amazing lip sync. So it works amazingly well for music videos. So this is another way that you can use the model. But you can also use it in many different ways. I think one that's very interesting that is that's maybe surely like not
available through through APIs is the ability to actually replace parts of classic rendering engines. So let's assume that, you know, we were doing like classic 3D modeling, right? So we're doing like the modeling in a low five way. And then we want to click on render and we need to, and let's say we just have like a very basic model, right? And we want to create a full world out of it. So usually what we need to do is, okay, create like all these like small models and like fill the background and then the whole scene and set up the lights and like do a lot, a lot of things that's related to modeling. And then also running like the renderer that runs like for a very long time. And you need to adjust like many settings and so on. What we can do is basically input that part like the initial like modeling part that's very basic, right? As as a depth map or some other type of control into the model. And then basically let the model like generate a video from
that. But it's not only that you can also learn the mapping between the initial like input to the output. So let's say you have we work to like with animation studios, right? And the way that animation studios work is that they have like an input, which is a blocking animation, which is very like rough, low five, like animation, low frame, low number of frames per second, very, very messy, right? And the output you have like final pixel, which is like what you actually air eventually, right? So we had some experiments where we were testing like 10 minutes of this footage, we were fine tuning the model on 10 minutes of pairs of inputs and outputs. And then the model actually learned this transformation, which can take, you know, it's very, very costly to do these things, right? You usually send, you need to do all the in between and many stages along the
way to generate the output. And you basically manage to train the model within like a day of a single GPU how to perform this entire like pipeline. So now you can basically have an animator just, you know, instead of waiting for a month until seeing the final pixel iterating on the first stage of like the blocking animation or even sketch, right? And immediately seeing the final results. So you can actually iterate on the first stages of the pipeline. So hopefully that like explains a bit what I mean when I'm saying that the inputs are not just, you know, prompts or just like initial images. It can be much, much more. And this is also the power of, you know, having the model in your hands. So you can actually use it in any way that you want because creative workflows are very, very flexible and can change between project to project between person and person. Turn some knob, adjust some dials, do a, yeah, yeah. I'm going to throw out like three scenarios and
you tell me if it's possible. And perhaps if it's not, if it's something that you would want to do or that you would suggest others do, we'll just throw it out there and see what's up. So the first scenario is let's say, because I've been seeing a lot of these tutorials where people are trying to teach their agents how to use blender and some of these other tools. Could you take a 3D asset and 3D model or scene that you've created and then generate that and use that work, use a 3D workflow in LTX desktop. The second one is, you know, could you make a screenwriting tool where basically, instead of, you know, using a video editor, you're writing in a script and you're generating local generations as you go. This is something that I've been experimenting with Google, but if you wanted to do this with Gemini, it'd be very, very expensive. So imagine doing it locally on your computer. And the third one is, well, let's say, you know, let's say I don't have an amazing computer. Like, let's say I'm doing all the pre vids that I want to do with, like locally on my
computer, like, you know, I don't know what the max resolution would be, but let's say it's like, you know, very, very minimal. And then I could then once I have the rough sketch of what I want in each of the clips, could I then connect it to the API and then get the high resolution based on that. Yes. So I think all are possible. So for the last, last scenario, then surely this is what we aim for. So basically, the way that the model works is that the space that the model, like, generates videos in is a compressed space. So it's not like raw video. It's way smaller than that. And basically, so basically what you can do is generate a low resolution video and you can do it even today. And we're going to have an API where you can send like these outputs, which are also fairly small to the backend. And then we can do like the the upscaling, which maybe you can do because, you know, you're running on a relatively low and hardware. So you need more memory and so on
and we can finalize that on the backend. So that's definitely achievable. I think that the story boarding, to me, it's more about like running yellow lambs locally because running, you know, image generation and video generation is definitely possible. So this is mostly about running those models. And this is also like possible today, right? And we have amazing open source LLMs now that are, you know, on par with the state of the art. So you can run them on your local machine. Even now, I think that, you know, when 3.5, a lot of people are got really excited about it. And even the small version of like 32 billion parameters, you can run it in smaller one. They came out yesterday too. That was like three at the time of recording, which was like, you know, nine billion to like zero point. Yeah. Yeah. Nine billion was the big one at that group. Yeah. Yeah. So they're even smaller now. They're still really good. Yeah. So I think that overall, like the ecosystem
of open source models are, we see that, you know, at some point, it wasn't unclear whether, like the open source models are going to close the gap with, you know, with the frontier models that are closed source. I think that, you know, we see now that it's definitely the case. Sometimes it's a bit, you know, lagging behind, but overall, I think we're in a very good place there. So I think that all of this, and you know, the things that, the thing is that it's very, very modular, right? So if you just want to run like an LLM, not locally, you can do that and then run the image generation locally as well. So you can, you know, do this mix and matching depending on your hardware, depending on the current state of the art. Kind of where you're sort of workshopping your idea before you're ready to say, go. What kind of idea? I don't know, like, I guess I'm thinking like, you know, the idea of, if you're ideating around, like, here's this character I want in this video,
I want to make that, like, you might use an image model there or specifically image generation to kind of refine what you're looking for before you come in and say, like, you know, bring this guy life. Yeah, I think that this is, this is basically part of what we do in LTX desktop where you can have your own, you know, media library and you can see all your generations. That's why it's also built for AI because, for example, one of the, you know, these are really small things, but when you build your own, you know, your product for yourself or for a specific use case, this is becoming extremely usable. So for example, when you have multiple generations, you just have arrows where you can, you know, scroll for these specific generations and not have like, you know, million like generations like flat in the same like media library. So you need to know which went to each generation and, and so on. And I think this is, this is why it's going to become like very cool because, you know, now you can build the application for yourself, but you can also use
like the same model. So everything that you learn about how to integrate with LTX to, how to use it to produce the best results, you can, the packaging is something that you can very, very quickly create depending on your needs. So in contrast to how we build software, you know, I want to say like a year ago, but it's even like six months ago, right? Where you can now build this thing, you can now build this thing for, you know, almost like a one off like project and then throw it away and build something else depending on your need. What is still going to be there is the model layer because the model layer is something that's, you know, still very, very expensive to build and require a lot of expertise that you, you cannot just say, hey, LLM, like build me a foundational model and then you, you get that. So we do the hard lifting for you guys and now you can imagine and build whatever you want on top of it. That's awesome. I think the other use case that we didn't
talk about would be like, can you actually hook up an agent to this and let the agent made video for you? Yes. So this is definitely something that's definitely on our mind. And you know, agent is like a big word for like, it's like a, it can be like, you know, the use cases again, like or endless, but one thing that's definitely very interesting is how you can use an agent to generate clips that are, you know, multi shots. So everything that you basically said about storyboarding, so you can do it like on the fly with an agent. I think the, the interesting part and like the potentially like missing piece for that is that agents need to somehow like close the loop. And what I mean by that is that, okay, they, you've done something, right? So you've done some action. Now you need to have some score on whether you actually managed to get the good results
or not, because otherwise, okay, so you, you even now can say, okay, generate like 10, 10 videos for me and, you know, just save it and save them somewhere. And the next loop evolve. And the next loop evolve. Yes, yes. So you need to be able to say, okay, whether like the, the result is good or not. And that's really depends on, first of all, like the quality of LLAMs, right? And like, what, how well they understand like videos, because these are now the best models to actually analyze videos and say, okay, it's good it isn't. But also, I think it leads to like much bigger question or like whether that's, there's even like a single objective function that you're trying to optimize. And my answer is like clearly no. And this is where creativity comes into play, right? So we're not saying, by the way, in contrast to, to other, many other companies who build video models, like, this is going to destroy creatives. I think the, it's like the other way around, right? Because what we've seen until today, like with both like with, you know,
the diffusion model, so like image, audio and video models, but also with LLAMs is that these models are not creative, right? These models like learn like a ton of concepts and they know how to mix and match them together in ways that we don't even like fully understand, right? But they're not going to write the next underground song that no one has ever heard that eventually is going to get adopted and becoming like part of them, I don't know mainstream music, right? So, so that's actually where humans like enter the picture and taste is going to become very, very important. I think this is going to be become like part of the IP of everything that's, that's going on, right? So, so you're going to be the one that's actually, we'll need to define even if you use an agent like what is a good output for you? And it may, maybe something that looks really bad
in terms of objective function to someone else, right? Because they like something that looks completely differently, but you need to be, so what we need to do basically is to be able for you to provide like this kind of taste, right? Whether it's through fine tuning or defining to agent like what you like, which is also like I think a big open question like how you even do that, right? So the the the productization like of this is still like an open question and no one is doing it, right? So that eventually we can with high probability generate content that that you like and maybe do like some kind of online learning where we ask, we continuously ask you questions, right? And you answer or you say, okay, like this, I don't like that. And maybe explain even why, right? So maybe yes, we have no questions or not good enough. So, and then we can somehow learn from that. And maybe, you know, this is only one thing that you like. So you can save a preset of that and and use an agent to to generate content. We've yeah, that's the key where it's like I like
this for this aesthetic for this given type of content, but I want to completely different aesthetic if I'm telling a different story like my advertising aesthetic is different from my action movie aesthetic, which is different from, you know, if I want to drama or, you know, whatever else. This all makes me think of Rick Rubin, you know, Rick Rick Rubin, the famous music producer, he's done the Beastie Boys, Randy MC, Tom Petty, Johnny Cash, everyone, you know, he famously doesn't play any instruments. He says he makes the big money for the confidence he has in what he likes. And the thing this makes me think of is I often think of AI as like some kind of a great enabler almost for people like that to be able to bring an idea out in some way that previously wasn't possible, but it doesn't mean it's not still a great idea, great picture. It also replicates the Hollywood, the Hollywood model where they're constantly doing surveys or prescreenings to see how people like getting feedback before they release any movie. And killing good movies. Yeah, I mean,
they do that quite often because it's very expensive, but this, I mean, the coolest thing about AI video generations, if you think about it from like a first principles perspective is it lets it reduces the barrier to entry for anyone who is trying to make movies. It is so I have a screenwriting background just so people don't know. It is so expensive to get anything made like today and and and and it has never been more expensive, even though it has never been easier for us to actually go out and film things, but it's just a little stuff to get a permit for a beach off down here, like somewhere. Yeah, and so you don't think about the, you know, the the the take where you look at this from the negative is it's like, Oh, well, the big studios are going to use this to like compress costs and and create less jobs. But on the flip side, the the benefit is that now anyone who wants to be a filmmaker could create with these tools, which I think is really awesome and you making it open source makes that actually possible. We don't have to pay Google or anyone else to do this. And sci-fi is going to be so great. That's for sure. Yeah, that's for sure,
100%. You know, I think that in in a sense, it's again, maybe a bit counterintuitive, but the fact that everyone can create now and we're we're big believers in democratization of creativity means that, you know, the content that will shine is the actual content that, you know, when you look at like the basic, you know, storytelling, the ability to convey, like feeling, right? And it's funny to say it like in the era of AI, right? Where everything is like, maybe like very mechanic. I think that people still haven't cracked it like in how to maybe the models are not at this point, like good enough to have like these proper expressions that you can actually make you feel, right? But we have some like we have people creating content internally that was also shared that I think it's like pretty good and getting like, you know, I actually felt something when I when I watched some of these of these videos. So I think this is it's actually be, you know, instead of like
paying million dollars for for production, like it's becoming about these core values, which I think do happen and do exist in Hollywood, right? Like there are there is a reason why, you know, these directors, these filmmakers are actually there because they're able to do it. And I think people will just get tired of watching like a ton of AI slop, like very, very quickly because it will become the norm. And models are already managing to produce really high end content, but and it actually makes you, you know, really like think about the story that is being told, how it is being told, like how it is being like directed and all these things that are, you know, up until now, you know, there were very much correlated, like if the story was great, then probably, you know, there was very good production like behind it. And now it's not necessarily correlated, but we will, you know, get through it. And we I think that everyone will,
you know, filter the content that they eventually want to watch and the content that we'll be getting a lot of traction will be the one that is still like about good storytelling, which is something that was that is still that is true before AI got very, very popular. I think it makes it changes the economics because if you look at it before it was this has to make money at a large scale for it to be worth this us making it, even if it's a great story. And this changes the economics were now, even filmmakers that actually are making movies in Hollywood right now can make a movie that perhaps they wouldn't be able to, you know, make or get produced in the current system, but it's a good story and they can go out and film it and they can just fill in, you know, actually still film with real people, but be able to use these tools to fill in the background and fill in all the things that they would have needed the, you know, extra hundreds of thousands millions of dollars to do. And now they can let the story, a good story can live on its own.
Things like expensive retakes over minor continuity errors and stuff like that. Yeah, so it's very cool. Sorry, Corey, you were going to say something. Yeah, I, something I want to make sure we address while we're on here and I haven't really brought it up yet is realistically, what is the, what do I need system wise to run LTX 2 on my computer? And specifically, you know, like HD level, let's say. Yeah, for HD level, so we basically currently the our requirements are RTX 5090, which is like the highest end like GPU for Nvidia, but again, as I said, like people manage and in this way, it will run like very, very quickly. Okay, so you can generate, I think I'm not mistaken, but with the fast model, which is like the still, you can generate like a 720p video of like five seconds in like 45 or 50 seconds, something
like that. So it's not, yeah, it's really not a big deal. Like again, that's why people are so excited about it because again, the model is like the fastest model among the open source models. I don't know about the closed source, right? But I believe also among like the closed source. So because we initially also built it for for scale for, you know, for, you know, massive apps, you know, with a lot of users. So it would make sense like from unit economics perspective, but now it pays a lot of dividends for, you know, for the community, for the ability to run it very quickly. We saw a lot of people in the community that managed to run it on like 30, 70 and you know, we've eight gigabytes of VRAM. Obviously, it will run, it will not run as quickly, right? And it's also like older hardware. We also saw some people managing to run it on Apple Silicon. So running it on laptops and Mac Minis, you know, now there's
the trend of everybody buying Mac Minis to run and open claw and running LLMs locally. So I would say if I'm personally like very much interested in that and what we can run on this hardware. And just because it's like very, very accessible to many people. So this is also the nice thing about open source. We have a ton of people from the community who just like, you know, trying different things. And if you produce something that's of value to them, they will also try to make it run on their hardware. We'll try, we'll tinker with it a lot. And eventually, they also contribute back, which is very, very important to us. And and eventually, I think this is like the, the nice like cycle that we have here, like this flywheel, where we release something, the community makes it better also learns how to use various things, even finds bugs sometimes. And then we manage
because of their input to create a better model. And then they enjoy it and vice versa. So so this is this is a very nice interaction that we have. It is I run most of what I'm doing through an RTX pro 4000, which, which should run it pretty well. Grant, you're an RTX 6000, right, secret? Yes, yes, on my highest, on my highest performance machine. But I also have Apple Silicon. And that's what I use for like my coding work that I do. So, you know, that's like, you know, 20, 24 gigabytes unified. So I wonder what that could, how that could do. So I'm not, so this is going to be like the model by itself is like 20 billion parameters. So if it's like, let's say you have, you need to have it quantized and then also leave some space to
to run it. So, but this, this is the nice thing. Like you can just, you know, try, try to see how it works. Probably there's someone from the community who actually managed to, you know, to run it like that. And there are people who are not just, again, like just vibe coding their own memory efficient solutions without even knowing what they're doing. Sometimes it's like, it's insane, right? Because they don't, they don't say, okay, okay, just make it, they like tell the agent, okay, just make it run on my GPU and then it's doing like crazy stuff, right? I just looked at like their solution, like, I'm, you know, it's insane, like, what it manages to do. And it works just last night. Code says me, I was working on a little project and it's like, I need to install CUDA. Can I, you know, can I go do this thing? I want, I need to be able to better leverage your GPU. And it's like, okay, and it goes, and it just does the thing. And it's a, a lot's changed even in just the last 60 days. It's really, it's a wild time. And you mentioned, you mentioned open models
yourself. Do you have any specific models that you like to use as a CTO, you know, actually building the technology? You know, do you have any preferences for your coding, for your personal use in terms of like language models? So for local models, I played with, with Quen right now, I think it's, it's a really great model and runs locally, like, extremely well. And runs very fast for images. I think the best models right now are both flux. The flux climb is actually a pretty good model and it's very, very small. It's actually an interesting trend to see, you know, also with Quen, obviously, like, they also released like the big model, but with flux now they, and also with the image, which is in another image, another good image model, Chinese one. So these models actually became smaller. So initially, you know, the open source models were, let's say, like, 20 billion parameters. If I remember correctly, more or less that size. And
these models are like six, seven billion parameters. A flux climb has a version of like four billion and nine billion. But they're actually becoming better as well. So this is happening. What's your thesis? So that's, that's a great question. I think that there are two main aspects here. So the first is that the data that the model, the models are trained on, it's just becoming better. You know, cleaner, cleaner, like the captions are becoming better. So when you train, you use like, you have like video and text caption together or an image and text caption. So as LLM is becoming better. So you can also use them to caption the images and videos better, which also improves prompted hearings eventually. So the prompt that you write, like the the model actually adheres to that better. So it's like one is like pure, like purely like just better cleaner data. And as time
goes by, like the teams are becoming better at filtering the data better, understanding how to caption it better and so on. The second part is that the models that we see eventually are not the original models that were trained. So imagine like a very big model. So for example, if we take like the LLM example, I don't know if like the original model that when train was like 30, 350 billion parameters. But let's say that this is like the original one. And then you you basically have a process called distillation where you take this model and take which is called the teacher model. And then you create a student model, which could be in this case like a smaller model. And you you basically let the student mimic the teacher through like another training process, right? And in this sense, you basically create some sort of compression, which in some cases
even creates better results. And the reason is that for large models, what usually happens is that they have a memorization, right? So they actually, for example, if we think about it visually, they actually remember like, I don't know, specific, you know, scenes from the specific clip specific actors. But actually what we want them to learn is how the world works, how the physics work, right? How styles, different styles work, artistic styles, and so on. And by actually creating this compression, where you say, okay, I'm just taking like less like an order of magnitude, less parameters. And I'm trying to teach like a smaller model how the the big model works, how the teacher model works, you sometimes even get better results. Because because you basically take the entire knowledge of the teacher and learn how to generalize the knowledge and avoid
memorization. So this is basically what what we're seeing right now. And we flux, you can also see this because you see, for example, the larger models that are like 20 billion parameters and even more. So you see like the progression there, it's not that they train the model from scratch, probably, they just took and distilled it. That's so awesome. I have one last question about something else before. No, no, I was, you go ahead, I was, I was going to let him go. No, no, I have one more, one more. And then, and then we'll, and then we'll be done. So, so do you have anything that you want to see in the development of video models going forward? I proposed an idea in our prediction episode where I was saying the thing I think would impress me more than HD quality. I think actually the distillation thing you're talking about is really huge. Like being able to do the same quality with much less gigabytes of VRAM. But the thing I really want to see is like almost like real-time generation where you can like draw on your screen and have it like fill, like fill in and generate around you. But what do you want
to see with the models that you're, that you're creating or that you'll work with in the future? Like, where do you want to see this go? I'll go. Well, there are many things. I think that one, you know, you referred to resolution. I think that resolution is key, especially as we cater more for professional use cases. So, professionals really care about every frame, right? They want to look at every frame. They wanted it 4K. They want HDR, right? So, you can, you can actually do, you know, post-processing. So, you need like 60, 12, 16-bit of color, but, you know, real high dynamic range. And they want it to look amazingly well. And I think that, you know, we're working on it, but it's still not like 100% there. So, we basically want to have a process where you can infinitely increase the resolution of a video. This is like the idea, like the, the vision there, right? And this is possible because the model by itself knows how to invent details, right?
And details are, they're actually fractal, right? So, if you learn details in small scale, you cannot, these actually repeat in finer scales as well. So, there are like technical details that basically we need to figure out, but this is really, really important for professional workflows. We want you to be able to generate, you know, 8K videos if you want and you have the budget for it, right? So, this is an upscale like these videos to this resolution. So, this is like one thing. I think in terms of, you know, control abilities is one thing that is also very, very important. We also like work in that direction and we want to have very, very precise controls, right? So, I gave you examples like in this podcast, but you know, can think of other modalities that are really, really important. You know, there are models now that offer like a very precise motion control, but I think it can be better. And we also, we're talking about like emotions. I think
this is extremely important. And as one worked like many years on FaceTune and dealt with, you know, with, you know, facial like selfies and so on, I know how delicate it is and how many neurons we have in our brain just to analyze, you know, it's one's expressions to actually understand like what you're feeling right now. So, every pixel here counts, right? So, this is extremely important. It's just like one example of a control mechanism. And everything that you said about real time is very, very, very true. So, we're actually, because the model is extremely efficient. So, one thing that we're also trying to push is making the model what's called autoregressive. So, having the ability not to generate, you know, five second clip or 20 second clip as we can do right now, but generating infinitely long clip video that just playing. And I think the use cases there,
you know, for filmmaking, maybe it's less interesting, you know, you have few films that are, you know, like one shot of like eight minutes, you know, but it's not extremely common. It's something that you may want to do, right? I think it's very, very useful for, you know, these real time experiences. And I'm not talking about the world models where you can, you know, press the keyboard and move around and like the gaming style. I'm talking about like having these experiences where, for example, I can talk to a teacher, right? And, you know, LLAMs are doing an amazing job today with, you know, with teaching me how to do various things like I post to LLAM like a research paper and I barely read papers today because I just say, okay, TLDR this and then it explains something to me. And then I ask, ask more questions, right? But I think this can completely change the way we educate ourselves and your my little kids are not going to school. So, and I'm thinking like, whoa, they can just like use like an agent, right? That is actually,
you know, a lot of streaming video. Let's say they're going, they're studying like history, right? So you can talk to a history teacher and it's not just like talking to an avatar, like all these avatar models today. You can just talk to a teacher, right? And you see like full body. And then you say, okay, I want to learn about whatever, like World War One. And then, you know, the background like switches to like a scene from World War One. And you can, and you know, the teacher moves around and show you things. And then it switches to like a map of the world. And you're having like this interactive like lesson where you can, where you use like LLAMs and real-time voice models together with LTX2 to generate like this entire experience, right? And I think this is going to be the future of education, right? So totally. And I also like very much like into like education and everything that's happening with AI today is going to completely change. I think the way we learn things. And so this is something that
I think we'll see very, very soon. And again, like the options here are quite endless. They are. This is Grant and I have talked about a number of times too, is the idea that I was like, I want to learn physics from Albert Einstein. I want to, I want to go back and talk to, yeah, I don't know, Plato. I, you know, and, and you bring up such a good point. You could like, go into black, you could like go to a black hole with Albert Einstein and he could explain it to you. Yeah. Yeah. I want to see this side of a tornado without having to die. I want to, you know, but that stuff that, you know, your kids will do that in school. I just absolutely believe that that's not 20 years away, not even 10. No, I can, I can see it. Yes. We're in this awkward, unfortunate position with education right now where I feel really bad for people who are like juniors in high school through like about to finish college. There's this window of people who
are in an education system right now that's like not really preparing them for what's on the other side of it, but it's also because they don't know what it needs to look like inside of it yet. But yeah, the future is going to be really wild when it comes to what you could do in a classroom. It's a think of how many people that would be engaging to versus, you know, a textbook and a black board and a, you know, kids, you just get lost in that system without having an immersive experience like to unlock all this stuff so much. Yarn, thank you so much. I would talk to you all day, so it's grand, but we're so grateful that you came to join us today and had a wonderful conversation is what's the best way to go check out what you all are doing, what you're building and where to go try it. Great, so you can go to LTX.io, check out the model, check out the API, and all the links you can go to hugging face and download the model, and enjoy and give us feedback.
We really like it. And by the time they're watching this LTX desktop will be available too, right? Yes, the links will be on the website as well. Yep, we'll make sure that if you're watching this, we'll make sure you have links. Trust me. And we'll link the repo and everything else, hugging face. Yeah, excellent, excellent. Well, everyone, that's all we have for today. Thank you so much for taking the time out to watch. Please take just a second to hit that subscribe button, so you don't miss the next awesome conversation we have with someone who's building the future of AI. But that's all we have for today. So farewell for now humans and lobster bots.
More episodes
More from The Neuron: AI Explained

BONUS: The AI Starter Kit: What to Try...and What to Ignore
The Neuron: AI Explained

Can AI Really Design New Drugs? Google DeepMind Spin-out Isomorphic Labs Explain...
The Neuron: AI Explained

BONUS: OpenAI Workspace Agents 101: Build, Run, and Scale AI Workflows
The Neuron: AI Explained

How Google's New AI Turns Anyone Into a Music Producer (Flow Music Demo)
The Neuron: AI Explained