
About this episode
a16z General Partner Jennifer Li sits down with fal co-founder Gorkem Yurtseven and Head of Engineering Batuhan Taskaya to discuss what changes when generative video becomes fast enough to run in real time.
They unpack the technical work behind H3 Max, fal’s post-trained version of MiniMax’s open-weight video model, and how combining model post-training with systems and hardware optimization significantly reduced generation time while maintaining quality. That speed has enabled experiments with continuous video, including streams that can remember previous scenes and respond to new directions while they’re running.
They also discuss why the next challenge may be less about speed and more about control, from camera movement and lighting to characters, motion, and lip sync. And they explore what those capabilities could mean for professional creative workflows, where artists and studios need predictable tools rather than simply generating a video from a prompt.
Resources:
Follow Gorkem Yurtseven on X: https://x.com/gorkem
Follow Batuhan Taskaya on X: https://x.com/isidentical
Learn more about fal: https://fal.ai
Follow Jennifer Li on X: https://x.com/JenniferHli
Stay Updated:
Find a16z on YouTube: YouTube
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Show on Spotify
Listen to the a16z Show on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Get every episode summarized
Each time The a16z Show publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
424 searchable segments. Every word is indexed and playable.
Full transcript
The a16z Show — The Next Frontier of AI Video Is Control. Machine-transcribed; use the interactive transcript above to jump the player to any line.
Generative Media is along with the coding agent market. What we call is token market fit. Everyone's waiting for a large consumer moment in AI. I believe H3 Max makes it possible. Were you surprised by the speed up and the gain you could get from post training this model? We have a version called H3 Max Turbo. That's public that can generate like a 5 second video in like 1.5 seconds. From a cost 10 points also like 2x less. People starting creating these beautiful scenes using an LLM model GPD H3 in Blender. And all of a sudden it unlocked the whole new workflow for Hollywood and professional people. We have been very very focused towards speed performance quality and now we have a really good pace model. The next month or two is going to be fully focused on. What happens when AI video becomes fast enough to generate in real time? A16Z General Partner Jennifer Lee sits down with foul co-founder Gorka Mured 7
and head of engineering Bhatuan Tashkaya to discuss H3 Max and the rapidly changing generative video stack. The unpack how post training and systems optimization made video generation significantly faster. Opening up new experiences where video can run continuously. Remember previous scenes and respond to direction as it plays. But speed is only part of the story. They also discuss the push toward greater control over camera angles, lighting, characters and motion. And why those tools could make generative video more useful for professional creative workflows. Welcome Gorka Mured 7 to our podcast again. We did the last one last year. This is long overdue and we have such a exciting model to talk about which is false H3 Max. The day when it came out I was calling it it's really in the league of its own. Like it's so funny to see the benchmarks where you have the dot of this model on the far left or far right and then everything else is on the other hand. And that graph is actually log scale. It's actually further but we had to fit it in.
We had to do log scale. That is hilarious. The time portion, the quality is not. For sure. The internet noticed for sure there are so many of our tweets about it. Like people really played around with this model. Maybe just give us the backstory of what inspired you to post trained this open weight model from minimax and how did you get the quality and speed to where it is? First of all the minimax H3 model is the first truly open source, very capable like latest generation video model out there. So even though we work with some of the other model labs to run inference for them, we never had this capability like had the right to add this capability on top of it. When minimax came up with their very capable open source model that is truly last generation can take references, very familiar architecture to any other video model, we thought this is a great opportunity to go all in and see what we can do. And again we did many different things that we are going to talk about that combined gave the
results that you show on the graphs. But the biggest reason why everything came together for this particular moment was because H3 was the first truly next generation video model that's open source. What is the idea given like thought has been known to be like a general media inference serving platform like what is the idea to get into post training open weight model like you talk a bit about it in the blog of like combining the system work with the model itself. Maybe talk more about the work behind that. Like general media is I would say along with the coding agent market what we call is token market and the way we define it is as can a single person productively spend a lot of tokens and the amount is like 10k a month something like that. So there is incredible amount of demand in the market to generate video to generate many things at the same time and a person who is doing this for their daily job they spend in front of a computer and do this all day long and they spend
thousands of dollars lots of tokens. And since around April the whole industry and fall itself we've been compute constrained. We are growing as much as we are adding compute like there are things we do here and there but the whole industry has been compute constrained and we've always been looking for efficiencies where we can relieve that a little bit so people can use this more. So that has been the idea behind everything we've been doing since April and this just came at the right time because this makes everything maybe an order of magnitude more efficient. So it gives more compute for other models or even like more tokens can be generated using H3 Max. I think Patu and would agree on that. Yeah, like we just from a system wide optimization which is what we have been doing for the past three or four years. You can maybe make the model 2x3x faster while producing the same quality right because it's at 10 of the day same model same architecture.
You have the same constraints you're just trying to optimize what you can get out of the chip itself and there is a roof liner and we have been approaching that roof line more and more especially like lately because our entire team has been focusing on how do we get out more video pixels from a single chip as much as possible. And this new set of like post-training related optimizations with system slash model code design enables us to go beyond that roof line by an order of magnitude. And like we just felt the pressure we have been working on it on top of open source image models before the video models with its one version with ideogram with its one version with flags. So we have been like experimenting with how can we build post-training infrastructure to take an existing model, build kernels and systems design around it to run it very very fast for a specialized version that can beat anything else that we would get just by running the model itself. And you know combination of that plus just like getting a frontier video model on our
hands and all this expertise we were able to call by like an order of magnitude in terms of speed. Incredible let's dig into that I may get some of that numbers wrong but there's like efficiency numbers there's cost numbers there's speed up numbers right not everything means efficiency but it all adds up to be very efficient yeah I guess what is stunning to me is there is like a magnitude lower cost and also much faster I think it was like 35x speed up compared to the original minimum accuracy at the point well at the Elow score you didn't really sacrifice quality so yeah just revealed more of the secret sauce behind of is this more of the type of system work you have done like did you have to do like model architecture change is this is more of that really brought down the cost and latency and how about like the next generation of chips like GB 200 fits into the whole story it's just like compounding the effect of like multiple different optimization variables that we have been targeting the first one is obviously okay you go from like the base model to a model that's like post-trained to be like more efficient for the fusion models this
is just essentially how do you go from running 50 steps to running something like 20 steps right like you're just trying to optimize that pipeline but as soon as you go from 50 steps to 20 steps you lose quality so you need to target in the optimization scene okay I want to improve the quality and then I want to apply optimization so we have like checkpoints of this that is significantly higher quality but obviously slower so what we initially did was okay let's run our post-training on our alp pipelines so that we can improve the models quality and then applied optimizations like on top of it so that the end result gets you the same quality or like even like higher quality than the original model but at the same time you're like an order of magnitude faster so most of the gains come from post-training this model to be like compatible that you can run this on like less amount of steps but on top of that you add like all the kernels and systems engineering work that you do that brings your like hardware utilization from like 30 40 percent which is like standard in like many inference workloads to like 70 80 percent and 70 80 percent on theoretical MFU which is like
impossible to reach so you're essentially at the roof line of what you can get out and then these models are not just like a single or you just give a prompt and you get a video back they're actually pipelines underneath you need to take a prompt you need to like run a lm like a very large lm to go expand that prompt to a format that the model was initially trained at generate the video in the latent space and then decode those latent in back into pixels and then depending on the workload there might be an upscale in component involved so there's like multiple components and every single component by default is unoptimized there's still like lots to be gained there and like we just looked at it from a perspective of we are gonna get the maximum out of every single component this made us go around lm at super high speeds right there is that component but for a different workload this is not like something like an agent decoding lm workload where you have very high cash rates where you have higher sessions it's a single shot give a prompt you get a prompt back and there's no caching you're operating at low batch sizes so there's like completely different set of optimizations on the prompt expression side complete different set of optimizations on the diffusion model completely different set of optimizations on the VAE that you take from latent
to pixels and you just combine all of these to have an effect at compounds from a hardware standpoint going from something like hoppers to black walls you see something like 2 to 3x improvement by itself but from a cost standpoint it's like pretty comparable because the cost is also like in that leak so I would say it only reduces your like wall clock time but not just the efficiency itself but it obviously helps if you want to go significantly beyond real time if you want to generate 5 seconds of video in less than 3 to 2 seconds then you need like some of these like latest generation hardware today to unlock that possibility maybe this is a detail of a question like is the model being served on single GPU or like it is like majority of the video models today run in a single node configuration which is 8 GPUs because once you start scaling beyond 8 GPUs the efficiency gets less and less because of the communication overhead and existing both like the existing minimax H3 and points as well as like other video models are probably getting served at like you know single node configuration same with this it's like running on in parallel across 8 GPUs and do you think there
will be more efficiency gains in there that you can either you know optimize more of the steps in between by sacrificing maybe some of the like narrowed down the user experiences let's say like the different type of inputs and outputs or like as you think of parallelism like is there more choose to squeeze maybe that's the we released a turbo version of H3 max so initial idea was calling this H3 turbo and like we're like we don't want to call this turbo because the quality is like better than the original one right like that's like this needs to signify how good achievement it is so we released H3 max but like a week later we had like you know like our team was like we can run this 2x faster at like 97 percent out of quality like we run a VALS they're like almost the same like there's still like there's like a notice like there's a small noticeable loss in quality but we we have a version called H3 max turbo that's public that can generate like a 5 second video in like 1.5 seconds which is like insane and that's also like 2x the 2x the like from a cost standpoint it's also like 2x less so there like depends on like how how okay you are
with like losing quality you know you can go down and like today these models are so cheap and so fast that I don't think people need any faster or like any cheaper like it's already like at a point where the from a cost standpoint compared to the front unit itself it's an order of magnitude cheaper compared from like a speed perspective it's more than an order of magnitude faster and like you just like enable all the experiences I think we would need to see like I think we wouldn't need to see what else levels that people would need but my bet today is we just need to improve quality more than like the speed at these speeds right like let's fix the speed and let's try push for quality and controllable to off these models which is like you know what we have been pushing in the past 2 or 3 weeks I think controllability is key like when we first did it we did uh text to video and then image to video and then references came later which which adds a ton of controllability and it's basically the default mode how people use these models these days references
and then we are now adding different uh lora's fine tunes of the base model as well we are working on like a lip syncing version we are working on a different different camera angle lora different style lora's uh like again open source adds a whole ecosystem around the model and it really really helps were you surprised by the speed up and the gain you could get from post training this model like like I I saw it adds a little bit of a surprise that like one day I think what's the Saturday or launch the model and the Sunday people put it on a twitch it become a real-time model like what I mean that's the interesting part of like when you release it we did e-vails like we spent a ton of money doing e-vails on our own I don't know like tens of thousands of dollars even and like the results were unbelievable and then like the plan was to just release the model without doing external e-vails and then okay we decided let's let's hold off let's not let's not tell people that this is like
so much faster and so much better before we have some external validation so so we waited like three four days to all these other like e-vail platforms to actually run the e-vails so we matched the results that we have externally as well and that's that's how we launched it because as you said the results are a little too good to be true and it was I guess well you're taking the best surprise that the real-time uh use case that came out of it or what are some examples that you think this model has like unlocked of the experiences this higher models could this happens that fall once in every couple of months where like the whole company gets gets hold of something and and the creativity just explodes and everyone is just working on a new little app or or a different optimization lawyer whatever it might like the whole company gathered around this this model and like some some front-end engineers started working on like interesting applications and we can talk
about our like world model accelerator team which is which is brand new they they started working on the live experience RTC the web RTC live experience so like there were five six different parallel little projects within the company and like I think we broke a record on Slack that day how many messages were signs in the company because like and and like we have a distributed team we have we have people all around the world like mostly in San Francisco but it's it's like incredible when you see like the 24-hour development like when people like work 16 17 hours and then someone else wakes up and picks up that like and that went on for like three four days and that's that's when we released all these projects take me into that it's so interesting because like you imagine like a model or product launch being like planned out like having again like having all these like evil vendors being ready
lined up and you know ship something out and then like you let the the world or the external users take it and then experiment and like build experiences put online yeah it seems like you know people internally who are very creative just like took this job everything they're doing like launched experience that got ready pop on Twitter do you want to tell us about that one yeah of course what one hour engineers Rehan and just just by himself completely and started streaming a live stream of continuous generations of H3 Max from his laptop like he was he's computer his computer exactly he was doing some like prompt tricks trying to keep like a coherent story and then like he started live streaming that on on Twitch in parallel levels I owe famous Twitter influencer at the spot had I had a similar idea and he reached out to us that he has a website ready already he wants to like host the streaming himself and and have have a website that does infinite streaming internally also
we had another team who was working on a continuous version of H3 Max so H3 Max is like the the Rehounds version and and levels IO version were independent clips it's still very fast but the clip starts it ends and then you take the last frame of the of the clip try to put it in the next one and try to create a continuous yes the last there's no memory like the the second clip doesn't really remember anything from the first clip other than the last frame but internally the the ML team was working on a version where it's the the transition is more seamless there is like two minutes of memory so like you're in a scene and when when you direct the the model or someone else enters the room it actually like everyone looks at that person entering and the scene is continuous so internally we were working on that and then another team was working on an experience we called for live for the continuous version so we had three parallel efforts going on that were
all independently going right along to it and these were all like spontaneous like you didn't like you didn't plan for it at all yes exactly and they just became products and experiences in the like in the following days like but don't let's talk about to how how we made them all more continuous that was very surprising to me because I've I've never seen that actually work on on a video model before so yeah so going back we have been like very very focused towards role models and essentially like action controlled or like you know action driven real-time continuous streams of video and the problem till like you know something like H3 Max was quality was not good enough at all it was just like you know you it degraded a lot it didn't remember the past before but we built the infrastructure we built the infrastructure that we can like go stream video have people control it in real-time being able to like multiply sit to multiple people very low latency and at the same time our ML team was essentially trying to
take every single video model and try to apply this set of like optimizations and tricks to okay how can we make this giant instead of a five second video 15 seconds read 30 second video but you were always like below the real-time factor where you know you you were always like you know you like you you never could generate like five seconds under five seconds once H3 Max unlocked that the ML team was like this is insane which are like separate teams internally we have a research team we have a inference team we have an ML team they're like they they saw this and like this is insane we can apply all these like set of learnings that we had in previous models where we attempt to do this very instead of trying to generate a five second chunk that's right to generate you know like a 50 like 10 second video and then the five seconds from previous one is still attended we still remember it and like as the video goes up we can like extend that memory up to two minutes and you need to do extremely clever optimizations because attending to a two minute video is just extremely extremely computing that's it and just like it goes up exponentially because from like a
compute standpoint so like we we did like lots of optimizations there but they're done of the state we were able to okay we can remember back to two minutes which is like generally good enough from a more perspective and then obviously with like prompt tricks you can still continue to see remember more finer the grain details above the two minute mark and you can essentially stream infinitely we capture that an hour from that perspective and then that team just like release that model under h3 max director which is public for people to use and I think it's the only model that can generate like you know up to 60 minutes continuous videos that is action control you can like you know start with a prompt say like there's like an office setting and someone is like you know working and then like 30 seconds later it's just imagines by itself 30 seconds later you can like say everyone walks through the door like it can take the prompt and reflected immediately which is the most one part and the office is still the same office same camera can pan back to the original person and the original person is still there in the same state yeah so you know we released that and it
got like we did this like fall live website to just like demonstrate it because it's like people need to see how cool this is right this is a new technology I don't think people are like really aware and it it could also like very viral but immediately because we also let people vote on what the next action is it was like you know like a formal process or like the chat was controlling whatever's happening which is fun but obviously you know we limited on like the options and then they could pick oh like a banana enters the office instead of a movement and it's like more fun and like people start like you know having this and we start adding more channels and like every channel had a concept there's like a channel very like full chaos there's a channel that's like cartoons from like 80s and like the model is like extremely capable and it just like remembers like so many different concepts and there's like it says like a big big memory from like a style and like you know concepts perspective so it just became like a very fun experience under me's again like there's so many reading promo experiences coming out of this like it's three max director was just another huge surprise to me it's like I found it interesting in the in the
gem media market that you it's not like you know like language model you have like this linear graph of like just continuously up compounding on like you know intelligence capability and so on like feels like in the field you're operating in it's always like a few months of like sort of quiet time but like a lot of things are bubbling but like in a very short period of time like everything first like all these things come in combination come together of like the the base model being good enough like you can get the latency down to the point where you can like references yeah yeah like get the real-time experience but also like apply controllability on top of that real-time experience like this just opens so many you know opportunities of like live experiences where like end user can control what's happening on the screen which is incredible like we have imagined a lot of these experiences but never being able to like really play around with it maybe just like tell us more about what you're seeing from the market of like
how are people using like the director capability like what are you seeing creators are creating that you haven't seen before and what do you think that unlocks as far as you know what people can do with the this this medium yeah it's been like almost three weeks since we released h3 max and already it is the most popular video model on the platform on the file platform by like double almost like more than double and in terms of like volume and so in a lot of other platforms it's also becoming the default model that people interact with because it's so fast so cheap it just makes sense if you if you come to a platform this is the experience that you want to see and so in terms of like popularity and volume it's it's taking over at least from our vantage point and for for max director again there has been I don't know tens of different
versions of these livestreams some of them are still going on and like becoming more and more popular we are trying to work with some like AI I P holders people who have like AI shows on Instagram and TikTok and and do train a Laura on on their style and do a live version of their show so we have couple lined up already so that's going to be very exciting and like the way people like if you talk to a creative technologist prompting with voice has already become something that like they use all the time is like using this proof law or the chat GPT voice mode and now like you can keep talking to the model and it like almost as if it's a real director in a real movie set directing like the camera directing people where to go you can do that and like our creative engineers started using these models like that so we'll see like a lot of interesting experiences
are built as we speak very interesting as in like the video is playing video is playing I mean you are you are like talking to the video and what's what's being displayed changes accordingly yeah that's incredible and talking about like how the the memory piece holds now like are again this may be a technical detail like the capability of remembering what happened in the last scene or in the last couple minutes of scene like are you remembering that through like the frames images or is it like through text it essentially it's essentially there are like it's remembers throughout video obviously very very compressed because you can't attend to the fall video but it essentially knows like most of the like the details happened in the past two minutes from its own generations and above the two-minute mark it has like think of it as like it has evolving system prompts on top of the two-minute mark from two to 60 minutes where it knows like the overall structure overall detail so it remembers like the last few scenes if you think as seen it's like 15 30 seconds then it's remembers like last 48 scenes and then on top of that
like there's like a continuous evolving gradually evolving system prompt that like keeps remembering like the overall coherence of the of the world like everyone's waiting for a large consumer moment in AI now it's like good enough and cheap enough that like a truly novel social AI experience can be built on top of it maybe let's talk more about the economic side of this like what is the I guess one just like talking about a serving cost for like same minutes of video with its three max and how has it changed your thinking around like your footprint of like inventory of chips like how do you want to have like different steps of experiences serving to the end user. But I mentioned this a little bit like everyone talks about how complex the next generation and alarms are but video models are actually very complex as well because the pipeline has different
components and sometimes they require different hardware configuration for efficiency things like that so the if you were to do this even more efficient it's let's call maybe even cheaper we would probably run different parts of the pipeline in in different types of hardware another interesting thing would be to run it on consumer hardware for people to run it in their in their own machines at home like optimizations don't translate 100% but translate somewhat somewhat close to that and then we can do extra work to translate more of it and so doing these optimizations in different types of hardware and combining the pipeline in a way that it's even more efficient I think that's what we are going to do in the in the next coming weeks amazing so you will have people like Rohan that can stream partially of the experience from the computer but also having like the director
and the control plane more living on the exactly on the cloud makes sense so we talk about all the consumer experiences this model could unlock and it seems like botan is happy with all the efficiency like squeeze out of the GPUs now we're talking more about how do we like improve quality and controllability of these models so that like you know the high end of the market the Hollywood creators directors can take this to the next level I saw some demos coincidentally like you know this model came out of came out the same week or week prior to Astro people were combining the blender experience with H3 Max from from fall like talk about how it's going to impact the the Hollywood world you using blender with one of these AI models together is an extremely popular workflow for for professional work basically you you render a low resolution of your scene what you want to do using using blender previous like
non AI technology and then once you add that video as a reference to an AI model you basically get close to 100% controllability and this this is an incredibly popular workflow for the FX artists people who are doing this professionally because they want exactly they want to get exactly what they put in into into the model and as as you mentioned a week after we launched H3 Max people starting creating generating these beautiful scenes using an LLAM model GPT Astra in blender and all of a sudden it unlocked a whole new pipeline using LLAM to create a blender scene and then passing that to the H3 Max model or any video model but it works very well with H3 Max because it's extremely fast and you can like try many things all at once in parallel and that that unlocked the whole new workflow for Hollywood and professional people and it gets you to like
close to 100% controllability as I said we have been very very focused towards speed performance and now we have a really good base model I think the next month or two is going to be fully focused on okay how much controllability we can add to these models so that professionals at studios professionals who want to actually produce like produce content that fits their use case perfectly can leverage these models the team has been working on an amazing you know like a lip synchronization model where you know you can just supply the audio you can supply your you can supply like a video or an image reference and then it can like synchronize the lips perfectly same with like motion controls you can just take a motion of someone dancing and apply to like your AIJ character and it fits perfectly and this like you can like get these results with like basic prompting and you're going to get like 80% 90% reliability what we are targeting is like 99.9% reliability in dot puts
so that you can actually trust the model did every single aspect of this generation perfectly and that's like what we've been pushing one big launch that we had last week was the camera controls which is essentially you can direct very the cameras going within the video perfectly to the to the degree and this is by like describing in the prompt or like generating the essentially underneath you give a JSON of like I want camera at like 0 0 0 at t 0 I want camera at like 90 degrees angle at t1 like you essentially supply a structured structured description of where your camera needs to be at any point in time and then the model is like perfectly conditioned to to regard it as like the only source of truth and it doesn't like hallucinate I'm like where the camera should go and I just like you can essentially reconstruct 3d scenes from a single input like because the model is itself is all very good video model but at the same time you know it's like perfectly adhered to the camera itself and this is because the base model itself already
has the understanding of the camera angle that you can it doesn't respect it's it just under like you need to tune them all you need to tune the model to a single degree and this is like what enables like at large scale you know post training infrastructure we now have the infrastructure to take H3 max at any capability to do it same applies for any new model right there's a new video model we we we we essentially spend most of the time building it as an infrastructure then just like one of training runs so that we can we can build like services around this for not just like you know open source models but for like frontier close source models as well because we see in the market this is like the biggest gap is just how controllable these models are even like first we start with text the video where you put a prompt you got a video back was good but like you never could describe the perfect character for you and we had image the video where you know use the image editing model and then generally like you know the first scene and then the model was like obviously much more fitting but you still couldn't like say oh I want this new character appear at like second three you need to put it to your first frame or like you can't like you can prompt it but
this must never perfect and then we added reference to the video where you know you can provide like an initial starting frame and you can also provide I want these characters with these voices like you know that's also like a big big I'm not very can essentially say this is the this is the voice for this character and now like you know we are I think oh within the scene I want camera to look at this degree at like T zero on camera to look at this degree at like T3 and then we are adding lighting controls where you essentially say where the light is coming from these are all all compounding on top of each other and like we just have the unified infrastructure to just apply to say any more at this point that's incredible this our fastest growing segment and and there there's a lot of noise about how AI might disrupt Hollywood but Hollywood usage was nonexistent a year ago and in in the past year it grew and now it's it's the fastest growing segment like Amazon MGM studios in their conference they released their NARA tool it's it's mostly backed by file infrastructure behind the scenes and we are seeing incredible incredible pool coming
from Hollywood and exactly what they need these like small point solutions rather than generating everything from scratch they want to be able to extend the video a little bit they want to be able to change the camera controls they want to change the lighting and someone has to build these solutions for them what what Hollywood needs and what the creators actually need and what the research labs are working on there's a little bit of a disconnect there and we believe we can come in and do these little post training projects to to close that gap because we work with all the Hollywood studios and we hear from them what they need and these are exactly the things they need these small point solutions that actually make them more efficient push out more more video and AI can actually close that gap very nicely maybe say in a little bit different way like we have been starting at this problem for the last three years as well like we see like companies trying to
like build a you know movie director like a video model like by either pre-trained or post-trained on the video side but what I'm hearing is like different people expressing the way they want the I'll put to come out very differently consumers talk about it and then like read the prompt and generate the results very differently from a Hollywood director which is obviously like professionals want to talk about like you know these camera angles they want to talk about the lighting like you sort of have built a library or like a collection of post training like I'd call it data and toolkits that can apply these any model that you can like grab the weights on so that they are adapted to like a different audience where they can express their creativity in a bit different fashion to control the model when it unlocks a lot of the Kibili underneath and have to problem was
capabilities all these all these models we are solving that the other half of the problem was legal and there is an incident sitting like that we we made a ton of progress there as well we now have a system of people can apply with with their own IP and we unlock their own IP in the models we are going to grow that and that's going to be a very powerful thing we do with Hollywood studios also we now have C-Dance US hosted as well we already had previously other Chinese models C-Dance was the missing part every Hollywood studio wanted us to have it you exhausted now that's available and so there are no no obstacles in front of these Hollywood studios now everything is ready and we believe they are going to 10 x 100x their AI usage in the coming months it's such a exciting world for movie lovers consumers people who consume a lot of video and creative
content and we have our conference general general to media conference next next week this is our second time we are doing it last year it was mostly consumer AI there were like maybe couple Hollywood executives here and there just curious about it and now it's dominated by studios new AI studios who are like offshoot of the bigger studios trying to do like only AI AI shows but also like the biggest of the Hollywood studios are also there because now they have big plans integrating AI into their workflows into their existing systems so you can see the change in the attendance of the conference as well that's awesome well for the audience check out the content coming out of the the gem media conference it's going to be very very exciting and thank you so much for coming in between coming on to our show it's a very exciting time for gem media thank you thanks for listening to this episode of the a16z podcast if you like this episode be sure to like
comment subscribe leave us a rating or review and share it with your friends and family or more episodes go to youtube apple podcast and spotify follow us on x a16z and subscribe to our sub stack at a16z.substack.com thanks again for listening and i'll see you in the next episode as a reminder the content here is for informational purposes only should not be taken as legal business tax or investment advice or be used to evaluate any investment or security and is not directed at any investors or potential investors in any a16z fund please note that a16z and its affiliates may also maintain investments in the companies discussed in this podcast for more details including a link to our investments please see a16z.com forward slash disclosures
More episodes
More from The a16z Show

The AI-Native CRM
The a16z Show

The Age of Body Futurism | Ruby Justice Thelot
The a16z Show

Greg Brockman on Why OpenAI Says We’re Entering the AGI Era
The a16z Show

World Models, Robotics, and the Future of 3D AI
The a16z Show