
Get every episode summarized
Each time AI Explored publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
AI Explored is made possible by:
“Before we jump in today, I've got some real big news. We just announced our first speakers for AI Business World 2027, the two day AI conference running inside of social media marketing world dedicated entirely to helping marketers implement AI.”From the transcript
Want a repeatable workflow for turning a simple concept into polished AI video content using tools like Seedance? I interview Ross Symons to discover a step-by-step process for producing professional-quality AI video, from developing a concept and building key visuals to generating clips with Seedance and assembling them into a finished piece.
- Why AI Video Quality Depends on How Well Creators Communicate With Models
- Develop the Concept Behind the AI Video
- Build Key Visuals Using the Subject, Environment, and Character Framework
- Build an AI Video Storyboard Using Keyframes
- Generate Video With Seedance and Assemble the Final Piece
Guest: Ross Symons | Show Notes: socialmediaexaminer.com/a123
Review our show on Apple Podcasts
See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.
Get every episode summarized
Each time AI Explored publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
545 searchable segments. Every word is indexed and playable.
Full transcript
AI Explored — How to Think Like a Filmmaker: AI Video With Seedance. Machine-transcribed; use the interactive transcript above to jump the player to any line.
Before we jump in today, I've got some real big news. We just announced our first speakers for AI Business World 2027, the two day AI conference running inside of social media marketing world dedicated entirely to helping marketers implement AI. Nearly two dozen speakers were announced so far across both conferences with dedicated AI sessions designed to take you from experimenting to actually implementing. Every speaker I personally recruited every session is pitch free. Go see who's speaking at AIBusinessWorld.live. That's AIBusinessWorld.live. Welcome to the AI Explored Podcast, helping you put AI to work. And now, here's your host, Michael Stelzner. Hello, hello, hello, thank you so much for joining me for the AI Explored Podcast brought to you by Social Media Examiner. I'm your host, Michael Stelzner, and this is the podcast for marketers, creators, and business owners
who want to know how to put AI to work. Today, we'll explore a proven process for creating AI video. My special guest is an AI educator who specializes in AI images and video. He's the co-founder and chief creative officer of Zen Robot, a studio that helps marketers and creators create AI visuals and videos. He's the head educator at Zen Robot Academy. Ross Simmons, welcome to the show. How you doing today? Hey, good new. Thanks, Mike. Thanks for having me. It's super awesome to have you. So let's start with your story. How did you get into AI video? Because I know there's a story there. It's a bit of a story. So I come from an advertising background. I worked as a web developer building websites, apps, anything that clients needed. But working in a corporate environment, I quickly worked up that I wasn't built for that at all. I love working, and I do love working with people, but working for somebody was just not something that really resonated with me. I didn't know that at the time. I just thought I was being full of it. But realize that it was a bit
apart for me to try and go do something by myself. So in 2014, I set out, this is while I was still working to do a project where I dedicated one little part of my day for an entire year to a single project. It was just to see if I could do something for an entire year. We'll dedicate a little bit of time each day to something that idea resonated with me since I heard about it years before that. And I thought, let's do this. And the only thing I really had as a pastime or a hobby at that stage was origami. I was getting into paper folding and doing that. So I set out for the duration of 2014 to fold something, take a photo of it and post it onto social media. I was on Instagram at that time. And just to fast forward through that year in April, so there was sort of four months into the project I quit my job, I became a freelancer. So I was doing free life, I had free lance work coming in. So it wasn't like I just dived out and like, you know, they had nothing lined up. I had some work. But by the end of that year, just through a series of very
fortunate and I guess timeiest events, I was able to build up a bit of a following on Instagram. This is just when Instagram was kind of picking up as a social media platform, which is becoming popular. And I just called it at the right time. There was no algorithm. You just had to post at the right time every day and you had to be consistent. By the end of that year, I'd amassed a sort of pretty big following of over 55,000 followers. And I was making animations at the time too. I was making stop-motion animations and, you know, just telling little stories with these, you know, just trying to get better at origami, but also telling stories. By the time that kind of happened at the end of the year, I was like, well, you know, no one else is doing this. I enjoy it. I've always got my way of development to fall back on, but let's just see how far we can push this thing. And I ended up doing that for over 10 years, I can tell 2024. But always through that process, or always through those years, I was always dabbling in anything technical. So learning new software, just trying to create content in as many
ways as possible, particularly digital content. So whether it was short stories or making images or anything. So when AI kind of re-edited, this was in the beginning of 20, oh, halfway through 2022. The first application I've used as an AI, image generation tool was mid-journey, mid-journey being a very popular tool now. And something woke up in me when I, the first time I was able to essentially code a piece of code or a piece of text and turn it into an image. The conversion from that, you know, coming from a web development background and understanding how tech works, to me, that was like voodoo. I couldn't understand how, like, how can a piece of text create an image? So I was obsessed from that moment on. It's later on that year, chat GPT came out. The API got released. I was dabbling and trying to make applications and just anything I could get my hands on. And at that stage, the tech was moving quite slowly. That was, you would get one new tool every four to five months, maybe. So, you know, as the tools progressed, stable diffusion came out. And then
the video model started coming out as well. And to me, that was another shift. It was kind of like, you're not only making images with text, you can now make videos with text. And I mean, in 2022 and 2023, the video models were like, dismal. It was actually, they were unusable. But there was still, I still knew that this was the worst. There was just something in me that I understood. This wasn't a sort of fly by night thing. There was a lot of money being pumped into it. And I just was guided by interest and curiosity and just kept up to date with the tools. As the tools got better, I got better at using them. I eventually started teaching people how to use them or my agency and marketing connections would reach out to me after they saw I was posting content onto both Instagram and LinkedIn. And they said, cool, would you mind coming doing a little workshop and, you know, just showing us how this stuff works. Fast forward to the end of 2024. I bumped into my, now a business partner. He used to be my boss. And he said, listen, you clearly have an appetite
and an understanding of how this stuff works. I've got some clients that might be interested in a bit of training. And Xenrobot was born. So we've been going for nearly two years in the end of October. It'll be two years. And we do content production. So it's both, you know, creating the content, beard video, audio and images as well as teaching brands and creative teams and marketing teams, how to apply those workflows and tools in their own business. Yeah. And everything you're creating is is fueled by AI images and AI videos. That correct? Very cool. So we're here to talk about AI video today. And you've got a background. You've been doing video for a long time. It's fascinating. Your journey into it. When it's done well, because we've seen plenty of it, not done well. But when it's done well, and we're going to talk about how to do it well, but when it's done well, what is the benefit? What is the upside? What's waiting for creatively minded people on the other side of this, if they pay attention to what we're going to talk about today? You know, firstly, I think that there's the big misconception around AI itself is that it's all AI video, particularly is that
it's quite easy to make. And it is easy. You can type a prompt in and you can make a video. Is that video going to be good? That depends. You know, how good are you at prompting? How well do you understand the platform? So, you know, with that misconception aside, once you do get into what quality video content in the form of AI video content looks like, it's, I mean, the upside is anyone with a creative story or looking for a creative outlet with their usually gate kept by the people that can make videos and animations and creative directors or artists, essentially. Like, if you're not an artist, then you can't create. Now, you don't need to be an artist. You just need to be someone who's technically, well, you don't even have to be technical. You just have to be inquisitive enough to want to understand how the tools work and you're able to create not just single clip videos, but you can start telling stories. I think that that's a massive upside. Creating social media content, as you know, video has and then will be king for a very long time. So creating content for with it's, you know, for social media campaign, you're trying to, again,
tell a story about pretty much anything. You're trying to keep a log of what's going on in your life. I think that it's, yeah, when done well, it's just that it's such a powerful medium. But to do it well is not an easy thing to do. I'm glad you brought that up because I accidentally skipped one of my questions, which is the misconception question. And you brought it up anyways. And I just want to reiterate it, hey, it's not easy to do quality video, but it is possible. That's the upside, right? And it's not just possible. It's ridiculously powerful when you do it really, really well. And I just want to start with like, what do we need to know at a base level or a foundation level before we start creating AI video? Because I think it is kind of a very new thing. I mean, it's like the last real frontier, I think on the creative side of things. So talk to me a little bit about what we need to be thinking about. There is a new way of communicating, I think, with machines and with AI. I use the term machines, but I think it is communicating with computers, you know,
the language which you use and the format of that language is a little bit different to how we speak to each other. But also in, you know, we've been introduced to chat GPT and Gemini and Claude and all of those tools speak English to you. And they talk into like you're a human, and it's very gruffing and it feels like you're speaking to a person that fully understands what you're saying. The reality is it doesn't, but it feels like it is. But when it comes to video and it comes to even images, like I think that there's new form of communication, the ability to articulate in a specific, I don't want to call it coded form, but it is almost in a structured, keyworded, sequenced format, which is something that we're having to learn. It's not a difficult thing to learn. This is something we teach in our master class and to a lot of corporate clients. There's a switch that has to happen and you can try to do it by yourself and you're going to hit roadblocks like you always do. But it's that understanding of that new form of communication
that needs to happen. There is also a level of understanding the difference between what and a diffusion model does. So a diffusion model makes images and videos and a large language model, which is an LLM like it's essentially a chatbot. There's a difference in the way that you communicate with them to get certain results. For example, if you go into mid-June, you say, hey, mid-June, create me an image of a cat walking on a beach with a cowboy hat on. It's going to look at that as a sentence, but it's only going to pick up the visual keywords that it's after. Where with chat GPD, you'll say to it because chat GPD can create images too. You can say to it, please create me an image of XYZ. I wanted to look like this and blah, blah, blah. It will then translate it because it has reasoning built into it, which mid-June has a diffusion model doesn't. Much like some of the video models too, it will then convert it into what it understands and then produce something. And the Jerry's out as to which tools are better is it better to make images with that or that.
It really comes down to preference, but also what's starting to become very apparent is that certain tools are better for certain jobs. If you look at mid-June, for example, as an image generation tool, it's a lot better. I keep bringing up image just so you know, because image is the foundation for video, which we'll get into shortly. And if you understand how to make images, but also how to guide the tool in a way that gets you the image that you want, you have a lot better success rate of creating decent quality video, as well as images in a sequence consistent series. This is really fascinating to me because folks that have been around for a little while might remember when chat GPT first came out with images, you would give it a prompt and it would rewrite the prompt and then you could see how it re-rided the prompt before it generated the image. And the benefit here is this large language model, chat GPT, cloud, Gemini, understands intent and understands kind of what you mean when you say it in like the way you would text somebody,
you know what I mean or in your short typing. But when you go to an AI image generator, it does not understand that because it's got its own method, it's looking for things, especially mid-June. Mid-June is looking for all sorts of weird little symbols and cues and I've had people come on to show to talk about mid-June. So it is absolutely fascinating that the way these things work is evolving. And effectively what I'm here and you say is we're moving towards an era where there's like this layer that sits on top of the tool that's interpreting what it is you mean and then it's creating the final output on the newer tools but on the older tools which might be even better, like mid-June that's not there. So understanding how they work is absolutely essential because what I heard you say is you might need to communicate with these things differently than you would with just chat GPT and that's kind of important for everybody to understand. So did I get that right? Yeah, absolutely. I think there's also another very basic understanding which is if you don't know how to communicate with mid-June or any new tool or all tool that exists, if you know how to
communicate and if you understand the power of a large language model, it gives you the ability to convert what you're actually trying to say. This is when I say it, I mean a large language model like chat GPT. If you tell chat GPT please write me a prompt or I want to create this image in mid-June. Instead of telling it to create the image because I personally prefer from an art direction and creativity and ideation perspective, mid-June is the boss and you speak to most creatives that are using a telegree. But mid-June, like you said, has a specific format that it needs to adhere to in order to produce better results but you don't need to know what that structure is. You can literally ask chat GPT say, I have no clue of how to do this. Can you produce an image that makes, you know, cap wargain to be to the cowboy hat? And it'll break it down into a structured format that is better accepted by mid-June to produce much better results. Excellent. Okay. We're going to break it down now starting now on how to actually go about doing this. And the first thing
that you and I agreed to talk about is the idea of a concept. So talk to everybody about what the heck a concept is because not everybody thinks like you think. We have a lot of great people here, but maybe they don't have like that art director hat on. So let's talk about what the heck is a concept and why it's so important. Well, look, I mean, a concept for anything is really an idea around what story you're telling. Let's call it that. So when you know, story is a loaded word, it's kind of like not everyone's going to be creating a full sort of cinematic experience. Like a Christopher Nolan style film every time you sit down in front of your machine. Sometimes the story can literally be you want to take off example like me. I teach something. You know, I'm teaching AI all the time. But the concept is how do I frame what I'm teaching in a format or a story or a narrative that makes it more palatable and cut through the AI slub. You know, which there's a lot of it out there. And it's just because of the ability to for everyone to be
able to create anything quite easily. So I think the concept is seriously, well, really just the idea behind what it is you're trying to say in the message. It doesn't have to be elaborate. It can be something that just pops into your head. It's literally just like cool. What am I starting with? What is what is my idea? What am I trying to say here? You might not be saying anything, you know, moving or emotional or deeply conceptual, but it can just be like, well, cool. What am I trying to say here? You got an example. One of the animations I did years ago was I used a a Red Bull can. So the types of animations I used to make were stop motion animations. And I'd you know, I just love the Red Bull brand and I wanted to do something. It wasn't a paid concert paid. It was almost like a spec ad. Let's call it that. And the idea was simply the can sliding in piece of paper sliding in piece of paper folds into an origami bull. The bull runs into the can, opens the can. The bull drinks some of the Red Bull liquid wings pop out of the bulls. It's sort of, you know, out of the side of the bull. The bull flies up and off it flies.
Is the whole play on Red Bull gives you wings. So it was a concept within a concept. And what I did recently is I did a side by side comparison of exactly the same ideas, but one was generated with C-dance, which is one of the most popular video models out to the moment. Definitely the most powerful. I kind of reverse engineered what I'd created as a stop motion animation using Gemini, broke it down into what the sequence of events is, fed that into C-dance to create a new animation using a couple of image references. I needed to be origami. This is what the Red Bull can looks like. And the concept, regardless of what it was I was saying, the idea was translated via the AI animation as well as the stop motion animation. And that could have been hand-drawn animation. It could have been a cinematic, you know, like story shot with cameras and real bulls and real cans. So the concept, regardless of what medium is being used to create that, the concept would translate
and be as good an idea, regardless of what you're using. So the idea of a concept really is, you've got this vision in your mind of what it is you want to create and you're effectively describing it in words. In your case, it was a Red Bull can sit on a table with a red piece of paper that turns into origami bull. And then it jams the thing and some Red Bull liquid comes out of it. I mean, that's the concept, right? So anyone who's listening, you might have kind of a vision in your head of something you want to create. It could be an ad. It could be YouTube shorts or reels or something like that. But you got to start with the concept. So once you have the concept, what is the next thing we need to be thinking about? Once you've got the idea and which again, you could use a large language model, even if your idea is quite simple, using a large language model to help you further extrapolate and develop it, I think is essential. And I'm sure some writers will think like, oh, you're being lazy, maybe. But in this day and age, I don't think we're all, again, not we're not all making Tarantino films. It's sometimes you just have a basic idea and a basic concept and you want to get that out as quickly as possible. So starting with, I guess,
once you've got the concept, then it's moving into, well, what are the key visuals that you need? What is the look and feel? What is the art direction? And again, these are all things, as I'm saying, these now, these are all things I would watch from afar as web developer and watch the art directors and the illustrators and whatever in the team of marketing teams that I'd work in. And I was just building the website and watching the stuff come up. But I now know that I can do all of that myself. If I know what questions and what phrases and terms to use to give to chat, GPT to help me develop that. So when we're starting with, you've got the concept, you move to images and the images can be like I said, the look and feel, the art direction, what you want to paper to look like, what you want the bull to look like in its final folded form to just to work on the red bull analogy. And once I've got those ideas, then it's like, okay, we've got our key visuals. What is the sequence step? So this is where a bit of understanding of animation comes in. It's like you have a six or 10 or 12 frame storyboard, which is like frame one. This is what happens. Then frame two, this is what happens. Frame all the way through to frame 12. Not to say that there's
12 frames in the final shot. It could be a 30 second ad, it could be a minute ad. But these key frames give you the angle, the lighting, the look and feel of what each of those frames is. It just gives you a high level idea, which if you're selling the idea to the client or you're just trying to get people to buy into what's going on. You have these sequence steps. And those frames essentially become the starting points for your video. So once you've got those frames. Let's back up the chain just a little bit here. So, okay, I just want to like, I don't want to skim too past, too far past the key of visual stuff because I don't think a lot of people don't even understand how to do this. So what are some of the things that we need to be thinking about for the images or the key visuals because you know, you do this all the time, you might take it for granted. But a lot of our audience has no idea what in the world. And I think maybe you have an example with a fragrance brand. I think that you could share a little bit. Maybe you could kind of walk us through like what those key visuals were so people can imagine them inside their head. To use that example, you know, I've got a fragrance bottle that I want to make an advert for. I want to help put it out on
social media. This is just an example. I've got the product already and there's been a couple of photos taken out of it. So one way of going about it is I could actually create a fake or a mock product using chat. You could do your mid-journey to create the product. So it's created this fragrance brand bottle. Now, once I had that bottle, which becomes my subject or my hero, yeah, you hear a subject or hero object or character. Let's call it that. Then you create the environment. So you'd like, what is the character? What is the focal point of what is the point of the story? Where does this live in the case that I used? I had this jungle environment. I just had this jungle. So I developed this arena. Let's call it that. This environment in mid-journey. Because I know how to use mid-journey better. Not to say that that's the only place to start. You could create images in a chat GPT. You could create the images in Gemini. And you could use reference images that you found on Pinterest or on chat to stock. It really doesn't matter where you draw the inspiration from. It's kind of like cool. This is the look and feel in the same way that you would
say to an art director. Look, this is the mood board. This is kind of the look and feel of what I want the overall piece to look like and feel like. So then I've got these concepts. Oh, I've got my character, which is the bottle. I've got the environment, which is essentially the area that is going to live in. And then bringing in a character, another character almost like a it's going to the, yeah, just another character in the whole thing. And this concept that I did have was bringing this black panther kind of walking in. So it was, so I thought this was playing. And this was really just me roughing with an idea. So, you know, it was a still shot zooming into the fragrance bottle. You have the fragrance bottle in there with the black panther walking into the shot. And black panther kind of finally staring at the camera and then jumping towards the camera. It's sort of 32nd ad. But all of those parts and all of those pieces are developed outside. And then through video concepts using keyframes, you then piece everything together, add some music, add some sound effects. And you've got something that could be quite a
compelling piece of content. I love this first of all because lots of people listening right now have products or maybe they are the character, right? So I would imagine if you are the character, then you could put a picture of yourself in here. Could you not? And you put yourself in a different setting. Have you done that for anybody before? Only myself. But for, I mean, we have, okay. Any tips on doing that with people? I mean, is there any special things we need to be thinking about? I think that, you know, getting a good photo, like a real photograph that's taken, it could just be you, you know, staying like I am now up against a wall with a selfie, you're having a selfie taken. I think it's important to try and isolate the character that you're trying to, don't have it with a lot of people around them or a busy environment. That's just one thing. It's almost, it just helps the AI tool SQL. This is the subject. And the more detail you've got of it, the multiple photos from different angles, wearing the same clothing and using that to feed the model to SQL, take these shots and turn this into a, let's say, like a corporate looking photograph or a
corporate identity portrait shot or series of photos, place this character in the jungle, place this character on the beach, place this character, walking through them all, whatever the case is, which you can do with all of this. So, and then again, this is just developing the idea and the concept, then you see, okay, cool, this is looking kind of like what I want it to be. I mean, what we have to do for our, for a lot of stuff is, you know, obviously under the face of Zen robot, so my face is in a lot of places, but we have one photo that we use all the time. And it's just taking that one photo, placing it into multiple places, like this is Ross teaching at a seminar, this is Ross, you know, against a kind of digital background, and just depending on what it is we're trying to sell, and what it is we're trying to, what story we're trying to tell, what's the narrative around it? Hey, I want to tell you something that's coming that I think you're going to be really excited about. We just announced our first speakers for AI Business World 2027. That's our two-day AI conference running inside of social media marketing world, focused entirely on helping marketers
and business owners put AI to work. Every speaker I personally recruited because they're getting real results with AI right now, and they're going to show you exactly how they're doing it. Quote, the AI speakers were amazing. Every bit of information they shared was completely new and valuable to me, said Sarah Heinbronner. The beginning lineup is now available at AIbusinessworld.live. Tickets are on sale at discounted prices, and here's something worth knowing if you grab a virtual ticket. You get access to the content from AI Business World and social media marketing world. Two conferences, one ticket. Go check the speakers out at AIbusinessworld.live. That's AIbusinessworld.live. You mentioned chat GPT and Gemini. Gemini's Nano Banana and GPT, I think Image 2.0 are as of this recording two of the leading image models. Do you have any tips on how to work with those, either to create environments, or do you recommend having a bunch of them mocked up?
Because not everybody is a mid-journey expert here. What's your thoughts on that? I think that firstly, it obviously depends on the concept, where do you want this person to be, and for really knowing where to get information like that, when I said information, where do you find these images? I always start with, if you don't know how to use chat GPT, not use mid-journey, which is most people don't, it's doing a Pinterest search back to the old school way of doing things, or going through a library of photos that you have taken already. Maybe you're a photographer and you like taking photos of the desert or whatever the case is. You have this bank of images which you can draw inspiration from. You like to look and feel that. But just really just finding reference images. I think that that's the best place, then, saying to collecting reference images of maybe not too many. I don't think it's the more images you create, the better the result is going to be. You really don't need that many. But if you can articulate to chat GPT or Gemini, I like this environment, or not I like this environment, saying what I like about this environment is X, Y, and Z. I like the time of day. I like how this
light falls on the sand. I like how X, Y, Z. The more context you're giving it about what it is you like and what it is you want to see in the final result, the better the actual result is going to be. You say, okay, cool, this is the environment, take this character, which is you as the person that you want to put into the shot, drop that person in and say, place this character wearing a certain hat, shoes, shirt, whatever the case is. And this is the pose they should be in. And it's actually, you know, it seems a lot more, sounds a lot more technical. But once you get into it, you really, and you understand that these models really do understand and they adhere to what you're asking it to do. And it becomes so fun because you kind of like, whoa, I can do anything with this thing. And it's like, I can, you know, change this and do that. And I can move them in this environment. Put them on a boat on a horse in the sky and doing pretty much anything. Interesting. And do you have any recommendations on perspective for lack of a better words because you know, generally speaking, images created by a lot of these air models have typically a direct on perspective, you know what I mean? And like, everything is, it's not isometric, but balanced. You know what I mean? Perfectly balanced. And do you ever
recommend like saying, or I change the change the the camera from up above or down below? I mean, you understand, I don't know what the phrase is, but like, do you have any tips on how to like change the view if you will of the camera? So I actually taught someone how to do this recently. And I think it just comes natively to naturally to me because I love film. Music is probably my first love film is my second. And just understanding a little bit about film, I think helps. And also paying very close attention. This is, I'm just speaking about how I see these things. Like what makes an emotional shot? What makes a character look more powerful? It's like a low angle shot. What makes a character look more submissive or weak? You know, it's a high angle shot. And then even if you don't know what all of those are, again, asking the model to say, look, I love these type of films. This shot to me shows two characters, one of them, one of them seems submissive and the other one seems very powerful. Help me extract from this image. What it is that's making that? Is it the camera angle? Is it the dramatic light? Is it the contrast? Is it the background? Is it the
the soft book, a lighting? You don't need to know all of that. All you need to know what how to do is or what to do. And what to say is to explain to the model that this is what you're trying to achieve using these new characters. Give me five examples of what these two characters in a be it an avant-garde or a, you know, a moody setting of a, like if you if you do know a bit about film and you know what certain directors sort of resonate with you. I mean, everyone knows a little bit about film, right? They know what films they like. So it's like, well, I like this film and I like that shot. You don't need to know anything more than that to say, you know, make this character look like Aladdin in the scene, for example. Oh, okay. And it will know. Obviously having screen or visual references is definitely going to help. But that's yeah. And then Saquel will give me, I did it the other day. I said in a single image, create six of guy Richies most used shots. So
it's like, that was like a, like very close up. There was a high angle sharpen. And there were terms in there that I never heard of before, which you don't need to know. That's the thing. But this is also a process of learning. So as you see these things pop up, you know, like close up is really over there. High angles, that low angles there. And just again, being curious about it and testing and seeing, like, cool. Now you let's go back to, you know, me as the character in my story. It's like, show me a guy Richies style shot of those characters in this environment. And then you're like, okay, cool. This is, this is starting to feel a bit more again. It depends. I mean, you're not going to do that for every social media post or, you know, promo post. But that is how you start expressing these different and getting different camera angles and and feelings in your images and videos. And you look different than what everyone else is generating because you're asking for these things, right? Because out of the box, it's just going to, exactly. Out of the box, it's just going to go for the basic direct shot kind of stuff, right? And the thing is, I think that's why a lot of people think that, oh, AI sucks. But it's doing
exactly what you've asked it to do. It's like, show me a cool photo of this person in a cool environment. It's going to be like, okay, what is your version of what is your definition of cool? I think that also what you learn and again, something we teach all the time is being able to articulate what exactly you want to see in the image and in the video. And less about the feeling, I mean, feeling, I think, or vibe is, you know, important. But it's, if you can direct exactly because you know that to create something more intense, it's going to be a close-up shot. If you want something, if you want someone to look vulnerable, you do like a zoomed-out shot of them in this sort of desolate arena where there's nothing around them. We'll make them look slightly smaller than the people around them. You know, it's these little visual storytelling techniques that help you de-eth that. But only over time do you learn those techniques. But if you're just starting out and you just want to tell like pretty basic stories and get a pretty standard, not standard, but tell a slightly different story to everyone else, ask the model, what is going
to make my content stand out, what is going to make the story more compelling? And test those up and see and then it's a B testing. You run one series of ads, it's like cool, that didn't really work. This one did. Was it the angle? Was it the clothing they were wearing? Was it the pace of the video? Was it, you know, it's all these things, like it's all metrics, right? And that I think just comes with time and iteration. It's really intriguing because I have now a new found respect for people that do storyboards, right? Because these storyboards aren't just the drawing of the character. They're showing them in scene and in a setting and in an action environment typically, which transitions into keyframes. You started there a little while ago and I want to come back to that. I'm anticipating that you could have a bunch of generic looking keyframes, like, okay, I know that I've got this, this can of energy drink, which is red bull. I know I have this red piece of paper and then it turns into origami. You know, I know kind of the basic steps. I could just start out with these images. I could start to imagine and get creative about how could I more creatively use these images? So are the keyframes almost like the storyboard or the keyframes
that foundation for a storyboard talked to me a little bit because as I think about this, I'm like, okay, well, we could change up the perspective of the red bull, you know what I mean? So it can of red bull. So it's scared of the bull, right? Instead of just sitting there in the middle, right? So like, talk to me a little bit about these keyframe things and kind of how all this works. I think that, you know, the example you've just brought up there to make the, if you want the bull to look more powerful and dominant, you want it to be slightly bigger, maybe not bigger than the can, but bigger than what it would be expected of it to be. So how do you do that? I think that with a key frame, essentially a key frame is just an image that you're using as a reference. If you're speaking about traditional sort of key framing or in the AI sort of space, it's a start and end frame. So this one technique, which is the start frame is going to be where the video starts and the end frame is where the video ends. When I say the video, I mean, that could be anything from a five to a 30 second clip. I think if you're going into 30 seconds, that goes a bit much, but if you're showing it, create a short between a little say between four and seven second long clip.
If you think about like somebody filming, you know, cinematography filming, if they're running the camera for 20 seconds of that 20 seconds, then I'd only get three seconds that's usable content. Maybe the actor does something strange of aeroplane flyers bars, but they capture that one little piece and they're like, oh, they've got it. Whether you know much about film or not, it doesn't matter. It's more about you're creating these clips and these clips are broken down into you can do it as a start and end frame. You could use those images as references. So you could say, this is the basic angle that I want it to start out. It doesn't have to be the exact frame. And this is kind of where I wanted to end up as a shot. And then the model, what's amazing about these models is it fills in what you don't know how to articulate. So if you do know what to say, you can. You can break it down into like into each second being like this from zero to one second. This is what this happened. This is what happens from from second two to second five. This is where it happens from second five to second 10 over 10 second clip. It can actually hit those and
it hits them wonderfully. That's getting very technical. But for the most part, you can say, okay, first shot, I want the Red Bull can. You just want to clean plates. There's nothing on there. So empty shot, that's shot one as your key frame. The second key frame will be Red Bull can in the middle of the shot with nothing around it. So that's your start frame. And between those two shots, it knows. If you say, and then using the prompt, you're guiding what needs to happen in the scene. So you have your start frame, which is an empty shot. The second frame where you're at end frame is where the can is. And you're telling it using the prompt to say, Red Bull can slides in from the right at a slow pace and stops directly in the center. Then you've got, it's almost like all three of these things are start frame, the prompt plus the end frame are guiding what's happening in the screen. And you rinse and repeat that throughout. That is one way of going about it. The other way is just dropping. You recommend going about that. I mean, that seems like the logical way to get started, doesn't it? To get started. And if you've never done something like this before,
absolutely, because it gives you a sense of what can be accomplished in five seconds, because sometimes I know people have done this way. They have a start in the end frame and a five second clip is they've selected. You've got a dial that you can select. You can select anything between five and 30 seconds with C-dons, at least. And you'll put a start in the end frame in and they type this gigantic like novel of what needs to happen in five seconds. And the model attempts to do all of it. And then there's hands and fingers and spaghetti and arms and limbs all over the place. And I can't like, oh, this sucks. Why is it not listening? But it's kind of like you're telling it between zero and five seconds. That much must happen. It's logically that cannot happen. So in following through with that process, you realize, okay, so the first five seconds, I can only tell it to do a certain amount. You cannot say, you know, Red Bull can comes in, piece of paper flies in as a jet. The bull turns into paper turns into a bull and things explode. Like there's just too much going on there. But if you break that down into your storyboard beforehand and you have
your keyframes, then at least you are guided. The storyboard is less for the model and more for you to understand what the sequence should be. And then once you've got that sequence, you're like, cool, that's working. Then you start with the start in the end frame. But yes, to your question, I think that starting with start in the end frame is a really great way of going about it. You don't have to have start in the end frame. You can just have just a start frame, which is also handy. So start frame plus prompt creates the video. So it's like you could say empty shot and you just have your start frame, which is your empty plate. And you could say Red Bull can slides in from the left. You could maybe make it a 10 second clip and say Red Bull can slides in from the left. A paper jet flies in from the top lands next to the can and unfolds into a flat piece of paper. Then you've got your first 10 seconds. Cool. Then you don't. Then there's steps and sequences, you know, depending on how technical you want to get of grabbing the final frame from the previous video using that as your next start frame. Okay. So now you building now you've got, okay, cool. So
you've got the first 10 seconds you extracting a frame up from the end frame as the end frame, which now becomes your new stop frame for the second five or 10 seconds on the video. You said there was another way. What is the other way just having a start frame as a reference? Yes. Okay. Got it. All right. You've already mentioned seed dance a little bit. So let's talk about seed dance and any other models that you recommend. There's a really good chance that my audience is not familiar with seed dance. They are owned by bite dance, which is the company behind TikTok. Why don't you tell us a little bit about what it is about seed dances, which is amazing and why you love it. From a prompted here and a reference image adherence perspective, it is, it is near flawless. It is insane. The latest model seed dance 2.5 is like nothing I've ever seen before. Whether you have been tracking, you know, the state of AI video or not two years ago, there was a very popular trend going around. It's said to an hour, a few years ago of Will Smith eating spaghetti, which
did the rounds. It was basically Will Smith just showing spaghetti and tears. It kind of looked like Will Smith, but it could have also been quasi-modo. It could have also been anyone else. And it was just him stuffing spaghetti into his face. And people are like, AI video cool bro, like good luck with that. And if you see the videos now that seed dances able to produce, it's like flawless cinematic. You cannot tell the difference. Like I saw a recent Will Smith eating spaghetti video where they're doing a mock thing like it's seed dance, like almost mocking those old videos. And the real Will Smith walks in eating like a place to begin. And it's fully cinematic. It's in a beautiful suit and he's commenting on, it's like, what's this guy doing kind of thing? You know, eats this spaghetti, puts it down, never fight. And it's just it's insane. And you're really just accept it's not the real Will Smith or is it the real? No, no, no, it's not. That's the thing. And you know, I think that it's a very advanced model. Just a note, a technical note on that if you go in Google, see dance now and look for a website to just use seed dance,
you're going to get a list of about 10 websites. So understanding that, you know, regardless of which video model you're using, pick a platform that has that built into it. So a couple like you mentioned, one is Luma AI. Luma is something I prefer to use. There's another platform which on node-based applications, there's one called Flora and there's another one called Figma Weave. But, ArtList, Imaginart, OpenArt, CREA, these are essentially platforms that allow you to generate everything inside it. So you would pull, you would just have an account on these platforms and you have access to seed dance, you have access to Klingon, Vio and Grock and all these other models. So just in case you are. Just for the record, LumaLabs.ai is that the... That's LumaLabs. Yes. Okay, so they are kind of an aggregator is what I'm hearing you say, right? Like they will work with different tools. Pretty much. Yeah. And this is the key thing, like you can't just download seed dance, like you can, chat GPT is what I'm hearing you say. It's a
model that is powered on the back end by some other tech stack, right? Exactly, exactly. And it's connected via what is called an API, which is just a little connector which connects the capabilities of what that tool can do into this aggregator or platform. What's the typical cost for something like seed dance? And you were mentioning the duration, I'm assuming you were talking about seed dance, right? Like you said, you can do up to five seconds, the 30 seconds or something like that. So what's the... On average, what is it going to cost to produce something like this? So one of the platforms I use actually gives the dollar amount of a 30 second. So you get tokens and the tokens can be broken down into time or actual dollars. And I think a 30 second seed dance clip at 720p resolution was $14. Okay, if I'm not mistaken. So as soon as you bump that 720 up to 180 at doubles, so that would go to sort of 28 between
$28 and $32, which you know, it sounds... I don't know if it sounds cheap or expensive, but from where we came from, where five second clips were used to cost less than 20 cents, it is relatively expensive. But what you're able to create it worth the money. If you know the steps... Honestly, if you've never used AI video before, I would not start by trying to create a 30 second 4K resolution seed dance clip. I think there's going to waste $50 on... Can you upscale it? Is there upscalers? Like if you start with... Talk to us about that. Most of these platforms, I know Lumen does, and I floor and and Figma, we've definitely do too, is they have upscale is built into it. And I actually did a costing today. I did a... I created a 480p clip, which is reasonably small. And I upscaled it twice, I think, so it came out to roughly between 3 and 4K. And the clip itself cost me what was it? $6, I think, and then the upscale cost me $3. So
that's cheaper than what I would have paid the $15 or $14. So that does... Was the resolution good? I mean, was it quality? Yeah, it was good, Ja. Yes, definitely. So the two upscale is the two that are available in terms of image and video. One is called Topaz Labs, which you would see built into all of these platforms, and the other one is Magnific. Okay, some people are going to want me to ask this question. Gemini does video. Does it make sense to ever start with Gemini and then go to seed dance and have seed dance enhance it? Just because maybe some people get some basic video for free, including with their Gemini or what's your thoughts on that? So Gemini's video model, the current one, which is VO3, VO3. You do get free generation. And VO3.1, when it em out, was state of the arcs, was the best. That's before seed dance and before cling 3.0. You know, it's such a tricky thing. It's not like you can practice with VO to become better at seed dance. If that makes sense, it's not like it's a... It is a more superior model, but I think you would
do yourself a lot more favors by understanding the prompt structures for the different models. So I know that seed dance from a prompting perspective, you can break it down. You can say, for a 15 second clip, you can tell it between 0 and 4 seconds, I want this to happen. Between 4 and 8, I want this to happen, between 8 and 12. So I'm always 3 to 15. And you include all your keyframes and you don't even need keyframes. You can tell it in the prompt itself, either use those as keyframes or you use them as reference images. The key difference there is, I think references work better with the model because it's not the... Like the keyframes are like it must look exactly like that at that point. And if you, which most people can't, can't visualize that is where it's going to end at second 4. And this is the first frame. What is happening between those frames, the model is trying its best to get there. And you're going to get weird stuff going on, the camera angles are going to be weird. There's going to be little floaty. But so using it as a reference rather than say,
this is kind of what I wanted to feel like. That's essentially what you're saying with reference images. This is basically what I want the feeling to be like, but I kind of want the person to end at second 15 holding the stick up in their hand or whatever. And then it does a much better job because it's drawing from references that make sense to it as opposed to it being dictated by the keyframes that you're using. If that makes sense, yeah. It totally does. Ross Simmons, we could keep going for a very, very long time. I know you have a class that you're teaching on this. So why don't you tell everybody, first of all, if they want to connect with you on the socials, where should they go? And secondly, if they want to potentially join or work with you or whatever, where should they go for that? So zenrobots.ai. You can... But like everything is there, but what I contact details are there. I am very active on LinkedIn. So my LinkedIn name is Ross Simmons. It's YMONS. And zenrobots pretty much anyway. So if you just Google 10 robots or I think you'll remember that maybe more than my name. But we do have a master class that runs every month. We have
three left this year. The last one will be in November. And that is a it's a seed based four-week program where we take people from the basics of learning how to communicate with AI to make images, move from images into consistency, from that consistency into motion, and then kind of putting it all together in what we refer to as a portfolio really piece. Awesome. Ross, this has been so amazing. I know there's a lot of people that are going to like go literally start messing around with creating model images and go and messing with seed dance as a result of it. So my creative juices are flowing. Thank you so much for coming on and sharing your insights with us today. For sure. Thanks. Thanks for having me, Michael. Hey, if you missed anything, we took all the notes for you over at socialmediaexameter.com slash A-123. Be sure to follow this show on your favorite podcasting app. And if you've been a listener for a while, we would love a review. Also do let your friends know about this show. You can tag me on Facebook, LinkedIn, Instagram, X, and even on TikTok now. I'm on TikTok at Mike Steelsner. Also do check out
my other show, the social media marketing podcast. This brings us the end of the AI Explored podcast. I'm your host Michael Steelsner. I'll be back with you next week. I hope you make the best out of your day and may AI help you become more successful. The AI Explored podcast is a production of social media examiner. One more thing we just announced the first speakers for AI Business World 2027 dedicated AI sessions, personally recruited speakers, pitch free training, designed for implementation, and virtual tickets cover both AI Business World and social media marketing world. See the lineup at AIbusinessworld.live.
More episodes
More from AI Explored

How to Get AI to Recommend Your Business
AI Explored

How to Use AI to Dramatically Improve Your Quality
AI Explored

How to Build AI Employees to Get Work Off Your Plate
AI Explored

Building an AI Creative Director: From Ideas to Finished Content With Claude
AI Explored