Skip to content
TrackPodcasts
technologyOct 8, 202616:32

#650 Neil: 4 AI Image Models Tested and the Results Surprised Me

AI Fire Daily

Get every episode summarized

Each time AI Fire Daily publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

About this episode

“Capital One's tech team isn't just talking about multi-agentic AI. It's called chat-concierge and it's simplifying car shopping using self-reflection and layered reasoning with live API checks. It doesn't just help buyers find a car they love.”From the transcript

I tested GPT Image 2.5, Nano Banana 2.1, Nano Banana Pro, and Nano Banana 2 across five image challenges. From text accuracy to realistic motion and YouTube thumbnails, the results weren't always what I expected. See which model won each round. 🏆

We'll Talk About:

  • Photorealism and AI Text Rendering
  • Prompt Following and Camera Control
  • Realistic Motion and Physical Details
  • AI Infographic Generation and Text Accuracy
  • AI YouTube Thumbnail Generation
  • GPT Image 2.5 vs Nano Banana 2.1 Test Results
  • Which AI Image Model Should You Choose?

Keywords: AI Image Models, Nano Banana 2.1, GPT Image 2.5, Nano Banana Pro, Nano Banana 2, AI Tools.

Links:

  1. Newsletter: Sign up for our FREE daily newsletter.
  2. Our Community: Get 3-level AI tutorials across industries.
  3. Join AI Fire Academy: 500+ advanced AI workflows ($14,500+ Value)

Our Socials:

  1. Facebook Group: Join 299K+ AI builders
  2. X (Twitter): Follow us for daily AI drops
  3. YouTube: Watch AI walkthroughs & tutorials

Hosts & guests

Transcript ready

190 searchable segments. Every word is indexed and playable.

#650 Neil: 4 AI Image Models Tested and the Results Surprised Me

AI Fire Daily

0:00
16:32

Full transcript

AI Fire Daily — #650 Neil: 4 AI Image Models Tested and the Results Surprised Me. Machine-transcribed; use the interactive transcript above to jump the player to any line.

Capital One's tech team isn't just talking about multi-agentic AI. They already deployed one. It's called chat-concierge and it's simplifying car shopping using self-reflection and layered reasoning with live API checks. It doesn't just help buyers find a car they love. It helps schedule a test drive, get pre-approved for financing, and estimate trading value. Advanced, intuitive, and deployed. That's how they stack. That's technology at Capital One. We have we have all been there. You know, you type out a prompt. Oh absolutely. A flawless masterpiece instantly appears on your screen. It feels like magic. But then you zoom in. Right. You always have to zoom in. You zoom in and the guy in the background has like seven fingers. The neon store sign is speaking complete alien gibberish. Yeah, totally. And the illusion just shatters completely. It shatters very quickly. Yeah. We see the messy reality right behind the algorithms. Welcome to the deep dive. I am incredibly glad you are here

with us. Today we have a very specific mission. Google just released Nana Banana 2.1 into the wild. Finally. Right. And it promises vastly better image quality. It claims highly accurate text rendering. It boasts significantly improved prompt following. And it is four times cheaper. Wow. cheaper than Nana Banana Pro. That is. I mean, those are some incredibly massive claims, but we really need to see hard evidence. We cannot just trust the marketing brochure. Exactly. We're testing four major competitors today. Nana Banana 2 is our baseline model. Nana Banana Pro brings the high quality visual details. Open AI's GPT image 2.5 is the heavy hitting heavyweight. And Nana Banana 2.1 is our new challenger. It's quite the lineup. It is. We put these four models through five rigorous tests. Real world visual tests. We want to separate the marketing hype from actual performance. Okay. Let's unpack this. We have to start with the

foundational elements, I think. Right. Right. Before we test complex physics, we start with reality. The basic foundations. We need to see perfectly natural lighting. We need to see believable glass reflections. We absolutely need readable text in a static environment. So for the first test, we wanted an everyday scenario. Picture a busy neighborhood coffee shop. It is a slightly cloudy afternoon. Very specific mood. Yeah. Exactly. Two friends in their late 20s are chatting. We requested a 35 millimeter lens right at eye level. We wanted realistic facial expressions. We wanted relaxed body language. Right. We even threw in a subtle curve ball. We asked for light condensation on the windows. We give the models very specific text challenges too. The main storefront sign must say good coffee, good people. The window lettering needs to read freshly brewed daily. To take away coffee cups, say morning club. And the poster inside. Yes. And there is a promotional poster inside. It says buy one, get one 50% off.

The comparison results here were incredibly revealing. GPT image 2.5 won this round overall. It really did. It gave us a highly accurate storefront. The human characters looked completely natural. The lighting perfectly suited that cloudy afternoon vibe. The smaller lettering on the cups wasn't entirely perfect. But the overall image looked highly convincing. Nano banana 2.1 had very mixed results here. Yeah, it was weird. It handled the large storefront text fairly well. But it completely failed on small coffee cup labels. The overall result just looked slightly less polished. Nano banana 2 is actually much worse. It repeated freshly brewed daily in multiple places. Just everywhere. Yeah, I just pasted the text everywhere. Some smaller signs contained completely distorted and complete text. And Nano banana pro made a classic mistake. It reversed the storefront name in a reflection. I'm looking at these results and I just don't get it. It's baffling, I know. We have a multi-billion dollar system that can generate perfect, photorealistic condensation on a window. But it spells

morning as M-O-R-N-N-G. How is it so smart and so incredibly dumb at the same time? Because it's not thinking. Right. Generating AI text is like giving a child a massive box of Lego bricks. You ask them to build a word they do not know how to spell. Oh, that's a good way to look at it. They know what the shape of an S looks like. But they might build a backroom. They might use the wrong colors entirely. They just do not know what the letter actually means. That is a brilliant way to picture it. People constantly forget what an AI model actually is. A math system trained to recognize and recreate complex visual patterns. Exactly. It is fundamentally just math. It does not actually understand our world. It does not think. Why do these incredibly advanced systems still struggle with simple spelling on a tiny coffee cup? Because they do not really read text. They process text as complex visual shapes. They do not understand linguistic characters at all. So they're just drawing. Yeah, the system just tries to draw detailed squiggles. Squiggles that statistically look like human letters. It is mimicking

the visual appearance of writing. It is not actually writing. So the AI draws text like shapes not as actual readable letters. Yeah. That shape matching works okay for simple textures. But it really struggles with strict underlying logic. So it cannot render text because it is just mimicking two dimensional shapes. But what happens when we force the AI to think in three dimensional space? That's where it gets messy. Right. Text is flat. Real life has deep depth, left, right, and dynamic perspective. That brings us directly to our second test. Let's move from static text to spatial awareness. We set up a rainy gas station evening. We were extremely specific about the camera placement. Super specific. Yeah. The camera is exactly 30 centimeters above wet pavement. It is approximately four meters away from the human subject. And it is tilted up exactly 20 degrees. This creates a highly dramatic low angle perspective. We gave the models very strict left and right rules. A woman is standing near the center. She holds an open red umbrella in her

left hand. And a white coffee cup in her right hand. Yes. Both hands must be anatomically correct. Which is always a gamble. Always. And there is a white foredoor sedan parked behind her. The driver's side is facing the camera. And there is a gas pump on the far left. Everyone struggled with this complex prompt. Every single model failed somewhere. GPT image 2.51 for overall cinematic perspective. It did look beautiful. It also got the hand placement perfectly right. But it faced the white sedan completely backward. Which ruins the whole shot. Yeah. Neno banana 2.1 got the low angle correctly. But it put the umbrella in the wrong hand. And it put the car backward too. Yeah. The woman appeared much too far away as well. Neno banana pro gave us great cinematic lighting. But the umbrella was in the wrong hand again. And the camera angle simply wasn't low enough. I have to admit something here. I still wrestle with prompt drift myself. Oh me too. You spend

20 minutes typing 50 careful details. The machine just casually forgets five of them. It is incredibly frustrating. What's fascinating here is how the failure happens. These systems can render ultra-photo-realistic rain. They can make wet pavement reflect neon light beautifully. Which is visually stunning. It is. But that requires intense mathematical lighting calculations. Yet they cannot grasp basic spatial geometry. Why does a system that can generate photo-realism completely fail at knowing left from right? Because spatial relationships are fundamentally different from lighting. Geometric logic is processed differently than surface texture. How so? Texture is just localized pixel prediction. Spatial awareness requires understanding deep three-dimensional space. These models lack a true structural grounding. They just guess. Yeah. They do not build a mental map of the scene. They just predict adjacent pixels. Right. Spatial logic is totally different from just recognizing static pixel patterns.

It gets even more complicated from here. Wait until things start physically moving around. We have seen AI fail at static left and right instructions. Let's escalate this further. What happens when we introduce gravity, momentum, and dynamic physical movement? Our third test is a typical backyard trampoline. A 25-year-old man is jumping very high. He is exactly 80 centimeters in the air. And we used a professional sports photography style for this. Right. We asked for a 70-millimeter lens and a blazing fast 1-2000 second shutter speed. The physical detail is really matter here. That fast shutter speed freezes motion instantly. His medium-length hair needs to be lifting up. His loose gray t-shirt should show realistic folds. A plastic water bottle bouncing nearby. The safety net must be completely intact. And he must be wearing sneakers. And Nano Banana 2.1 finally takes a clear win. It absolutely nailed the dynamic physical motion. It really impressed me.

The bouncing water bottle was physically believable. The sneakers looked completely correct in mid-air. The entire jump felt incredibly realistic and grounded. Gpt image 2.5 really struggled here. It made him jump way too high. The water bottle just floated awkwardly in space. Like you had no gravity. Exactly. The hair movement wasn't entirely convincing either. Nano Banana 2 was even weirder. It made the guy completely barefoot. We explicitly said not to do. Right. His jumping pose was terribly unbalanced and awkward. The water bottles movement felt highly unnatural. Nano Banana Pro left the bottle standing perfectly upright. The safety net also looks very broken and incomplete. Generating all of this from scratch is just wild. It really is. Whoa. Imagine scaling to a billion queries. It is actually staggering when you think about it. It really boggles the mind. Every time you hit enter, this model is hallucinating reality. It is effectively hallucinating the physics of a bouncing bottle. It builds it from

scratch based solely on latent space patterns. And millions of people are requesting this simultaneously. Exactly. How does the AI understand the physics of a bouncing water bottle without actually having a physics engine? It relies entirely on massive data sets of frozen motion. It has seen millions of high speed sports images. It essentially hallucinates the correct physical trajectory. So it's just guessing where things should be. Yes. It predicts where pixels usually go during motion blur. It does not calculate mass or real gravity. It mimics visual patterns of falling objects instead of calculating real physics. Spot on. Mid-roll sponsor read placement. Welcome back. So we have conquered dynamic physical movement. Let's pivot from the physical world to abstract data. Can an AI organize complex scientific information? This is a tough one. We wanted to see a vertical four by five layout. The topic was decompression sickness. It is also known as the bends. We wanted an educational infographic aimed primarily at beginners.

The requirements were very strict. We needed accurate science on human nitrogen absorption. We needed the facts on clear bubble formation. And it had to be readable. Right. We wanted simple accessible English and absolutely no invented statistics or fake medical claims. GPT image 2.5 wins this round. Clearly the infographic was highly detailed and structured. The text rendering was incredibly strong. The scientific content was actually surprisingly solid too. It was. Nano banana 2.1 had a very clean visual layout. It was extremely easy to read. But it had a massive glaring problem. Oh yeah. It was bad. It gave overly specific hallucinated medical advice. It gave context free ascent rates for scuba divers. That is a terrifying safety issue. Nano banana pro was very messy overall. It duplicated a whole section entirely. And it had blatantly inaccurate nitrogen absorption facts. And Nano banana 2 just distorted several key scientific

terms completely. Here's where it gets really interesting. Think about the last time you searched for a quick explainer online. Oh all the time. The prettier the design is, the more easily we believe the data. We trust a beautifully clean layout instinctively. We just drop our natural guard. We do. We totally do. We assume the math is correct because it looks professional. Which is a very dangerous psychological trap. It makes the hallucinations a much harder to spot. We lower our critical thinking for good graphic design. Based on these hallucinations, is it safe to trust AI for educational graphics right now? No. It is generally not safe at all. Visual aesthetics often mask deep factual inaccuracies entirely. AI outputs often look highly professional and polished. They look like a real textbook. They do. And that professionalism bypasses our natural skepticism completely. You have to verify every single medical claim manually. You cannot outsource your scientific accuracy to a diffusion model. Always fact check. Because it prioritizes pretty layouts over actual scientific accuracy. That brings us to our final most

personal challenge. Right. For our final test, we move from general data to personal identity. What happens when you put yourself into the AI's hands? This was fun. We used the standard 16 by nine YouTube thumbnail format. We uploaded your actual personal portrait as a reference. Yes, you did. We wanted to preserve your exact facial identity and structure. The prompt was quite ridiculous. I am wearing a bright yellow shirt. I'm reacting with total exaggerated surprise. An anthropomorphic boxing banana is punching me directly in the cheek with shiny red boxing gloves. The banana wears shiny red boxing gloves. Yes. And the bold text must say nano banana 2.1. It needs vibrant colors and highly dramatic YouTube lighting. Nano banana 2.1 takes the crown here. The colors are incredibly vibrant and punchy. My face is highly recognizable and accurate. It really looked like you. The dynamic punch is visually clear and hilarious. And the text is perfectly readable and well placed. Gpt image 2.5 had a truly great punch effect.

But it badly distorted your human face. It also overcrowded the text completely. His way too cluttered. The text was just far too massive for the frame. Nano banana pro made the banana look much too friendly. The punch lacked any real visual impact or kinetic energy. And nano banana 2 made the background far too dark. Your surprise expression felt completely disconnected from the actual punch. It just looked weird. This raises an important question. Why did the newer model excel at simplicity here? Why was the layout so much cleaner? Why did the older gpt 2.5 over complicate the thumbnail compared to the newer nano banana 2.1? Because newer models are tuned very differently today. They are better tuned for modern visual hierarchy. They understand the deep importance of negative space and design. Less is more. Exactly. Older models just throw massive pixels at the canvas. They desperately try to fill every single empty gap. Nano banana knows exactly when to pull back. Gpt over complicates the effects while nano banana just understands basic visual hierarchy.

If we connect this to the bigger picture, being the newest model doesn't automatically mean being the best at everything. We saw wildly different strengths today across the board. Google's newest release did not dominate completely. Gpt image 2.5 remains the absolute champion for pure photo realism. It is best for complex camera instructions and strict framing. Yeah. And it completely dominates with detailed text heavy infographics. But nano banana 2.1 is the new go to for dynamic action. It handles realistic physical motion beautifully. And it is absolutely perfect for highly stylized YouTube thumbnails. You absolutely need a multi tool approach. You get it. Reliant. Just one single generator for everything. You have to pick the right tool for the specific job. You really can't. The technology is too fractured right now. So what does this all mean? The funniest part of this entire deep dive, even the absolute best most expensive models still make bizarre mistakes. They really do. They fail in completely unpredictable hilarious ways. Like backwards sedans or entirely mutated

human hands. Exactly. Mistakes that a quick glance might completely miss. Getting a 90% perfect image is incredibly easy now. Getting a 100% perfect image is still incredibly hard. Very hard. The illusion is incredibly fragile. The magic breaks the second you look closely. Now we're curious. Based on these rigorous tests, which AI image model would you trust with your next big project? Thanks for joining us. Right. See you next time. Capital ones tech team isn't just talking about multi agentic AI. They already deployed one. It's called chat concierge and it's simplifying car shopping using self-reflection and layered reasoning with live API checks. It doesn't just help buyers find a car they love. It helps schedule a test drive, get pre-approved for financing, and estimate trading value. Advanced, intuitive, and deployed. That's how they stack. That's technology at Capital One.

More episodes

More from AI Fire Daily

View all episodes →