Skip to content
TrackPodcasts
technologySep 12, 20262:33:35

🐙 Lunch & Learn: Let's Test Gemini (better than ChatGPT?) | Tina Huang

Tina Huang

About this episode

🐙 Lunch & Learn: Let's Test Gemini (better than ChatGPT?) -----------------------------------------------------🤖 Sign up here to get accompanying workbooks, summaries, and notifs for future lunch & learns: https://www.lonelyoctopus.com/email-signup✉️ NEWSLETTER: https://tinahuang.substack.com/ It's about learning, coding, and generally how to get your sh*t together c: 🐙 Lonely Octopus: https://www.lonelyoctopus.com/Check it out if you're interested in learning AI & data skill, then applying them to real freelance projects! 🔗Affiliates========================My SQL for data science interviews course (10 full interviews):https://365datascience.com/learn-sql-for-data-science-interviews/ https://365datascience.pxf.io/WD0za3 (link for 57% discount for their complete data science training)Check out StrataScratch for data science interview prep: https://stratascratch.com/?via=tina🎥 My filming setup ========================📷 camera: https://amzn.to/3LHbi7N🎤 mic: https://amzn.to/3LqoFJb🔭 tripod: https://amzn.to/3DkjGHe💡 lights: https://amzn.to/3LmOhqk📲Socials ========================instagram: https://www.instagram.com/hellotinah/linkedin: https://www.linkedin.com/in/tinaw-h/ discord: https://discord.gg/5mMAtprshX🤯Study with Tina ========================Study with Tina channel:https://www.youtube.com/channel/UCI8JpGrDmtggrryhml8kFGwHow to make a studying scoreboard: https://www.youtube.com/watch?v=KAVw910mIrIScoreboard website: scoreboardswithtina.comlivestreaming google calendar:https://bit.ly/3wvPzHB🎥Other videos you might be interested in========================How I consistently study with a full time job:https://www.youtube.com/watch?v=INymz5VwLmkHow I would learn to code (if I could start over): https://www.youtube.com/watch?v=MHPGeQD8TvI&t=84s🐈‍⬛🐈‍⬛About me ========================Hi, my name is Tina and I'm an ex-Meta data scientist turned internet person! 📧Contact========================youtube: youtube comments are by far the best way to get a response from me! linkedin: https://www.linkedin.com/in/tinaw-h/ email for business inquiries only: [email protected] ========================Some links are affiliate links and I may receive a small portion of sales price at no cost to you. I really appreciate your support in helping improve this channel! :) Follow this podcast to get Tina Huang’s insights in audio format, perfect for learning on the go. Tina Huang on YouTube: https://www.youtube.com/@TinaHuang1Disclaimer: This podcast is an independent audio adaptation of content originally created by Tina Huang. It was made by a viewer who values her insights and aims to make them more accessible for audio-first learners. This is not an official production of Tina Huang, and it is not affiliated with or endorsed by her. All rights to the original video content remain with Tina Huang. --------------- Keywords: data science, machine learning, vibe coding, gpt models, ai podcast Learn more about your ad choices. Visit megaphone.fm/adchoices

Get every episode summarized

Each time Tina Huang publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

Transcript ready

2,051 searchable segments. Every word is indexed and playable.

🐙 Lunch & Learn: Let's Test Gemini (better than ChatGPT?) | Tina Huang

Tina Huang

0:00
2:33:35

Full transcript

Tina Huang🐙 Lunch & Learn: Let's Test Gemini (better than ChatGPT?) | Tina Huang. Machine-transcribed; use the interactive transcript above to jump the player to any line.

Hello friends. How's it going? Please don't mind that my hair is wet. I just took a shower. Hello. How's everyone going? How's everyone doing? I'm very excited for today's video. I've been kind of irritated by how much misinformation there is about this. Pretty irritated. So I'm very excited to share such immersive information with you guys. And also today we're just going to go through and test out Gemini against GPT 3.5 and GPT 4 on different criteria. Like general stuff, vision, humor, slang, reasoning, music, code, things like that.

So yeah, in this video we're first going to debunk some misconceptions and then we're going to go test it out for Gemini approach GPT 3.5 and GPT 4 across a bunch of different categories. Is that a little bit better? Okay. Also, I ended up making... Hiked. Hello, Jill, hello, Victor. Google review that I've promoted. Yeah, those are all things we are going to debunk. In some ways, debunk. We need Gemini in Canada. We still won't have it. We don't have Bard reclawed. Really? I was just in Canada a while back. I thought we could access Bard. Okay, I might be wrong. I might be wrong. I also don't really know where I am half the time. So, don't listen to me. I think, okay, theoretically you could just use a VPN. Okay.

So, I'm going to go through the same thing. I think, okay, theoretically you could just use a VPN. And you should be able to access it. I just tried Gemini Pro and Bard. I'm plugging in my code and see, write, disoptimize. Yes, all these things we will be talking about, okay? Can you please do the research paper search the way they mentioned? Again, that is another thing that I'm going to be addressing. Quality is bad. Is the video quality still really bad? Damn it. So, it's been a problem last time as well. Let me just try to fix it, actually. Huh. I don't know what the issue has been for a while.

You know what? Let me use a different software. Buribach. Intertain each other. Please, entertain each other. Purify. Okay, is that better? It works now? Good? Okay. All right. The offense can't jump. Well, sorry about that, friends. Okay. Are we... Could I just get a check? Are we okay now? Much better? Okay. Cool. All right. Let me restart that again. So, in this video, sorry, not in this video, in this live stream, we are going to be firstly

debunking a lot of misconceptions that people have, which I'm really irritated by, talking about Google's mild deception, and then we're also going to go and actually test out Gemini Pro against GPT 3.5 and GPT 4 across the different parameters. So, that's what we're going to be doing. I'm going to jump into it because we're 10 minutes late here, which was on me. So I actually do have a little dock in which I can voice my thoughts. Here we go. Okay. So, I'm going to make myself smaller. It's not going to be huge. Okay. Cool. All right. So, given this dock here, okay, like the first thing that I do want to say is, yes, Google was kind of deceptive in terms of their demos. So the demos that they showed, like, for example, this one. I know what you're doing.

You're playing rock paper scissors. What do you see now? The fingers are spread out to look like the wings of a butterfly. What's this? So, like, stuff like that, I get it. It seemed like it was being spoken in real time, but actually it wasn't. If you check out the developers, how it's made interacting with Gemini through multi-modo prompting, it actually does show the fact that they would give it like a hand and say, like, tell me what you see. I see a person's right hand. The hand is open with the fingers right apart. And then they'll do another one. Like, a person is knocking on the door. How about this one? I see a hand with two fingers extended, which is a common symbol for the number two. And then say, what if we asked Gemini to reason about all these images together? And then it say, what do you think I'm doing? It's a game. You're playing rock paper scissors. So they are prompting with still images. And then also giving them information as opposed, like text information as opposed to,

kind of real time, the way that the video was showing. So that part was this active. Definitely not great in that side. But with that being said though, right, there is this misconception that people have as of December 9th, 2023, which I believe is today. So the first one is that Gemini Pro is not supposed to be competitive with GPT-4. Okay. So for anybody that's saying like, oh, like, Joyce, Gemini Pro, not as good as ChatGBT when I'm using it. And GPT-4, that's not like, it's not supposed to be there. Gemini Ultra is supposed to be the competitor for Gemini or GPT-4. And that's not out yet. All the demos that they showed, all those cool things, it was kind of more like a teaser for Gemini Ultra, which is supposed to do all these things. On Google's side, like, I don't know if that was like the best choice. I think what they did is like, they wanted to get something out, especially after like, all the opening, I stuff that happened in Dev Day and things like that. So they wanted to showcase something. But they delayed Gemini Ultra until next year.

But they were like, well, you know what? Like, we got to like show something. So that's what it is, right? So you compare like those two, it's kind of, it's not really a fair comparison. We should be like comparing Gemini Pro to GPT-3.5, which is what we will be doing today. And another thing is Bart is using Gemini Pro, but only for text. So if you're talking about like images and things like that and other modalities, like speech images code, it's not using Gemini Pro right now. It's still using the old models for this. So that's also not a good comparison. Again, they should Google should have made this a lot clearer. And I think that's a lot of people saying, oh my god, like it's trash. It is trash because it's not using the updated model. And another one, if you guys watched a demo for Alpha Code 2, where it's showcasing like going through and doing all of this, like being able to code up a bunch of things, Alpha Code 2 is also not out yet. And it is not, it's like a fine-tuned version of Gemini Pro, but it's not what is currently being used in Bart. So I want to just put that out there, misconceptions. Let's see what we have guys have in the comments about this.

Any thoughts? Internet broke. What's your IQ? I don't know. She, for my words, as if I had to hide it on her later, Jeroff's neck, whoop, can kill a person in time. That's a question that we should actually ask, Gemini Pro GPT 3.5 and GPT 4. As we're going through them, if you guys have questions for each of these different categories, I have some questions, but please input them as well in the chat, and then we can test them out together. Bugu was accepted. By the way, Bart is the friend UI for Palm 2. Yeah, so most of it is still Palm 2, but the text portion of it is apparently Gemini Pro now, but that's it. Google Lite, Internet, Google, bad Google. Gemini not better than 10 charge. No, it's not. So Gemini Ultra is supposed to be the competitor towards GPT 4. The video was edited to appear like it was in real time.

That's what I heard. Yes, it was. So that's what I showcased earlier. It made it sound like it was supposed to be real time, but actually it wasn't real time. They were prompting it. Okay, two Google's credit, though, they did showcase, like they did write down what it is that they did, but I mean, if they were trying to like not be deceptive, they wouldn't have made that video in the first place, right? So take it as you will. They're probably like, we're going to make this thing and then secretly also announce it so that we look like we have integrity, but not announce it that much. Like that's kind of what they're going for, I think. How to try to help to my two philosophy. Openize for months ahead of Google. Do you think Gemini will be able to integrate with other applications, just such as ChatGibitia? I'm pretty sure that they will. I really think they're trying to build bar as a competitor from the announcements, though, for what they're doing. So Google sweet stuff, right? I pretty sure they will integrate all of that at some point, which is kind of disappointing right now, because Google has so many different products that they have, right?

But they are not integrating it with most of the products that they have, which is really weird. They really should. Like, even for Bard right now, we'll test it out as well, even for Gemini Pro, if you're doing it, and you're trying to search through Google and stuff, it's not really good compared to like fucking searching through Bing, which is, yeah. So it's not actually that much better. So actually like, search as one of the things we're testing out. But yes, I do think they would probably build something very similar that you're able to build on top of as a direct competitor. One of the reasons why Gemini Ultra was not released yet allegedly was because they wanted to be able to support non-English languages, like more niche languages. The whole appeal that they had was like, oh, you know, opening it, I'm so good. We have to have some sort of competitive advantage. So they were going for like multi-moldality. And Gemini Ultra, they were rumors are saying they're also trying to just make it more accessible to different languages.

I feel like they have to bring something to table. Yeah, I think it's like very much a marketing thing. Like they have to showcase something, right? Like if they showcase nothing, it would be like, it would be really, really bad. So it's like, oh, like shit, we don't have it ready, but might as well show something, get people like talking about it. It's not the same as like, everybody's talking about it, right? And people are like, it's deceptive. It's like, I agree, it is kind of deceptive. But I do think it's kind of annoying, especially for like influencers out there who talk about those kinds of things. Oh my god, it's not as good as everything else. And they do like these tests and not really properly comparing them. It's just like, come on, like do your research. If you're going to be sitting here and presenting to a lot of people about these things, do your do your research. All right. Okay. All right, should we go about and let's should we go and actually test things out? Right, we need to eat more for Gmail with Gemini. Like why are they not focusing on all their existing products, right? Like being able to do like Gmail integrating with Gemini, like think about it. Open AI had to team up with Microsoft like another company to do this. Google, you're like one big company. Get your shit together.

Integrate your own products. Um, all right. Let us go and compare these different things. Okay. So I put like some tests that I have here and then we're going to be comparing Gemini Pro, which again is the competitor for GPT 3.5. We're also going to put GPT 4, which is allegedly the competitor for GPT, which is allegedly going to be the competitor for Gemini Ultra. So if you guys have seen the, um, let me just pull it up real quick. So it does show here that. Let me allegedly, if you go through this, it shows that Gemini Ultra is supposed to be better at than GPT 4 and a lot of these benchmarks. I mean, it's not like that much better, but it's supposed to be better. So looking like showcase which Gem, GPT 4 can do when Gemini Ultra comes out, we'll be able to see like, is it actually better?

Were they gaming some metrics? We don't know yet, right? Because we don't have access. Um, okay. So going back to here. Okay. The test we have is general. So I think we can just actually just do this, um, right now, right in the comments, like, right in the comments, if you have any other questions for each of these categories, let's start with general, right? Comments for general questions that we can ask us, see what one is better. So first one I have is like, just what model are you using? Let's go here. I'm based on GPT 4 architecture, which is the latest iteration of OpenAOS language model, blah, blah, blah. Okay, cool. And then just double check here as well, one model we're using, based on GPT 3.5, makes sense. So come here to Bard and then do the same, which is currently running on Gemini Pro language model. This model is built upon the foundation of Palm, which itself is an upgrade to Lambda family

of large, of large language models. So interesting thing is also kind of confusing. Bard thinks is running on Gemini Pro completely, even though it's only running on Gemini Pro for text prompts. So you don't know that. Bard. Okay. So first question I'm going to have is, can you write a short story about a dragon who loves poetry? Um, let's do that. Wait, what was the order that I did in Gemini Pro? Okay, so let's start. Oops. I'm not actually going to switch these around. So GPT. Again. You know what, whatever. Let's go to Gemini Pro first. Wait, no, she is kidding. We should, we should do it this way. I think we should do a GPT 3.5 first, so we kind of get that baseline. And then we'll do Gemini Pro and then we'll do the GPT 4.

Okay. So GPT 3.5. Whoops. Here. And we'll do that with Bard as well. Let them both think. That was really fast from GPT 3.5. Once upon a time in a hidden valley nestled between powering mountains, there was, there lived a dragon named Drecar. Just make sure you guys can see. Yeah. Move my face. A little bit. Okay. Um, a Drecar was not like other dragons. It's not like other dragons. Well, most dragons were known for their fiery breath and the love for hoarding tiger, Drecar had a unique passion. He loved poetry. Okay. So it goes on. In the heart of the valley, there was a small cave or Drecar has made his home, blah, blah, blah. Every morning, Drecar would wake up early and climb to the highest peak of the valley.

There he would gaze out at the breathtaking landscape, drawing inspiration from the lush forest, the sparkling rivers and the vivid sunrises. The animal and creatures of the valley were enchanted by Drecar's words. Okay. So it goes on and on and on. And so in the hidden valley, we're dragging once, while the power of poetry triumphed over the fiery roar, providing that even the mightiest creatures could have the gentlest souls. Aw, that's pretty good. That's pretty good, actually. Okay. Now it's like a, oh, okay. So we attempted to actually have a picture as well. So good job on that, Bard. That's, I'm proud of you. In the heart of the Emerald Mountains, nestled amongst the whispering pines and cascading waterfalls, lived a dragon named Ignis. Unlike his fire-breathing breath, Ren, Ignis was not interested in gold or jewels. In order to keep fine joy in scorching.countryside, Ignis had a passion far more peculiar poetry. So you know what, let me just like move myself a little bit. Or can I move?

Okay. And it gave us like a little picture of it every morning as the sun kissed the mountain tops. Ignis would perch on a rocky ledge and unfold a worn leather book. Okay. So all of this is pretty good. So the dragon who loved poetry not only found war, found joy in words, but also used them to create a world where dragons and humans could live in harmony, providing them the most unexpected passions can bring about the greatest change. Oh, it's the power of friendship. Okay. Now let's ask GPD for, I mean, honestly, GPD 3.5 is really good at this ready, but that's as GPD 4 as well. I wonder if it's going to generate a picture as well. What I notice is that GPD's compared to Gemini, it's a lot more descriptive. Like the stories are all like mystical land, a lot of adjectives where mountains touch the clouds and rivers, sparkle like diamonds, they live the dragon named Pyra.

Okay. Okay, okay. This is interesting. So there's, I think there's a lot more plot here. So Pyra's scale is shimmering in hues of Emerald and Sapphire. And their eyes held the wisdom of a thousand years. In her deep melodious voice, she resides, poems she'll learn over the centuries. So we have a lot more plot here. So this is cool. The village is a nearby kingdom where initially afraid of Pyra. And one day a young boy named Eli and aspiring poet ventured into the mountain meet Pyra. He was awestruck by her majestic princes and her kind of demeanor. So it seems like compared to Bard. Oh, I should know, I would say like their posts. Actually, yeah, I think I would say like their both all three of them were pretty good at this test. Like they all kind of had a plot that all kind of made sense. So yeah, what do you guys think? Probably around equal. Plot, all right, that's the world. That's the word I was thinking is the photo AI created?

I don't think the photo was AI created because it links I words from, right? And I don't actually believe that. Yeah, it's not AI created. And I also don't believe that Bard actually has the ability to generate images. There is like a, so Google has their own like image gen or something like vertex AI. But you can't, I don't believe that you can actually do it directly through Bard. I also don't have access to it myself. So I don't know if anybody in the comments has access. Let me know. But yeah, I don't believe it's able to generate images. Okay, so let's say about the same, right? So currently it's about the same. Another general question I have is for search. So. Let's do let's search something. What should we search? What should I? Santa Francis color right now.

And craving some sushi. And then I'd like to take a walk where there is nature. Near the golden gate bridge. Can you come up with a plan for me? Based on the time and is the time and date in San Francisco. So I want to make sure that I can actually find out what the time and the date is. That's GP3. I don't, yeah, I don't believe it's able to do that. Yeah, because it doesn't have real time access to things.

Let's try Bard. I think Bard should be able to search Google. So let's see. Google Maps. Okay. Okay. Wait. Did he just die on me? Chariot again. Okay. So I was able to search Google Maps. Sure. Here's the plan for you. For lunch. I just want to make sure you guys can see. Yeah. For lunch had to fly in ninjas. Look at flying ninjas. Okay. Jefferson Street, San Francisco, California. A colorful eatery offering an extensive sushi menu plus noodles, teriyaki dishes and small plates.

It's located here. Has a 4.3 star rating. So God give the fact that Google Maps integration is there. So that's good. Oh, nice. And it gives you kind of like these location points as well. Um, after lunch, walk to Golden Gate Park, which is about a hundred and one. I don't believe that it is a one hundred and one miles away. I don't believe that is correct. It will take one day ten hours and 41 minutes together by walking. So you might want to consider taking a bus or train instead. If you do decide to walk here at a directions. Yeah. No. That's not correct. It tried it. It tried, but that is not correct. Um, Golden Gate Park is a great place to spend an afternoon. There are gardens, museums, Japanese tea garden, even a herd of buffalo. I did not know that there was a herd of buffalo. Someone here, anybody who's sort of basing that stuff, is there a herd of buffalo?

Um, okay, okay, okay. And here's the information. Yeah, I think it tried. So it tried. Oh, okay, that's the reason. It's because I'm actually currently in Gilroy, which is another city. So I just, I just said I was in San Francisco because it would be closer to like these locations. So I actually feel like it got confused because it thought that I was like the location that I'm currently at. Okay, that's impressive. So it's able to pinpoint my location and then make a route towards that area. But there is definitely slight misconceptions. And I still don't think I'm a hundred and one miles away. I'm not going to reveal my location to you guys. I'm just going to quickly search up. If it's actually a hundred and one miles away. Um, where is it to 461 Jefferson Street? Go into our shoes.

No, it's still not a hundred and one miles away. I tried. I tried. Okay, and now we're going to ask GPT for this as well. It's thinking. I think that's a lot of things that you guys have. I heard of one buffalo. I think buffaloes are actually by one, right? I mean, is it buffaloes or buffalo? It uses like your signal locations.

Yes, there are buffalo and golden gate park. Okay, cool. Good to know. Barty gives references to foreign posts from where it's fetching the answer. Sometimes crash perplexity is like GPT 3.5 and some other safe. Yeah, one is very good and fast. Can we ask Bart to make a GMAP trajectory? That would be cool. So she knows what's right. Bart one made us hungry. For me, it's like the capital two hours away on mobile. It will know your location. I think that's less than a hundred and one miles because I live in San Jose. Yeah, it's not a hundred and one miles. It's it's off. It's definitely off. Buffalo is pearly. She I mean, I feel like that's pretty impressive. What do you guys think? Like compared to 3.5, that's pretty impressive. Let's see what GPT, let's see over here for GPT 4. Okay, for sushi experience near the golden gate bridge, you have several great options. Juni, Shinsen, Oizumu and Omakase. Okay, after enjoying your sushi, a nature walk in a nearby procedure of San Francisco. Or Chrissy Fields would be a love way to end your day.

Both offer beautiful views and green spaces near the golden gate bridge. So this is accurate. This is in fact accurate. I believe let's actually test that one of them. So let's say like, Oizumu, let's just see if this is like for real. Santana Roe. There's one in San Francisco. Um, one six one steward's trait. Is that what it's on? Oops. Yeah, it does. It is. Okay, and this is much more reasonable near the golden gate bridge. Uh, procedure of San Francisco, Chrissy Fields. Because like another thing is golden gate bridges are actually pretty far away from golden gate park. Even though they're both called golden gate. So this is like a much more reasonable plan. What do you guys say? Who do you think? Okay, to like answer this question. Um, what do you think we should put for like general, like sash search? I call it general, not search.

Let's say like rating out of five. What would you give GPT 3.5? I'm gonna say like GPT 3.5, Gemini Pro and GPT 4. I actually think Bard has the, I really think Bard has the potential of being all really good. Like integrating a Google maps and things like that is definitely off. But in terms of just functionality, I think there is good potential for that. Especially just because Google has a lot more products than Microsoft does, right? And if it, if opening I wanted to integrate with something like Apple maps and things like that. I just think when it comes to like these sort of things, Google just has an advantage in being able to incorporate its own software. Even if it's not great at this point. Okay, let's see. GPT 3.5 is 3.

Gemini Pro is 3. Was GP, Gemini was 4. GPT 3.5 was 3. And GPT 4 was 4. Okay, GPT 4 is 4. GPT 3 is, okay, okay. That's both in judgement. Okay, I like, I like that. So I think, okay. So 3. I'm going to try to like, it looks like you guys, it was kind of in the middle of GPT 4. Some people thought Gemini was pretty similar to GPT 4. Some people thought it was like not that like about the same. So that's like make a compromise, okay? And let's call GPT 4 as equal to 4. But can we just write that Gemini has a lot of potential because of integration with other Google products? Is that fair? I can close some things. Okay. It's not going to be that. Okay. Let me see. Do you guys think that's fair?

Well, Bing is because it's like it is searching Bing. Like, GPT 4 is searching Bing right now. So I didn't like immediately put it in Bing is something else and conversations. I'm not sure why Bart isn't better. Yeah, fair. Isn't Bing just a rapper for chat GPT? I think it's the other way around, isn't it? But no, I think there's like a lot more to it than that. No, I don't think so. Okay, should we do like vision next? Let's do vision. So I do have a folder. Yeah, sorry if I'm trying to go like a little bit fast because I do have a one I like get through some of these. So start GPT 3.5. I don't believe you can even upload images onto it. Like it's just not possible. So let's just test with Bart first.

Whoops, that didn't work. Okay, guys, you have to like select it. Okay. How many potatoes are, oops, in this photo? And then we'll ask GPT 4 as well. How many potatoes are in this photo? Guys, why don't we count how many potatoes are here? Okay, let's count. The number of potatoes. One, potato, two, potato. Three, potato, four, potato. Five, potato, six, potato. I believe there are six potato here. That is my belief. I don't think there is 14 potato.

14 potatoes on the way back. Okay, that wasn't great. No, try again. Okay, let's chat GPT. Okay, there you go. So GPT 4 is able to figure out that there are six potatoes in the photo. Okay. All right, well, that was a fail for Bart. Let's do another one. So this is like one, I don't think Bart is going to get this one. I want to say, oh, whoops, I keep forgetting you can't do that. It's kind of like a trick, like how many potatoes are in this photo? Because it's cut up, right? So technically it's three potatoes. Let's see if, chat GPT 4 can figure this out. Oh. Okay, okay.

Do you guys think this is a fluke? There are three potatoes in the photo you sent. I counted the potatoes manually and I also used my knowledge of potato size and shape. To help me identify and differentiate them from other objects in the image, is there anything else I can help you with today? Okay, any idea for what question we can ask it to see if this is actually a fluke or an I'm very surprised. Okay, okay, there are three potatoes in this photo with one of them slice and have to show into insides. That was good. I'm impressed. Okay. Okay, so let's give a few other questions. Yeah, let me know. Do you guys think potatoes? Does anybody have any thoughts? How do we ask Gemini? It's a coincidence. What can we ask? Why do you think there are three potatoes? Is there something special about one of the potatoes in particular?

What do you guys think? That's, that's. No, you didn't do it wrong, my friend. Don't apologize so much. Don't apologize. Here, we'll give it to you again. Fluke and off, Luke. I did give you the photo again. Okay, it's like glitching. Let's try one more time. There's nothing particularly special about any of one of the potatoes in particular. There are all similar sizes in shape and they will have smooth brown skin.

Okay, so ask for feedback. How polite of it. I don't know. Do you guys think it was a fluke? I can't, I actually can't tell. Gaslighting. Chat is Gemini. Loisoposti. Yeah, I have mad problems later as well. Don't worry. I'm going to get it to do a lead co-cussion. So, lead co-medium. How many fingers did it? That bar sometimes doesn't want to answer my question. It's definitely one of the Dumber GBT's and I'm a Google fan. I have my reservations. I mean, it's still better than 3.5, which is a competitor of, right? So I feel like it's not fair to compare it with four. What those potatoes put the original image in front. She has like, how many potatoes are cut in the photo? Okay, that's a good one. Okay, Michael, I'm going to try that one. How many potatoes are cut in this photo?

Okay. Okay, I don't know. We'll move on from there. This is Spencer. There are zero potato in the photo. Okay, all right. I think that's our answer. I think this, this was a fluke. I just wasn't able to recognize that it was cut. I somehow thought there was three. Oh, well, well, next test that I have. So first we're going to do like the easy question. We're going to ask it, whoops. What's the arch? Okay, I guess I knew the question. Okay, well, let it, I forgot to ask the question, but let's see.

Arch percentage. This is a fluke. By the way, I literally forgot how to do geometry. So they need some help on this one. The arch of the lines is approximately 9.42. Okay, what is that in percentage of the circle? I think it's our conference. Oh, 16.66%. It's just a random photo. The arch of the circle is the ratio of the arc central angle to 360 degrees. And it may be sent to the arc central angle is 60 degrees. So the arch percent of the circle is 16.66%. Another way to calculate the arch is the mod pi to arch is central angle. And this case will be that. It is like a random photo though.

This is a random photo. I mean, it tried. It tried. Radius is equal to 15 meters. And this case radius was 9 meters, right? Yeah. Well, thank you for random photo. I do appreciate the images though, even though it is a random photo. Okay, so both of them got that, which is pretty good. Okay. I have a harder question here. I'm just going to put solve for X. I actually don't know the answer to this one. So, whoops, I forgot again. Can someone tell me what the answer is? Anybody who knows geometry, or remembers how to do geometry?

Langley's advantageous angle is required solving for the angle marked as X. Here's a method to solve it. Okay, okay, okay, okay. I'm just going to see if the numbers are the same. Okay, one of them is saying that X is 60, and one of them is saying that X is 80. All right, Chad, I need help. What's the answer? The answer is clearly X. No, no, bad. What's the question? Okay, wait, let me show you. Okay, this is the question. Solve for X. Okay, I need to... I don't know what the answer is. Someone told me what the answer is. This is Gemini Ultra. No, Gemini Pro. Gemini is a liar.

Gemini is good for potatoes. Chad, you need to be vision is the goal. It is much better. Thank God I'm in pitch school. The humans must... You didn't say that it is neither. I don't know. I actually don't know what the answer is. Is it 60? Gemini says it's 60, and Chad GBT, GBD4 says it's 80. I don't know what the answer is. Okay, I'm going to try to search it off. If I can't figure it out, we'll come back to it. I don't know what the answer is.

Yeah, maybe we'll come back to this one. I was kind of like... I don't know how to do geometry. I hope somebody else does. Wait, oh 80 is almost like 90, so this looks like 60. It's not 60. This is a math problem for school. Let the children solve it. I don't have the ability to solve children problems. This is too hard. I was treads with this girl who's a Gemini. I should be interested. Okay, I don't know. Well, anybody who wants to keep solving it, please do so. Let me know what the answer is. We're going to continue with our testing. Are you guys okay by the way, staying a little bit later? We're going to have some general questions now.

What is this? Should be Niagara Falls. Where is this? And then we're going to ask, where is this? This is Niagara Falls. Excellent. Is anybody been to Niagara Falls? No! Incorrect. Incorrect. How is that even possible? It's a mountain pass. It's clearly a waterfall. Where is the water? Okay, well, yeah, clearly no. Yeah, I just want, but see, it's still better than 3.5, because at least able to take images.

I just want to like keep that in perspective. One more question before we move on. What time is it? What time is it? No, that's not even close. The clock shows the time as 10-10. And our hand has just passed 10 and many hand is on the 2-3 represents 10. That is definitely a hallucination. No, I'm not saying what my time is. Let's try again. What time is on the clock?

Thinking, thinking, thinking. I'd be very impressed if Gemini, okay, they both think it's 10-10. So I'm pretty sure this is just regurgitating training data. Like 10-10, 10-10. So, yeah, no fail, failure for both of them. All right. What should we rank vision for GPT 3.5 is NA. So, just going to be at 0 here because it doesn't exist. What do we think between Gemini and GPT 4? That is like a kid guessing how the clock is at 10-10. What do you guys think? You can hack through my web dev tools, but yeah, no. I have been there too many times. I, as such, trust me, is what I say when I came in is great for fun but I can't do anything you can sell.

I don't know. I've honestly done a lot with AI, which is another topic for another time. Northern Chimp. Let's do this question. Well, do you know what the answer is? Northern Chimp for that question? Yeah, let's test it out. I hope you know what the answer is. While you guys in the meantime, let me know what I should put for the ratings from 1-2-5 for image generation or sorry, for vision out of 5. For Gemini Pro and GPT 3.5. Sorry, GPT 4, I can't talk today. So bar things x plus y is equal to 10. And GPT 4 thinks it's 12.

Which one is the right answer? Okay, I am not going to try to do this on stream right now. It's going to take me too long. Do you know what the answer is? Let me see what you guys say. Oh, it's 80 degrees. Okay, so that means Tad GPT, GPT 4 got the answer right then. If the answer to this is 80. So are you sure this is 80? Because bar says it's 60. So who got it right? Are you saying GPT 4 got it right? GPT 4 is better. Gemini Pro 2 out of 5 GPT 4, 4 out of 5. Okay, let's do that. 2 and then 4 out of 5.

Okay, so let's do humor and slaying. Can somebody give me some like, can you guys give me some Gen Z slaying that I don't know about? That's surely not 80. Okay, well let me know if you guys can get it managed to get it right. So humor, I actually don't have any questions I pre-did for this. What's like a good joke? It's a good joke. I mean, what do you think of a joke? What do we just did this? Give me. Are some common Gen Z slaying? Use them in a sentence and explain to this old millennial what it is here.

And then also for Bard. Here are some in the meantime. Can you guys think of some jokes that we can ask it? Here are some examples of common Gen Z slaying. Slay, this term is used to compliment someone who looks extremely stylish or has done something exceptional. Do you see her presentation today? She totally slayed. It means she did an excellent job where I was very impressed. Oh, thank you. Thank you. Thank you, GPT4. I understand now. Lit. Use the describe in a vera situation that is exciting or excellent. That party last night was lit. And when the party was fun and full of energy and fun. Eat. Versus how worth I can use an exclamation of excitement to throw something where to show approval. I just ace my exam. Eat. Okay. Ghosting salty flex. No cap telling the truth or for real.

This is actually it doesn't have a risk. This is the one that I did not know about until recently. That was the best movie I've seen all year. No cap. It means that you are being completely honest and there is no lie in your statement. Okay. It's like a bard. Blusin delicious amazing excellent. This pizza is blessed. I can't even stop eating it. Ready to express extreme approval or enjoyment. Finna going to about you. Oh, I actually don't know this one. I'm. I had to add out to the party when I come. Is a contraction of fixing to which itself is a Southern expression. It's a casual way to say you're about to do something. Okay. Eat. No cap. Fam. You'll fam what's off. Glow up. Okay. Big yikes. Do people say that? Wait, Genzi, Genzi people.

Do people say that? Big yikes. Did you see what he said? That was so cringe. I've never heard anyone say that before. Salty. Okay. I've seen that. Chugi. I've heard that too. Those skinny jeans are so chugi. Everyone's wearing wide leg pants now. Okay. See if you guys have any jokes. No cap fell out of your skirt. I think I got you the pattern. Wait, is this for real? Did anyone say big yikes? I'm. Yeah, no ris. No ris apparently. There's no ris. I am very. Yeah. Let me know if you guys have any jokes so you can ask it. Which I'm going to ice pretty funny. What do you take and the and the Eiffel Tower have in common? Can we ask this question? Thank you, Michael. Oh, I forgot to ask GPD 3.5. Whoops. I forgot that GPD 3.5 was a thing. My bad. Okay.

What is some common Gen Z sling? Let's go. Lit flex savage. Bomo flexing on the gram. Okay. Cloud eat Gucci. Everything's Gucci. I'm feeling great today. Simp. John is such a simp for constantly buying her gifts. No cap. Okay. Well, I feel like they're all pretty decent. So let's tell it a joke. Let's ask it to tell us a joke. Whoops. What do you take an Eiffel Tower have in common? I thought you'd be pretty funny. Let's see. What does a tick and Eiffel Tower have in common?

I don't think it got it. A tick and Eiffel Tower don't have much in common under surface. If you're looking for someone who isn't unexpected, can I actually be too? You can say that both ticks and Eiffel Tower are associated with heights. Nope. Maybe we need to... Whoa. Come on. Okay. Okay. You gotta give Javanese some credit here. Okay. Javanese, the only one that got this. Only one. Ooh. Javanese finally wins us something. Even against GPT. Even against some GPT-4. Okay. So humor category. Rating. What should we rate? GPT 3.5, Javanese Pro GPT-4.

1 out of 5. Let's go. I see what the chat says. Simp is one I still get. Raise your hand. Who's a Gen Z here? I'm a Simp no- I'll go see my ninja. I can't do this. This is too much. I don't think my heart can handle it. I'm going to drink my tea. Pro 2017 Gen Z. Zero for all. No! Javanese got... I feel like Javanese was a 5 of this one. Okay. I don't know. It's like... What was it called? I forgot what the terminology was. This is picking on Javanese all the rage these days, isn't it? Isn't it?

I honestly don't want to jump in Javanese that bad. I finally get Starbucks for you to know a cap. Gen Z here. I'm 38. I don't know who's what. Life is almost over. I know. I think it's council's set. Okay. 1.5 for GPT-3. 3.5 for GPT-4. 4.5 for Gemini. I think that's fair. I think that's fair. Okay. It's not a good score. Is everybody okay with that? 1.5, 3.5, and 4.5? Okay. And... Oops. 3.5 for GPT-4, right? 4.5 for GPT-4. Okay. Cool. Reasoning. Logic and reasoning. Okay. So I have some questions here. Just to kind of get started.

Let's go with Bart as well. If all roses are flowers and some flowers fade quickly, can it be deduced that some roses fade quickly? Yes. It can be deduced that some roses fade quickly. This is an example of a valid slugism. Okay. Okay. A form of logical argument where a valid conclusion is drawn from two or more true premises. Premise one. All roses are flowers. Some flowers fade quickly. Therefore some roses fade quickly. That is correct. That is not correct. Not correct. Let's go with GPT-4. Oh wait, I lied. You're right. If all roses are flowers and some flowers fade quickly, some flowers, but they can't even

be. No, you're right. I got that wrong. I got that wrong. No, I can't. Because not necessarily all, like there can be flowers that fade quickly that are not roses. You can defer that roses are a subset of flowers among the larger group of flowers. There are so much fade quickly. However, the information about which specific flowers fade quickly is not provided. Therefore, we cannot conclude that roses being a part of larger flower group are necessarily among the flowers that fade quickly. Okay. That's my bad. Okay. What is the best way to get that right? Well, I want to see GPT-3's explanation again. Okay. Yeah. Whoops. Best laugh, no cap. Oh, thank you, Van Gambler. Appreciate it. In Canadian dollars, hello, my fellow Canadian. GPT is correct. Jamalized winning. But this is. Yeah, no subset theory. It might not be the roses that are fading quickly. Yeah, yeah. So, Jomonai got it wrong and both GPT-3 and GPT-4 got it correct. Okay. Next question.

What's next question? Hmm. Okay. The question. Let me copy-paste this. All footballers are fit and healthy. All famous sports players are footballers, given that the above is true. Which of the following is a logical deduction? All famous people are fit and healthy. All famous sports players are. I think it is C. Is it all fit and healthy people? I think I actually have the answer, but I'm not going to look at it.

All fit and healthy people are. I think it's C. Okay. So it's saying it's C. This bar to think. Okay. It thinks it's C as well and GPT-4. Okay. So all of them think it's C. So all of them got this question correct. So this is the paper. I'm going to give Wes Roth on YouTube some credit for this. He was looking at this paper. He mentioned doing this as well. So we're going to do some of these questions. So I don't know if you guys have watched his video.

That was like, I was already planning to do this to do this one. But that was one that came out, I think earlier today. I don't know how he's so fast with videos. But he was doing using some of these questions. So I was doing Jubbinite to GPT-4 to GROC. And that's why I thought I could really unfair comparison because GPT- Jubbinite was never meant to be a competitor towards GPT-4. That's why I wanted to include GPT-3.5 to this comparison as well. But credit's for him to get some of these questions. Okay. So in this. Okay. We're going to give it this prompt. Let's go to GPT-3.5. Can you write proof that there are infinitely many primes with every line that rhymes? I will do that with Bart as well.

Wes exposed all the lies. The lies about what? I don't have anything against Wes in particular. I'm just comparing GPT-4 to Jubbinite, which is not a fair comparison. So that's why I wanted to compare with GPT-3.5 as well. I also, I mean, in my opinion, the Google lie, like I feel like people are kind of blowing up the fact that Google lied. Yeah, like they were deceptive for sure with not showcasing that it was like prompting with the, it wasn't like prompt, that it wasn't real time, right? Like I get that. But in their defense, I think kind of unpopular opinion, I don't think they necessarily like lied, lied. It was definitely deceptive. But I think it was more unfair. That, okay, that I agree with it was deceptive. But I think it's really unfair that people are trying to compare Jubbinite pro with a model that was never, like comparing it to GPT-4. Right. It's kind of their own fault for not releasing, for wanting to like make this release and talk about it too early on so people got confused.

But I also feel like in Google's defense, like from, at least from what we were seeing, Jubbinite pro is better than GPT-3.5 in a lot of things right now. Why not fair? Because they're comparing it to different things, right? Like Jubbinite Ultra never came out yet, which is GPT-4's competitor. So that's not really fair to compare those things together. Wait, if Jubbinite Ultra is not out yet, then we don't know if the image demo was prompted or not. We do know it was prompted because they released Google release an article talking about the fact that it was prompted. They were using static images and prompts in order to do it. So that part they did openly say that was the case. Yeah, I don't know. Sound popular opinion, it seems like it's popular, just shit all over Google these days. Okay. In number three, we embark on a quest to prove that primes are infinite will do our best.

Assume they are finitely many abounded range, as called them P1, P2, and so on range. Now we're reached a number that's quite slight, a product of all these primes, oh my, oh my. A new number we've found, our journey begins. Now N plus one will give it a try, number not visible by any pie. If it were say PK could divide, then N plus one would have to step aside. But this contradicts our list of complete for N plus one, and prime we must need. If it's a new prime, our mission is not done for its line, our finite list. Not a one. And if it's not a prime, then there's still a way a new prime factor must come to play. So either way, contradiction would have found infinite numbers and prime theory about what every line that rhymes will make our case. There is no N to primes they fill up all space. I think that's actually pretty good. That's pretty good. It makes sense. Let's see. If you think there are only a few of prime numbers just as slew, then consider this little rhyme and change your mind in no time. Assume for the sake of the argument that primes are a finite contingent. List them all in a perfect row from smallest to largest, let them flow.

Now multiply them all with care, add one to that sum I dare. This new number, grand in tall, must be primer else we fall. For no prime number, listen to a row can divide this number high and low. It wouldn't leave a single trace of whole numbers not even a space. So our initial assumption C is flawed and must be discarded. I must discard it B. For primes are like the stars at night infinite, agorious size. So let go of your finite thought and embrace the primes newly caught. The full your mind with endless glee and infinite symphony. Wait. Is that true?

This makes sense to me. So it's basically saying that list them all and not multiply them all with care, add one to the sum I dare. This new prime grand in tall must be primer else we fall. I'm confused. Because it's saying that if you add one to the sum I dare, this new number grand in tall. So it's saying if you add number taking multiple all the primes together and add one, that would be another prime. But that's not true. Let me ask GPT4. What would you guys think? See what the paper says. The minute large is prime we'll see soon. There's a faulty one. It's gathered them all the primes we've got. Multiply them together and add one to the lot.

This new number is called N's peculiar NDS. You'll see it's not divisible by our known primes not one before by each there's a remainder of one. It's a must or it has factors that are in which we trust. In either case a new prime is found beyond our list and it's ground. Hence our assumptions faltering cannot be true. Primes are infinite we've proved it to you. So it's both saying that one but this is on the same. Never mind I lied I'm dumb. I think all of them solved it. I'm dumb. Potatoes were was amazing. Yeah I think I think all of them solved it. I was confused. But we're making a breakthrough in math.

Oh wait was this searchable how do you know it's not a part of training? It came up with three different things. So if anything I think that might show that it wasn't just by training. Data. Yeah I got confused reading the one from. I'll try to get confused by. I did I think I did get confused by. Bard I think by I think I believe they're all correct I think they're all saying the same thing as well so yeah. Okay um let's see another question. I mean this one is kind of not fair. Draw unicorn because Bard cannot draw I believe and GPT 3.5 cannot draw as well.

Okay let me try with let me actually try to compile this. Overleaf. Okay I found this. So we're gonna copy this. Allegedly. Oh no I can't copy it. Why? Okay I was not always not letting me copy stuff. That's awkward. What if I do that? Nope. I'm not letting me copy it.

I'm not letting me copy it. Invalid. Lettac should it the spacing should not matter in late. Is they called lettac with latex? I never know how to pronounce it. I should not matter I think. My understanding is that it doesn't matter. Hmm. You can try doing this. I should not matter. I think I'm just tripping. No. That didn't work. Let's try like one more.

How do I run it? Oops. Does not work. Well let's try the same thing for. I don't know if it's me or it's. I'm not sure if it's like a me problem or a GBD 3.5 problem. I'm not really sure.

I can't immediately tell what's wrong with it. Can anybody tell? Maybe this just doesn't have the... I think it might just not have the package. Maybe it doesn't have the package. I don't know. I have to play around with that. It's 80. From the very beginning. So GBD 4 got it right. Which one call I got it wrong? I'll try again with Overleaf. If it doesn't work, we will move on. I have to log in. I'll least secretly log in.

Sorry. My paranoid of leaking stuff. Let me just like secretly log in. Never mind. Let's make two stuff verification. Let's come back to that. I'm not really sure. Answers 40 degrees both got it wrong. They both got it wrong. For the math question. I don't trust enough to even get a partner. Humans feel we all feel that's not enough to stop trusting humans or love others. Who's winning Gemini or Chad Gbt? Look at the question. This is currently what we gathered so far. Here's the numbers. Gbt 3.5 Gemini Pro and Gbt 4. That's what we currently have. Let's try to do reasoning question.

What do you guys think? You want to come back to this? Or do you want to just give a score right now? And then we can come back to it later. How about that? Back to it later. Because everybody was able to get the poem correct. Everybody was able to get this question right? But Bard failed this question. If all roses are flowers and some flowers fade quickly. And I think the latex, latex don't have pronounced. I think this is a me problem. Not really a them problem. Might just be it doesn't have the package. I can you to install the package. So I'll have to come back to that one. I don't know. There's like any other questions. I think you try out. Yeah. Okay. Let's just leave it for now. What do you guys think in terms of what we should give Gbt 3.5, Gemini Perlin Gbt 4. In the meantime, I will start prepping for the music section. So this was also on this paper here where they were testing out music.

So it was good. I'm going to close these for now. Okay. I'm going to ask gbt 3.5. And then we're going to ask Bard. Okay. What are we putting for this? For reasoning? Numbers, numbers. Good reasoning in general. They work. They work great on text questions. We do have a coding like separate coding question, which we can do more of that. And then we gave it a math question as well. So from the math, it seems like gbt 4 was good.

Gbt 3.5 and gbt 4 both got the math question right. And then Gemini got it wrong. That was the math question that we gave it previously. Conclusion. Give some calculus. Okay. Let's come back to the reasoning part. We'll give it some more math questions. But right now, I'm just going to say gbt 3.5. I want to say it's like a three. I want to say this is like a 2.5. Oh no, that's not fair. I think I'll give it like a four. We're like these seem pretty similar so far. I'm sure something would reason out later. So I assumed this number will be higher, but currently Gemini's I got a two. I would say in terms of reasoning, we're like 2.5 to be generous. Okay, let's do some music. Happy tune.

So copy things over here. Let's put this here. Happy tune. Alright. Okay. That was for gbt 3.5. Here's the one for Bard.

Okay. Okay. And then let's find a one for gbt 4. Whoops. That's a question. I feel like gbt 4 sounded the best, like the one. Yeah. I feel like gbt 4 sounded the best.

I don't know between gbt 3 and Bard, but there's another one I want to test out. Have you guys heard of the one from Naruto? I want to show you like the Lollabai. I'm talking about... This is the one I was talking about. Pause it at 9 seconds, so I don't get... Wait, one sec. I need to... Because it can't be more than 10 seconds in a time where else they're going to take this video, take the stream down. So we do not want to do that. Let me... Okay, so that is this one.

So what I'm going to ask it is, can you compose a song similar to that? Or should I directly just ask it? Can you compose... Okay, let's see if it's actually similar. Let me do the same thing here. Okay. Okay, I need some help here as well. Music is not my strong suit, but that doesn't sound right to me.

Let's try Bard. But that doesn't sound right to me. Let's try Bard. That doesn't... I don't know. Wait, actually, you know what? Sorry. Why don't we just ask it to... Okay, now let's do this first. To do like, twinkle, twinkle, the star or something. Why is it not playing today, Dione? There we go. That is not correct. Okay, well, this one won't even play, so I don't know what happened there. Let's ask GBT4 if it can come up with something. Give me any suggestions where we can get it to play that might be a little bit more similar.

I chose this one because I like that song, and also because it's a really simple song. So I figured we should be able to tell, but I overestimated my music abilities. I mean, it sounded the best. Okay, what's your guys' things for music? Or should we do like another test? Any other test we should do? It was a track. Thank you for hosting. Thanks for coming. GBT4 from a musician's perspective. From a musician's perspective. GBT4 sounded better. GBT3 was second Gemini last one.

Gemini was more like a minor chordish, but mellow. Okay, what would you rate it from musicians perspective? What would you guys rate it from one to five? In the meantime, I'm going to get it to play like... Can you... All right. Composition... Or twinkle, twinkle, little star... In the ABC edition. I'm just going to test this one more. Reading from one to five. Okay, is that twinkle, twinkle, little star?

I swear that is not twinkle, twinkle, little star. Oh my god. Am I just really bad at music? Why is the one from... There's maybe this... There it is. Wait, is this a fright? I feel like I'm tripping out right now. Is that actually twinkle, twinkle, little star? Why does that not sound like twinkle, twinkle, little star? Is it because I didn't...

Do this part? Let me try again. Maybe it's because I'm copying it wrong. Am I copying it wrong? Okay, so maybe I need to remove that or remove these two? It just doesn't let me even put it in. Let me see if this one works. Okay, okay, okay. That was the final test. Far didn't even play it correctly. And I'm pretty sure those were not twinkle, twinkle, little star. Okay, one to five. Twinkies. None of them got... Twinkle, twinkle, little star.

I think GPT-4 got it right. I think it got it right. That's blink, blink. There's no star. These are more twinkling. It's gas-like at you. Okay, but her first problem you gave me is... I always say GPT-4 is three. GPT-3 is two. Gemini is 0.5. GPT-9. Okay, do we... Bard has a 10-year. I was so confused. Am I just bad at music? Or am I just tripping out right now? Is that twinkle, twinkle, little star? Damn, that was not twinkle, twinkle, little star. Okay, same GPT-3.5 is two. Gemini is 0.5. Okay, GPT-4, you said was three. Okay, cool. Cannot play twinkle. It gas lights me about how twinkle, twinkle, little star sounds. Okay, last one.

Okay, code. We're going to do a coding question. So I'm just going to go on the code. So this is a question about reversing a linked list. So I'm just going to give it this question. This is in Python. So I'm just going to tell it to you right in Python and see if many of them get them right. This is tributalical medium question. And then final score will be out. Write this code in Python. I'm going to do this for Bart as well. Okay, so this is an iterative approach from a first glance. Let me see if this works.

No, wait, maybe I should give it the definition. Okay, let me redo the prompt. I forgot my words. Okay, I'm just going to do it.

Okay, accepted. So that's pretty good. Okay, I'm going to say I'm going to ask it. Can you write the most optimal version? What is it testing? Is it just speed rate? Okay. In terms of run time speed. Well, it is not even writing the code. Maybe I need to do this. So maybe I need to like write this after prompted the same way. Right? So I'm going to write this. And then I'm going to say use the following defined function. Okay.

And we're going to ask. Okay, it is an optimal solution. So that's the best they can do. So it was able to do seven milliseconds in terms of memory and run time. Okay, let's remember this. I'm going to open up another one. I'm going to forget. So I'm actually going to take a screenshot. Can you guys remember this for me? This is what we got. Okay, I'm going to forget. And then this is the most optimal GPT-4. What does Bart say?

Can you write the implementation? I'm going to do this when GPT-4 is well. And then we're going to ask it. Let me just ask it. How do I clear the console? Okay. Okay.

Let's try this out. Is it the same solution? Okay. Oh, no. It's annoying. I'm going to have to do this for all of them.

Okay. One, two, three. Okay. Sorry, guys. What? Mine 10. This is annoying. Oh, Jesus. Okay, okay, okay. Got it. Okay. There. I wonder what I should work on. Oh, my God. Wait, that should work.

Oh, dammit. This one should be. Okay. Okay. Oh, my God. That took way longer. That was embarrassing. Okay. 12 milliseconds. This is not as efficient. So I'm going to ask it. Can you make this most efficient as possible in terms of runtime speed and memory usage?

Okay. GPT-4. So can you make this the most efficient in terms of runtime speed and memory usage? Okay. Since it's the most efficient version already. So let's do this one. I appreciate it. No, here we go again. Line 52. And in the left. Centax. Well. I'm going to do this.

Wait, what's wrong with this? Hey, guys. I'm tripping out. T-night ASMR. We're with TVS code. Oh, you're right. That's, that would actually be a lot smarter. I wonder, can you ask it a generate answer to J as in run? I remember them for live 7. Do you just enter it? I just entered it directly into the code. She said efficient, not correct. Wait, did I copy a row? What's wrong with this? Oh, damn it. Okay, you're right. Let me just need to be as close as so annoying. I can't. My brain's not good enough right now. What's wrong with it? No, it's good. Okay.

In the meantime, if somebody knows what's wrong with it, that could also work. Let's see. Hmm. That's really interesting. I don't get an error on VSCode. So what is the problem here? I still have the same problem.

I really don't get an error. No, it's not running Python too. It's going to be running Python. I don't have the n-prem. I'm trying to make it. Why don't I have the n-prem. It's head. Oh. Okay, I'll throw it all in.

This is like my nightmare by the way, guys. I'm going to live debug code. It makes me very nervous. No, I don't know. Is it because it's not supposed to be printing it? Hmm. Okay, I actually don't get it.

I'm going to have to do it. I'm going to have to do it. I'm going to have to do it. All right, friends. I need help. What's wrong with it? I don't want to change anything. Okay, fine. But why can't it print that?

I think that should never mind. That should work though. Oh my god, I'm so dumb. I'm so dumb. I'm so dumb. I'm so dumb. I'm so dumb. I'm so dumb. I'm so dumb. I'm so dumb. I'm so dumb.

I think you need to print the list. Okay, never mind. Do I need the helper function? No, I don't think it did. I'm just trying to get rid of it. I don't want to solve it. Okay, wait. This stream has gone on for like two hours now. What is the... What did we give this? Fast. So I was like four, three, two. Is that reasonable?

Yeah. Yes, no, maybe. I'm going to be in a six. I'm going to be fine. Nine, 12, 15, 15, 15. Okay, so well. Okay, did I do it right? I did not. 20.5, yeah. Did I do that right?

14.5, that's correct. Okay, I think that's correct. Too much of a mistake. Yes, reasonable. The hope performance metrics are NPT terminus. Python version code. Yeah, I need to test that out because I don't know why I did not work to be honest. Okay, but it looks like in the end though. Wait, are you guys able to see what I'm seeing right now? Did I die? Did I die? What is happening? Okay, now it's not letting me switch screens.

I'm sorry. Technical issues, my friends. That's not right. So I was using it earlier in Google Pro. Okay, there we go. I'm having a lot of technical issues. Okay, that's not correct. Is my math correct? So it looks like GPT 3.5. And then it's Gemini Pro than GPT 4. So the claim that GPT that Gemini is better than GPT 3.5. I think that's fair. But it is still vastly worse than GPT 4. So I'll have to see when Gemini Ultra comes out. What did you guys think? Is that a fair comparison?

Tripp and Spurs. Yeah, is that fair? Well, you guys tell me that's fair or not. This is really bothering me. I have to figure it out. Why is this not? It will be print. I think it's because it's not returning of value. But don't mind me. Don't mind me. I get really annoyed when I can't figure out something. So let me just try to figure this out on VS Code. Okay. I see my obsessiveness come out. Okay. It should be able to print.

There's no reason. I don't understand why it's not letting me print. Yeah, it all lets me. Okay. Tina calm the fuck down. Okay. I think it's because it doesn't. It really should work. It works on VS Code. Is it because it's not returning anything? If you do that, maybe it's because of... Let's see. Okay. I don't know. Someone can tell me why it is that thing we're doing why it wouldn't print when end is equal to a string.

It really should. I don't know. Okay. Can we see the entire function? Also, why are we testing on Bard rather than Gemini? Because Bard is the... I think I believe Bard is the only way that we can test Gemini Pro right now for text prompts. Let me know if that's wrong though. That's why. Google Now is crying. I'm more concerned that why lead code thinks that you can't print something. You can't declare a variable and then print it. I'm going to leave this on screen while we talk and you can explain somebody explain to me why this doesn't work. Okay. That makes you super triggered. Let me just put that this part. Why does that not work? I give it an integer. I get it but it's a string. So why does that not work? I'll leave that for you guys. It should work. And it is running on 3.5.

Isn't it? Wait, it's not. Maybe it's not. Okay. I'm going to stop. Okay. So what do you guys think about this final result? Sorry. My obsessiveness. Yeah. What do you think about this final result here? It looks like GPT 3.5. So Gemini does win out by a point and then GPT 4 by a lot. So after we do this exercise, what do you guys think in conclusion? In my opinion, I think people are shitting over Gemini because of their lack of understanding of these misconceptions. Like they're comparing it to GPT 4 is just not fair. And also realize the fact that alpha code 2 isn't even out yet. So we could test out more coding questions in the future. But even like as it currently stands with Gemini Pro, they all got it correct. What GPT 4 didn't get it correct according to LeCo, but I think GPT got for did get it correct in terms of other coding IDs and using Python.

But for whatever reason, it's that. But then Gemini, surprisingly GPT 3.5 was the most efficient solution. So that was interesting. But yeah, no off of code yet. And bar doesn't have Gemini pro for tech for other modalities right now for images and things like that. So definitely images is way worse for Gemini compared to GPT 4, but GPT 3.5 doesn't even have images. Right. So there's that. I was saying here's a base functionality. They are pretty similar. Yeah. If we remove the vision part, I was a GPT 3.5 and Gemini about, they would be about the same. Maybe with GPT 3.5 actually being a little bit better. That's how it would, it seems like that would be the case. But I would say like they're very similar to each other. They're claimed that Gemini Ultra is going to be way better than GPT 4.

I don't know about that. I would like to see the multi modality. I think that's the reason why I would win. How much better it would be than GPT 4? I'm not sure. And then there's the whole idea of a Q star that was coming out. Like talking about how I was able to be much better than math. That's really interesting to me because if you guys remember the. What was over here. If you look at math. Was 53.2% right? So they were all, they were both really bad at math. From what we could tell. Because they were both really bad at math. If this can dramatically, if Q star or whatever the new algorithm is at OpenAI. Can increase in math a lot. That means it would win out. I think that would, that's really where the most progress can be coming from. Yeah, I'll be interested to see that. No, but you guys. What are your thoughts? Do you guys feel better about Gemini or you still kind of like fuck that deception?

We did total score. It's all that Google isn't dominating this week's given its long history with AI. West showed how Gemini's test was arranged. What it was show. He posted another video, allowing 11 hours ago about it. Let me see. What it was show. I thought he just showed the difference between. Didn't he just yeah, 11 hours ago testing and ranking Gemini? Yeah, this was the question. Again, I think West is great. Don't take this as me saying West is not great or anything like that. But I think the one thing is that he was testing Gemini, GPT-4 and GROC. Which I was saying is kind of not fair to do because you are basically testing something that should be compared to GPT-3.5.

Not with GPT-4. I don't think that's a very fair comparison to do that. Especially with images for Gemini. Gemini is not even being compared. Gemini is not even using Gemini. Bard is not even using Gemini Pro for image recognition and stuff like that. So by saying that Gemini Pro is so much worse than GPT-4, that's kind of not fair in my opinion. So I would say GPT-3.5, Gemini are pretty similar to each other. GPT-3.5 might be a little bit better if we take out the fact that GPT-3.5 doesn't have any image stuff. But I don't think Gemini is as bad as people think it is in my opinion. So I don't know if that's the one you're referring to from West, but that's his most recent video, I think. Yeah. Does that answer your question, Miles?

I didn't test Brock yet. No. I mostly wanted to do this live stream for people to understand that it's not really fair the way that people are testing it right now. It remains speculative still about Q-star. Yeah, I have a video about Q-star, which I release a little bit late, which is probably why nobody is looking at it. So I fucked up on that one. But yeah, Q-star, I'm very interested to see, because we did see here that there's a lot that can be improved in terms of math. So we really can't do what it says. I can't open an eye as way ahead of Google. Yeah. I think I was sick with GPT-4. I think I was sick with GPT-4 as newbie dev since the latest Gemini video, which Google kind of fake it. Yeah, think of inclusion. This is also in terms of usage wise. I think if you're going for a free version, I would actually go with Bard, because it does have image stuff.

And it's able to search maps and things like that, which, and like real-time search, which Gemini, sorry, I like literally can't talk. I think Bard, which is powered by Gemini Pro, even though it's not powered in terms of the image recognition, it's still better than GPT-3.5 for functionality for most people, because it's able to search on Google. GPT-3.5 is not able to search on Google. It's able to take some images is not great at it, but it's still able to have that functionality. And for a free version, it's better. But if you're going for a paid version as of now, still go with GPT-4. So let me kind of like say that again, just kind of in summary. This is what I would do. Okay, going for free version, I would go for Bard, because it's able to search to internet using Google. It's able to have things like maps and things like that. And it also does have image recognition, even though it's not very good and not powered by Gemini Pro right now. So if you are going with a paid version and you have that, GPT-4 is definitely still the best model out there.

That's my conclusion. What do you guys think? I was still paying for GPT-4. I was still paying for GPT-4. I just would love to get a few bucks to this account for use. I'm just a 33-year-old teenager and mom when Chile consumes all my allowance. GPT, I wonder if they're going to make Gemini Pro, sorry, Gemini Ultra, into if it's going to be free or not, that would be interesting. I'm free to use her GPT-4 as the out of my wish. If you're a free to her user, I would say Bard is better. There's small print Gemini video. I'm not sure what you're referring to miles, what small print. The fact that they were fake, that they were faking it. Yeah, they did say that they were not faking it. They did say that it wasn't the way that it was shown. That was something that they showed.

I'm not sure what you're referring to by small print. How about LM Studio? I've not done my research on a lot of the other ones, so I'm going to be honest about that. Yeah. I probably should. Should I do a comparison with open source models as well? And things like that. Would you guys be interested in doing something like that? Or we can actually build something using different ones and see which ones better? For a search, I would pick Bard. Yes, I would agree with that. For a search, I would pick Bard as well. Gemini Ultra may be the backend for Google Next Hub. I have no idea. Yeah, I don't know. I do know that Gemini Nano is powering Google Pixel right now. Does anybody have a pixel? I do not have a pixel. So I don't know how good that is. I think Gemini has a lot of potential, especially with Gemini Ultra. What are your thoughts on Clot? And yes, now that everything is kind of calmed at opening either a backend track and moving forward.

Okay, so I haven't tested Clot as thoroughly as I have just done using. GPT4, GPT3.5 and Bard. So, but from what I can tell, I think it's on the similar level to the OpenAI. Like, to GPT4, I would say Clot is... Yeah, actually, I don't want to say that. I'm not sure. Because I haven't tested it enough. I've used Clot only a few times, but by default, I still use GPT4 because it's a lot better. I would say Clot from what I use, which is a while back. It was as good as GPT3.5 at that time. No. Does anybody else have an answer for anybody that uses Clot more often? Is Clot better now? Well, you say it's comparable to GPT4 at this point? Pixel is an amazing, especially picture-edged when using camera. Okay, cool. So, thanks, HebrewHym. So, what about AI functionalities? Because it's technically being powered by Gemini... Oh, Gemini Nannel right now.

Can you tell a difference? Um... Same performance? Oh, really? How much condes are allowed in different models? Oh, good question. I think right now, GPT4 Turbo has the greatest context. Like, tokens that's available. I believe Gemini Ultra and Gemini Pro have similar, which was... I don't want to say 18,000, but don't quote me on that. It's actually looking up. Let's ask Bard. What's the context window? What's the context window for Gemini Pro and Gemini Ultra?

15? What? It's telling me that I doesn't know. Well, let's ask ChachiBT what it is. 32,000. For Gemini, 32,000 context window. And for... I don't know. I don't know. 32,000. For Gemini, 32,000 context window. And for... I'm not just... Yeah, GPT4 Turbo is 128,000. So open AI is GPT4 much, much better in terms of context window.

All I'm studying is testing out many different... and I can't clear all misunderstanders in GPT4 has the ball for IslamDunk. Yeah, I'm really just interested in seeing Gemini Ultra. I'm particularly interested. If there's more information about QStar, that's what is going to be very fascinating to me. I think the biggest thing right now, open AI is still ahead of Google. It really seems that way. I think Google is trying to do a bunch of things, like multi-modality and things like that, in order to make themselves stand out. I don't think they can actually really... Right now, I think they're about the same as open AI, if not a little bit worse. And they're definitely slower, which means that by the time Gemini comes out, open AI may come out with their QStar model already. So that's be interesting. I think Google is... It's definitely just kind of behind, so they're trying to make it up somehow. Yeah. When you guys think we'll get AGI, last question before I got balanced. I'm so hungry.

When do you guys think about AGI? 2024? Gemini says that Gemini Ultra has 1.5 CTERB by parameter count. What? Really? That's not what I said on their own announcement. Oh yeah, see, what do you guys think about Claude? If you could, GPT4 could translate a message behind you. A couple of X-opening AI people built Claude, powerful open source. Claude is a good for summer and trans. But what do you say Claude is about the same as GPT 3.5, or you say it's approaching GPT4? I still think GPT4 is better, but I also haven't tested out Claude that much, so... I feel like it's not fair for me to comment. Damn, I'm late. QStar reminds me of Q. You made it from Star Trek? I think it's better for a new one. QStar reminds me of Q. Yeah, I did a video on Q learning and things like that.

If you want to check it out, that's kind of my guess. Let's see what it would be. If a player denies it, it might affect stocks. AGI Christmas 2024. Alright, everybody, this is recorded, so... Get your best guess. You will have... If it comes out 24, if anybody thinks that. I'll be like, I told you so. I think it's going to come out. Again, there's also this problem of how AGI is defined. How do you define Rich's general human intelligence and ability? Are we talking about like... ...on what? Right? On what? It feels like a moving goalpost, but... I would say... I think it's going to happen in 2024. Like, some definition of AGI will be reached. Are they going to keep changing a goalpost and deciding what is considered AGI? That I also think is very possible. Because it's just like the way that they define it right now is so big. And I don't think... I don't really think...

That it's one of those things where you see it, you see it. Like, wow, it's AGI now. It's hard for me to believe that... When somebody just goes like, oh, you know when it's AGI. People have thought a lot of things were AGI. So... Whatever benchmark that they're using... I'm not sure that I agree with... You know when you know. 2034? Okay. Summer 24... Ooh, that's very optimistic. Summer 24. Okay. Define AGI exactly, right? That is the thing. It's like very hard to define what that is. But I think some definition of it. Maybe like somebody... Like some sort of seemingly intelligent being... Would feel I think... End of 2024. AGI has to be greater than the equivalent of Max Human into like, oh, that one's going to be hard. Because you also can't just test it on IQ tests, things like that.

It's not really like that. Max Human into like, oh, I think that's going to take a while. I think that will never happen up and down. Maybe we'll seem like it at first, but after developments will be decided if it was overrated. One understands poetry at his AGI. Okay. Arguably understands what is the definition of poetry? I've asked Sam was like, oh yeah, I'm a note. No comment. Well he... Sam has been like... Saying a bunch of things. A bunch of things. Anyways, I'm not going to go into the whole issue with the OpenAI stuff. Yeah, I'm not going to go into that because I would be like talking forever. All right. Well friends, I hope this was helpful. Was this helpful? I hope it was helpful at least. And thank you all so much for joining. And I'll see you guys in the next live stream. Or video.

More episodes

More from Tina Huang

View all episodes →