
About this episode
Get every episode summarized
Each time Sidebean publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
322 searchable segments. Every word is indexed and playable.
Full transcript
Sidebean — Why this thumbnail keeps changing - Startups 101 | Sidebean. Machine-transcribed; use the interactive transcript above to jump the player to any line.
I have no clue what thumbnail you've clicked on. We're shooting this video in November 2022 and we created the system that changes it every day. And what we're doing here is bringing this practice to the very extreme. This is an exaggeration. It's an exaggeration of a practice that more or less has defined the web as we know it today. It's this practice that started with a clever American advertiser born in 1866, but it has defined the internet. The reason why YouTube thumbnails change is A-B testing. Essentially, creators are testing one thumbnail against the other to find a winner. But that test can be repeated again and again and again to optimize through the very last consequences. In this practice, it extends to newsletters and email subjects and landing pages and websites and entire applications. So I went on a hunt for examples and stories from founders and product managers that could tell me even more about what they've done and how it has impacted their work. And I discovered, of course, how we are all essentially getting things, of thousands
of experiments that run across the web in every page you've visited, experiments ranging from simple color changes to changes that tapped into the very core of our human psyche. So in this video, I want to tell you a few of the stories that we've discovered, successful A-B tests that companies have run, the platforms they used to run them, and some of the best practices from our own experiences doing this ourselves. Club. Club. Hopkins is known as the father of the A-B tests. Before tech, before computers, before screens, before calculators can figure out that by creating coupon campaigns with two different versions of the same campaign, he could essentially improve results dramatically. He called it scientific advertising. I had earned him incredible recognition and a salary in 1907 that was equivalent to about $5 million in today's money. But translate that to today, A-B tests again have to find the web as we know it. Now before I go into these examples, there's one very key concept that you need to understand. It's simple, but if you don't get this part, nothing in the rest of the video is going to make a lot of sense. What an A-B test does, essentially, is take the amount of traffic going to a page or to
a thumbnail or to an in-app experience and split it into two or more groups, the control and the variant. You can have as many variants as you want, but for the test to be valid, those groups need to be mostly random. There needs to be a minimum amount of people going through the test and enough difference in the outcomes to determine that the test was valid. This is called statistical significance and you don't really need to understand how it works in the background. We actually made this tool. It's free. I'm going to link it in the description and you just need to plug values there to determine if the significance is there. And you have to do that before making any crazy decisions with your results. Why so? Because improvements based on an audience that's not large enough or based on a difference that's too small may not be relevant enough to declare a winner. And this can be very dangerous. Now remember, A.B. tests can be used on anything from ads to websites to buttons to YouTube thumbnails. And you'd be surprised how much difference a very small change can make. Some of these stories I knew already, some I discovered while researching this video,
but listening and reading through them meant all my team got bombarded this week with ideas about how we should be doing to change thumbnails and approach our blinding pages and newsletters differently. Also, a very last disclaimer, these are all success stories that the teams behind them may public. Most A.B. tests are not successful, they're just not relevant enough to talk about them. Out of the ones that we run a slight bean, something like half doesn't end up improving anything, but testing our pages is just a routine part of the marketing team's job. Okay, examples. Let's start with Groove. When preparing their first website to come out of their beta, they did a full market research exercise to educate the design of their new page. And that included doing a market gap analysis to understand not only what their competitors were doing, but how they were selling it and how they could stand apart. And then based on this, they worked on this full marketing site design with proper mockups, wireframes, finished designs, they launched it and it damped. Two weeks in, there was an end page conversion, that's the percentage of people that visit the website and signed up for the platform was 1.87%, which is just catastrophic by most
standards. Imagine if a click to the website cost them $3 or $4 on ads, a lead or a sign up on this page would be costing about $200. That's a lead, not a paid customer. Now this whole project took months, and according to Alex, their CEO, it cost them around $50,000 in designer fees. The work that they did was good. The pages look good. The market gap research sounds to me like a valid experiment, like a valid document. The problem was that this was too expensive of a project and it wasn't lean. They developed dozens of pages spent weeks and months before testing anything. So two weeks in, they were willing to throw all that work in the trash, swallow their pride, and start over with a new page, which not all companies, not all product managers are willing to do that quickly, especially when you spend so much money. Their approach to the second time was something that they called copy first, which doesn't ignore the value of good design, but again, according to Alex, it lets design take you back seat. They went and talked to hundreds of customers to understand what was driving them, what was the pain that they were looking to solve, and why they had signed up for the platform. And their biggest breakthrough, their biggest discovery was that many of them were frustrated
with ZenDesk, which is their competitor, this customer support tool. So they had a customer support platform already, but hated it. And the simple insight educated the copy for them, and it has echoes even through today's landing pages that have been optimized many more times. The fact that the page was not so design heavy and didn't need to be perfect, allowed them to ship it within a couple of days and it meant that future changes were a lot easier. And after A-B testing it against the original one, conversion rate increased to 4.3% from the 1.8 that they saw on that launch version. Now the difference between A-B tests are usually very small, their gradual improvements, and that's why you need to double check statistical significance. The handling conversion rates with a variant is a massive breakthrough, and you don't see that every day. And again, it doesn't stop there. Key conversion parts of the website, like a landing page that gets a bulk of the website's traffic, they need to be in constant iteration even today as we post this video, their landing page has completely changed. The best tool to run these landing pages is a tool called Google Optimize. It's free and lets you run one page against the other, or even change parts of that page with just installing a little snippet.
Remember, make sure that you give your tests enough time to make them statistically significant, because it's very tempting to just draw conclusions before the data is reliable. Unbounds is another solid platform that you can use to quickly build landing pages and test them against each other. This was an example of a complete overall, but A-B tests can also cover small details and make small but incremental and important differences. Pubstip for example, designed just a small corner of their blog, and it was enough to get them thousands of extra leads, because you see Hubsbode's game is content marketing. They publish content on many axes, and want you to read and engage with that content as you get nurtured to eventually buy their product. Last year, they discovered the best of users that use the search bar that they have in the right corner. They discovered that if they use it, there's a 1.5X extra chance of them becoming leads. In other words, if you use Hubsbode's search bar, you are very likely to sign up for their newsletter or to download one of their free resources, which of course requires you to give them their email. This will make sense once you see it, but finding this connection between search and conversion
to becoming a lead requires a lot of data analysis. I'm not going to bore you with that in this video, but let me know if you want us to cover some of that. Of course, getting more people to search was critical, means more leads. Now they played around with some designs, with some small changes, with different copy inside the search bar, and they found that this version had a 6% increase in people that ran a search. Now that's nowhere near, grooves 2X improvement, but in a page like Hubsbode with thousands of visits they get, 6% is a breakthrough. Still, just looking at how many more people ran a search is a dangerous fall success. You have to look at your core objective. Their objective was more people becoming leads, giving them their email, and that improvement was 3%, not the whole 6%, but still a small success, essentially just changing the copy inside the bar. Of course, by the time we published this video, Hubsbode's page will have likely changed, and it's been overhauled again because optimization never ends. Testing experiences like this requires much more advanced tools. I've worked with Optimizely in the past, and they've pivoted from competing directly with Google Optimize into much more advanced tools for data scientists.
Now, before I move into some of my favorite YouTube examples, I want to talk about today's sponsor, which is Chartmobile. Chartmobile is this platform where startup founders, where revenue leaders, where data analysts can work together to speed up their growth and find the levers that they need to pull to make that growth happen. They have a beautiful UI for SAS, for subscription companies to dive as deep as they want in the metrics of their customers, and we've been their customers for much longer than they've been our sponsors. At Slightbean, we've been using Chartmobile for years to track our core growth metrics around MRR and Churn and Expansion to decide on our pricing strategy and to look at the results of our AB tests. For thousands of B2V SAS companies, it's been super powerful to look at their data through a single pane of glass to get these clear insights based on your existing billing setup, and then analyze that based on plan, or location, or customer size, or cohort. Even very yet, it's free. If your company has under $10,000 in subscriptions, it's totally free. And after that, it has a pricing that scales with you. You can sign up for free using the link in the description. Thanks again to the Chermobile team for helping us make cool videos.
And now let's talk about thumbnails. Because Atlantic Page might get thousands of views per month, but for any channel of a decent size, their thumbnails are getting thousands of impressions per day. Moreover, on a desktop, you're competing with dozens of other thumbnails for a viewer's attention all on the same screen. So your thumbnail is almost like a movie billboard here. We have struggled with thumbnails a lot, which is what I've been pushing myself and seem to really get our shit together on thumbnail and title making. By the way, for the purposes of this video, consider thumbnail and oversimplification of thumbnail and title, because they are really both part of the same thing. Anyway, one of the most shocking revelations I had as part of this journey to improve this stuff was this light on Mr. Beast's presentation at Vid Summit. Jimmy sent this to me. Oh, OK. Send this to all my friends before I upload a video. And I'm like, hey, which one do you like? You don't have to go this extreme. I'm probably the most extreme with thumbnails. So now on average, we're probably making like 20 versions of video. But you don't have to go that crazy. Now, Jimmy approaches this process first qualitatively. He debates with his team or text fellow trusted YouTubers to get their insight on a thumbnail.
And this is a bee testing of sorts, but it's not really the analytical type of tests that we're looking after. The problem with YouTube is that a bee testing isn't necessarily easy, because YouTube doesn't have the tools to truly compare one thumbnail against the other. In Jimmy's case, their approach is launching a video, giving it a few days. And if the performance isn't what they expected, they switch it for another one. We're much quicker about this, by the way. We know that most of the videos that we release should have a 7 to 10% click the rate when they get exposed to the most loyal viewers, or the ones that get it first. So a click the rate of 8 is generally a sign that the video will perform well in the future. If we can't hit 7% click the rate on the first few minutes, we change the thumbnail in as little as a couple of hours after the video has gone live. This is why we always make two or three thumbnails for every video, so that we're ready to change it without having to go back to the drawing board. We certainly don't have the budget to make 20 thumbnails and then pick one, but one of the best lessons that we learned from Jimmy was that they think of the thumbnail before they actually produce the video. And they only move forward with shooting the video if they've been able to come up with
a title and a thumbnail combination that you just can't help clicking. Now this all works. It obviously works for Mr. Beast. But both of these approaches are flawed. Now I can fairly say that Jimmy knows more than me or more than anyone for that matter. But while these thumbnail changing approaches are working, they are far from being true av tests. And to explain why I'm going to have to geek for the next 60 seconds or so. It's just 60 seconds. I promise. Beware with me. Because yes, you can look at how the click the rate has changed. You can look at how views change and draw conclusions on whether or not the thumbnail is better. But this isn't perfect because it isn't a truly randomized test. Most of how YouTube distributes traffic, the circumstances in which people see the thumbnail are too different. They create massive differences. If you look at the YouTube click the rate, you get this average number. That's useful. But it isn't great. Jimmy released some of his own stats recently, which are completely insane. But the point is that this number is just an average of people who saw your video on their homepage or after searching it or after seeing it on their sidebar. Even new events can create ripples on conversion rates. For example, our video on the dot com bubble tends to get a boost in views and in click
the rates when there's new talks of crises and inflation and economic turmoil. Now the better number requires you to go deeper and look at the traffic by source so you can at least compare apples to apples. And honestly, few YouTubers have the time to look at these metrics that closely. And even so, for example, on YouTube search, they click the rate various YouTube factors that you don't understand, like in which search position you racked. The higher, the better click the rate. But you don't know what the search position is. YouTube won't give you that data. Also, a thumbnail might perform great on search because it stands out from the other results. But it may perform poorly on the YouTube homepage because it's not interesting enough to fight against the other 12 thumbnails and get in that page. I told you this video was an exciteration of the experiment because changing this thumbnail daily is a bad AB test. It will not provide statistical significance on any of the outcomes. But if you're here, it's because we at least made you click. Now a partial solution to this came from a company called TubeBuddy that tries to hack YouTube's limitations by running thumbnails against each other on 24-hour sprints. It's still not perfect though. And they have this list of disclaimers and limitations on how they run these experiments.
The simplest reason here is the day of the week. And I'm just not going to bore you with the rest. It's still better than nothing, but again, it's not a true AB test. And we're not going to get one unless YouTube implements that functionality. Now, if this is all starting to sound like a blur, it's because it should be. Because this is us, YouTubers trying to understand numbers and metrics that only YouTube's own data scientist team can fully understand. And part of that relates to Google's, I mean, or Alphabet's secret algorithms, which we actually covered on our video last week. But the epitome of these tests are the ones that large tech companies perform in their apps. Because when you have the team to do it, and when your platform depends on a deep understanding of human behavior to monetize humans, AB tests become key. Let's look at TikTok, for example. Most platforms in the world would want to get your email first. That's the first thing you need an email so you can send them spam. I mean, we're engagement campaigns if they fall out of their path. Part of TikTok's genius is that they ditched the sign-up part altogether. All you need to do is pick your interests and you're ready to go. You don't even have to press play on the single video or sign-up. Make no mistake.
This efficient onboarding flow is the result of dozens, maybe hundreds of iterations, all to get you to spend more hours on the platform, which may or may not be related to psychological or even political implications. Texas Governor Greg Abbott has banned the use of the social media app TikTok on all government issued devices over cybersecurity concerns. In videos tailored to their interests, certain sections of the brain involved in addiction were lit up. The most successful video I have is like 68.000 views right now. I'm sorry, 68.000 views? Yeah. A less thin foil, but still important example is Instagram's initiative to remove the like count from the feed. In 2019, in the midst of this worldwide crisis on teenage depression, researchers and governments looked into social media as a possible cause. Facebook or it wasn't made it already. These guys were pressured to do something about it and they responded with an A.B. test. In specific countries and with specific sets of users, they started hiding the like count on other people's posts. The social media platform is conducting a test here in Canada, hiding the number of likes on photos and videos.
That meant that you could only see how many people had liked your own publication but not anyone else's. According to them, it was this effort to remove the sense of competitiveness from social media. One people that worry a little bit less about how many likes they're getting on Instagram. It's been a bit more time connecting with the people that they care about. It turned out that it didn't actually change nearly as much about how people felt or how much they used the experience as they thought it would. This is of course a vague and shitty answer and like counts are about a tiny part of the whole Instagram experience and how damaging it could be for young minds but it was an A.B. test and it was enough to ease the pressure they had. The reason I tell you this story is because the answer to these A.B. tests wasn't clicks or views or minutes on the platform. The results of a test like this translate into deep data into behavioral changes in human being is from what kind of content they post to how often they do it to whether they're depressed or not. When the only people that can understand this data work for a company whose sole motivation is more revenue, what are we really optimizing for? Well that got a little deep. So let's zoom back out into the realm of tests that we can understand.
A.B. tests are great. You should do them. Don't forget about statistical significance and don't forget to subscribe. We'll see you next week.
More episodes
More from Sidebean

How Startups Shaped Silicon Valley | Sidebean
Sidebean

HP’s long term plan reveals what AI job loss really looks like | Sidebean
Sidebean

Google Maps isn’t just a map anymore, it’s your personal scout. | Sidebean
Sidebean

Foursquare: the startup zombie that still eats your data | Sidebean
Sidebean