
About this episode
Get every episode summarized
Each time pplpod publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Transcript ready
464 searchable segments. Every word is indexed and playable.
Full transcript
pplpod — The Mathematics of Close Enough. Machine-transcribed; use the interactive transcript above to jump the player to any line.
When we look at the modern world around us, there is this profound illusion of absolute precision. Well, absolutely. We like our reality to be crisp. Right. I mean, we think of skyscrapers, engineers down to the microscopic millimeter or digital bank transfers, moving exact pennies across fiber optic cables. We just expect things to be perfect. You punch a calculation into your phone and the screen just confidently says, here is the exact answer. Yeah, it's a very comforting illusion. We inherently trust the numbers because they feel binary, you know, it's either right or it's wrong. But then, then you peek behind the curtain at the foundational math, running our physical and digital lives, and suddenly that pristine calculator starts glitching. You realize we're looking at a computational landscape that is honestly just an entire universe of close enough. It really is. It's the mathematics of approximation. And navigating a world built on tiny, necessary compromises is what it's a lot more complicated
than most people realize, which is exactly why we're here. So welcome to today's deep dive into the source material. We've got a fascinating stack of research today, custom tailored just for you. We do. Our mission today is to explore the anatomy of being wrong or more specifically how we mathematically quantify and control those tiny compromises in the systems you use every single day. Our grounding source for today's exploration is a highly detailed text on approximation error. And, you know, whether you are a scientist measuring volatile chemicals in a lab or a programmer building an app or just someone wondering why your GPS suddenly thinks you're driving into a lake. Right. Which happens way too often. It does. But understanding the hidden mechanics of error is absolutely crucial for all of it. Okay. Let's unpack this. Before we can understand how these tiny errors actually ruin algorithms or crash-grip computer systems, we have to define the basic language of a mistake. Right. According to the source, there are two primary ways we measure discrepancy between an exact
quote-unquote true value and our approximation of it. Yeah. And what's fascinating here is that the language of mathematics gives us very specific tools to bound our mistakes. So the first tool is what we call absolute error. Okay. This denotes the direct numerical magnitude of the discrepancy. It doesn't care about the scale of the universe. It just cares about the raw distance between the truth and the guess. So it's just the straight-up difference. Exactly. Formally, if you have a true value and an approximated value, the absolute error is bounded by a positive value, which is typically represented by the Greek letter epsilon. Meaning, if I measure something, the absolute error is just how far off my measurement is. Plain and simple. And whether I overestimated or underestimated doesn't really matter, right? Which is why the math uses absolute value bars. Precisely. It's just the magnitude of the gap. The text has a great, simple example for this. Imagine you are measuring a piece of paper. The actual true length of this paper is exactly 4.53 centimeters, but you're using like
a standard plastic ruler from a school desk that only lets you estimate to the nearest tenth of a centimeter. So you write down a recorded measurement of 4.5 centimeters. Right. Because that's the best your tool can do. Exactly. Error there is exactly 0.03 centimeters. And that seems tiny, right? I mean, 300 of a centimeter is barely a speck of dust. Yeah, it's nothing. But relying solely on absolute error leaves out a critical piece of the puzzle, which brings us to the second measurement. And that's relative error. Okay. So how does that work? Relative error provides a scaled measure. It takes that absolute error and considers it in proportion to the exact data value. Mathematically, you divide the absolute error by the magnitude of the true value. Oh, and if you multiply that by 100, you get your standard percent error, right? Like most of us remember from high school chemistry. Assuming the true value isn't zero, yes. And the source brilliantly illustrates why relative error is often the much more important metric.
I love the comparison they use for this. So let's say you are approximating the number 1000 and you make an absolute error of three. Okay, you're off by three units. Right. Relative error of 0.3%. But what if you are approximating one mix in with that exact same absolute error of three? Well, in that case, your relative error drops to a mere 0.0003%. The absolute error is the exact same raw mistake in both scenarios. You are off by three. Yeah. But the relative error provides the context like being off by three when you're counting a thousand dollars out of a cast register, that's a noticeable problem. For sure. Your boss is going to notice that. Exactly. But being off by three when you're counting a million dollars, that is basically a rounding error that no one is going to lose sleep over. So relative error is the context dependent assessment of how much that mistake actually hurts us. If we connect this to the bigger picture, this contextual understanding is vital, but it also introduces some massive mathematical traps if you apply it to the wrong kind of
measurement scale. Ooh. I am always fascinated by mathematical traps. What is the first one? The first one is pretty simple. Division by zero. Relative error becomes mathematically undefined if the true value is zero. Right. Because you can't divide your absolute error by zero, it just breaks the arithmetic. Exactly. But the second caveat is much more insidious. Relative error only truly works and is only consistently interpretable if the measurements are performed on a ratio scale. Hold on a ratio scale. I know there are different types of measurement scales. A ratio scale means a scale that has a true non-arbitrary zero point, like a zero that signifies the complete physical absence of the thing you're measuring. That is spot on. Think about temperature, which is a classic trap detailed in our source. Let's look at the Celsius scale. Okay. Zero degrees Celsius doesn't mean no temperature or no heat energy. It is literally just the arbitrary freezing point of water. So it's an interval scale, not a ratio scale. Right. It's just a convenient benchmark.
Exactly. The true temperature in the room is two degrees Celsius, but your thermometer approximates it as three degrees Celsius. Okay. So your absolute error is one degree Celsius. Yes. Which means your relative error is one divided by the true value of two. That's 0.5 or a 50% relative error. A 50% error sounds catastrophic for a scientific measurement. Like you completely watched the experiment. It really does. But let's take that exact same physical room, that exact same physical temperature and that exact same physical mistake and measured on the Kelvin scale. Okay. Because the Kelvin scale is a true ratio scale. Exactly. Zero Kelvin is absolute zero, the complete theoretical absence of thermal energy. So what is two degrees Celsius in Kelvin? I am definitely not doing that mental math. What is it? It translates to 275.15 Kelvin. I agree. Dig number. Right. And because the increments are the exact same size, an absolute error of one degree Celsius is exactly equivalent to an absolute error of one Kelvin. Right. The physical distance on the thermometer didn't change.
Exactly. So now we divide your one Kelvin absolute error by the true value of 275.15 Kelvin. Okay. One divided by 275. That's going to be tiny. It's roughly 0.00363, which is about a 0.36% relative error. Yeah. You go from a 50% error to a fraction of a percent, describing the exact same physical reality just by changing the scale. That is wild. So if you use relative error on a scale without a true zero, the percentages are essentially just meaningless noise, completely meaningless. And the text points out another quirk about how arithmetic operations interact with these two types of errors, which feels almost counterintuitive until you really break it down. Oh, you mean the multiplication versus addition quirk? Yeah. Let's say you have your true value in your approximation and you multiply both of them by a constant number. Well, when you multiply by a constant, the absolute error changes directly. Like if you multiply everything by 10, your absolute error becomes 10 times larger.
Because the raw distance between the numbers just scales up. Exactly. But your relative error stays completely identical. Right. Because the constant cancels out in the ratio, you multiply the top of the fraction of the bottom of the fraction by 10. So the proportion remains exactly the same. But if you add a non-zero constant to both the true value and the approximated value, the reverse happens. Absolute error is completely insensitive to addition. Let me make sure I follow. So if the true value is 10 and the guess is 8, the absolute error is 2. And if you add 100 to both, the true value is 110, the guess is 108. The absolute error is still simply 2. Exactly. The raw gap didn't change. But the relative error gets completely worse. Oh, I see. Originally, it was 2 over 10, which is a 20% error. Now it's 2 over 110, which is barely a 1.8% error. Exactly. Just by adding a constant, you've artificially deflated your relative error, creating this false sense of accuracy without actually improving your measurement at all.
That feels like a trick someone could use to manipulate data. Oh, it absolutely is. Which raises an important question. What happens when we move from theoretical math on a chalkboard to the physical tools and digital machines we use every day? Right. How do real world instruments actually handle these errors? The source dines into this, particularly looking at instruments like analog voltmeters, pressure gauges, and thermometers. Manufacturers of these indicating measurement instruments frequently guarantee their accuracy, not as a percentage of the actual reading you're looking at, but as a percentage of the instruments full-scale reading capability. These are known as limiting errors, right? Or guarantee errors. Exactly. And they create a very dangerous dynamic for the user. It implies the maximum possible absolute error is fixed based on the top of the scale. Meaning the relative error can spike dramatically when you are measuring small values. Precisely. The text uses a laboratory beaker to illustrate this point. Yes. The beaker example is perfect. If you have a beaker that holds a maximum of 6 milliliters and you measure out 5 milliliters
when the true volume is actually 6, your percent error is roughly 16.7%. Okay. Great, but manageable. But if the true volume is only 1 milliliters and the instrument maintaining that same full-scale absolute error guarantee reads 2 milliliters. Your relative error is now 100%. Exactly. Using an instrument at the very bottom of its scale is a recipe for disastrously high relative errors. Because the physical margin of error on the glass doesn't shrink just because you're measuring less liquid. Exactly. I imagine this physical limitation translates directly into the digital realm too. It absolutely does. Digital systems and computers cannot represent all real numbers perfectly. I mean, they have finite memory. You can't store an infinitely repeating decimal in a machine with fixed RAM. This leads to unavoidable truncation or rounding errors in almost every single calculation. We call that machine precision, right? The computer literally just runs out of space to hold the numbers and has to mathematically
chop the end off. Yes. And this introduces a crucial concept in numerical analysis called numerical stability. Okay. I remember this from the text. Numerical stability measures how much a specific algorithm allows those tiny initial rounding errors to propagate and amplify into substantial errors in the final output. So a numerically stable algorithm is robust. It suppresses the noise. Like if the input is slightly chopped off or malformed, the output is still pretty close to the truth. Exactly. It dampens the error. But a numerically unstable algorithm exhibits dramatic error growth. It's like the butterfly effect of mathematics. Oh, wow. A microscopic change in the input, just a tiny approximation error, can cascade through the algorithm's internal steps and render the final results completely unreliable. So what does this all mean for computer science? If algorithms are inherently saddled with these rounding errors from the very first nanosecond, how do programmers mathematically guarantee that an algorithm will spit out a reasonably
accurate approximation before the heat death of the universe? That brings us to the realm of computational complexity theory and polynomial time approximation. For an algorithm to be practically useful, it needs to be able to compute an approximation within a specific time limit. This time duration must be polynomial, meaning it stales at a reasonable manageable rate in relation to two things. The size of the input data and the encoding size of the error bound. Whoa, hold on. Before we lose the listener, let's back up. What does encoding size actually mean in plain English? Because the source throws out terms like big O of log 1 over epsilon bits. And that sounds incredibly dense. It sounds super intimidating, but it's actually just about computer memory. The encoding size is simply how many bits of memory it takes to tell the computer how precise we need it to be. Oh, okay. If you want the absolute error, meaning the smaller your epsilon is, the more bits you need to encode that tiny, tiny fraction, and the exponentially longer the algorithm takes
to run. That makes perfect sense. We say a value is polynomially computable with absolute error if we can algorithmically find an approximation within that reasonable time limit for any specified maximum error. And the exact same concept applies to relative error. Okay. Here's a question that tripped me up while reading the text. Let's hear it. The computer has an algorithm that can successfully solve for a relative error in polynomial time. Does that automatically mean it can use that same efficiency to solve for an absolute error? Like are they mathematically interchangeable? It is a brilliant question, and the answer is a strict one-way street. The source provides fascinating proof sketch for this. Okay. Polynomial computability with relative error implies polynomial computability with absolute error, but not the other way around. Okay. I need a visual for this. How does solving the relative side guarantee the absolute result? Let's use a metaphor. Imagine you are throwing darts at a wall in the dark, trying to hit a microscopic bull's eye.
Okay. I'm with you. That bull's eye is your target absolute error, your epsilon. You start by throwing a special relative error dark. You basically tell the algorithm to just find an approximation with a massive, loose relative error of one half, or 50 percent. Okay. So you're just asking the computer to get you into the right bullpark. You aren't asking for extreme precision yet. Precisely. The algorithm spits out a rough rational number approximation. Let's call it dart number one. Because we know this dart landed within a 50 percent relative error of the true bull's eye. A mathematical principle called the reverse triangle inequality allows us to deduce a crucial boundary. Okay. What does that boundary tell us? That the absolute magnitude of our unknown bull's eye cannot possibly be larger than twice the magnitude of where dart number one landed? I see. Even though we don't know exactly where the true value is, our rough, sloppy guess just allowed us to draw a firm circle on the wall. We now have a mathematical ceiling on how big the true value could possibly be.
Yes. And because that first throw was so loose, the computer calculated it incredibly fast. Right. Didn't take a million years. Exactly. Number two, we invoke the exact same relative error algorithm, but this time we give it a much tighter specific target. We set our new target based on our desired absolute error divided by that ceiling we just drew. Oh, that is so clever. You use the rough boundary from the first throw to calibrate the precision of the second throw. Exactly. And because you proved the true value is trapped inside that boundary, the math cancels out perfectly. The final distance between the true value and your second dart is mathematically guaranteed to be less than your absolute error target. You basically used a relative error tool to perfectly solve for an absolute error bound. It is an incredibly elegant workaround, but as I said, it is an asymmetrical relationship. Absolute does not imply relative. If you only have a dart that computes absolute error, you cannot guarantee you can compute
relative error efficiently. Why not? Why doesn't the trick work in reverse? Because relative error is a proportion based on the true value. If you don't know the true value and you don't have a floor beneath it, your absolute error algorithm might be hunting for a proportion of a number that is infinitely approaching zero. Ah, and the calculation would just take forever. Exactly. There is one significant exception though. You can do it if you can compute a positive lower bound on the magnitude of the true value. Meaning you have to be able to mathematically prove that the true value is strictly greater than some positive number. If you know the floor is solid, you can reverse the math we just did. But if you don't have that floor, the absolute algorithm is flying blind when it comes to relative proportions. This profound asymmetry brings up a special class of algorithms the source mentions, the fully polynomial time approximation scheme, or FPTAS. I notice this. For an FPTAS, the time complexity doesn't just scale with the logarithm of the bits. It scales polynomially with the reciprocal of the relative error itself.
Yes, and the text points out that this specific dependence is the defining characteristic that makes an FPTAS uniquely powerful compared to weaker approximation schemes. And if we really want to blow this wide open, we have to recognize that everything we've discussed so far, rulers, thermometers, absolute value bracket startboards, has been entirely focus on scalar numbers. Single one-dimensional values. Exactly. The real world and the software that runs it is rarely one-dimensional. The final climax of the text generalizes all of these definitions into higher dimensions. We move from single variables to massive end-dimensional vectors, matrices, and normed vector spaces. When you are quantifying the distance between a true complex matrix and an approximated matrix, you can't just slap simple absolute value brackets on it anymore. Yeah, this is all well and good for measuring a single straight line, but what happens when an AI is trying to compress an image with millions of pixels simultaneously?
How do you measure the error of a million different points at once? The source says we have to replace absolute value with vector norms. Yes, and we have several different types of norms depending on what kind of error we care about. Okay, walk me through them. The text lists the L1 norm, which is just the sum of absolute component values. And there's the L2 norm, also known as the Euclidean norm, which measures the straight-line distance through the multidimensional space. And the source specifically highlights the Frobenius norm, which is used heavily in image processing. Yeah, let's stick with the image compression example. When you take a giant high-resolution original image, which is mathematically just a massive matrix of pixel values, and you compress it into a small JPEG file, you are creating an approximation. So how does the norm measure the error there? The Frobenius norm acts like a mathematical blanket thrown over the entire image. It calculates the square root of the sum of the absolute squares of all the differences. Oh, I see.
Essentially, it gives you an average measure of the overall multidimensional error across the entire matrix. It tells you if the JPEG generally looks like the original. But then there's the L infinity norm, which works completely differently. Okay, how so? Instead of a blanket, the L infinity norm acts like a highly sensitive alarm system. Like it's looking for the worst case scenario. Precisely. It doesn't care about the average. It scans the entire matrix and looks exclusively for the single largest absolute difference. Wow. It finds the one pixel that is the most wrong and defines the error of the entire matrix based on that single worst case scenario. That is fascinating. Having these different tools allows computer scientists to define absolute and relative error across an entire landscape of data simultaneously, whether they care about the average performance or, you know, protecting against the worst case outlier. Which is fundamental to modern statistical modeling, artificial intelligence and machine learning. So let's bring this all together for you. We started this deep dive looking at a simple plastic ruler measuring a piece of paper
to the nearest millimeter. We unpacked the profound difference between the raw physical mistake of absolute error and the crucial proportionate context of relative error. We navigated the traps of interval scales, exploring how measuring temperature in Celsius instead of Kelvin can warp a tiny fraction of a percent error into a 50 percent disaster. We saw how the physical tools we rely on from analog car speedometers to laboratory beakers inherently introduced limiting errors that become incredibly dangerous at the bottom of their scales. And finally, we followed the math into the digital realm, exploring how algorithms use polynomial-time boundaries and vector norms to guarantee that their multidimensional approximations won't spiral completely out of control. Which leads us with one final provocative thought. I like the sound of that. It's grounded entirely in the text's brief but chilling mention of numerical stability. As we noted, numerically unstable algorithms may exhibit traumatic error growth from incredibly
small input changes. That mathematical butterfly effect we talked about. Exactly. Consider the vast and imaginably complex digital algorithms running our modern world right now. They dictate global financial markets, trading millions of times a second based on predictive matrices. They optimize the flight paths of thousands of aircraft currently in the sky. They balance the real-time load of power grids across entire continents. And all of them are relying on approximations. All of them are utilizing floating-point math, matrices, and bounded errors. So the question to Ponder is this. How many catastrophic cascading failures in our world today didn't start with a massive obvious mistake? Oh, man. How many systemic crashes, flash crashes in the stock market, or inexplicable regional grid failures, began simply with a microscopic, mathematically unavoidable approximation error. An error that an unstable algorithm quietly caught in the dark and relentlessly amplified
until the illusion of precision shattered entirely. It definitely makes you look at the calculator on your phone a little differently, doesn't it? Thank you for joining us on this deep dive into the source material. Keep questioning the numbers, keep exploring, and we'll catch you next time.
More episodes
More from pplpod

How Nirvana Accidentally Changed Music Forever
pplpod

Whiskey Myers: How the "Yellowstone Effect" built a multi-platinum southern empi...
pplpod

George Jones: How an 8 mile lawnmower ride & a bridge crash built the greatest v...
pplpod

Molly Tuttle: How a prodigy shattered the "Guitar God" glass ceiling & hacked he...
pplpod