Skip to content
TrackPodcasts
historyApr 7, 202619:33

The Mathematical Geography Of Word Embeddings

pplpod

About this episode

The Mathematical Geography Of Word Embeddings

Get every episode summarized

Each time pplpod publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

Transcript ready

518 searchable segments. Every word is indexed and playable.

The Mathematical Geography Of Word Embeddings

pplpod

0:00
19:33

Full transcript

pplpodThe Mathematical Geography Of Word Embeddings. Machine-transcribed; use the interactive transcript above to jump the player to any line.

So if you sit down at your computer or you know, you pull out your phone and you type a question into a search engine or a large language model, there is this kind of everyday magic that happens. Yeah, it really does feel like magic sometime. Right, because you type out a sentence in just plain English and the machine just gets it. It understands you. But I mean, computers do not read English, they don't understand words at all. No, not even a little bit, they only understand math. Exactly, they only understand numbers. So how do you bridge that massive gap between human thought and cold hard computation? Well, it is arguably the most important bridge we have ever built in computer science. And to give you an idea of just how mind bending the math on the other side of that bridge actually is, think about this. If you take the mathematical coordinates for the concept of a king, subtract the concept of a man and then add a woman, a computer will instantly spit out the exact coordinates for the word queen. Oh, wow, that is absolutely wild. It's literally doing arithmetic with human concepts.

Yeah, it is. Well, welcome to today's deep dive where we are going to explore exactly how that happens. We are looking at a massive stack of research on a fascinating natural language processing concept called word embedding. And this is really for you, the learner out there who just wants to know how the digital world ticks. It's a great topic. It is. Our mission today is to demystify the invisible mathematical architecture that allows computers to quote, unquote, understand human language. Because this is the secret sauce behind the entire modern AI revolution. Okay, let's unpack this. I'm ready. It really is the foundational pillar of modern natural language processing. So at its core, a word embedding maps a word to a real valued vector. Right, which sounds very technical. It does, but to translate that from computer science we can just plain language. A vector is essentially just an ordered list of numbers. And what these numbers do is they define a specific coordinate in this massive mathematical space.

The fundamental goal here is to encode meaning into those numbers in such a way that words with similar definitions are physically grouped closer together in that vector space. I love trying to visualize this because a list of numbers sounds so dry. But if you think about it spatially, it's just incredible. I like to picture this vector space as like a giant multi-dimensional map of towns. Yeah, that's a great way to look at it. Right, so if we have a town called Happy, the town right next door is gonna be joyful. Exactly. They share the same local weather, the same neighborhood. But if you look for the town called Sad, well, that's gonna be all the way on the opposite coast. It's geographically distant because it's semantically distant. Yeah. We are literally turning meaning into geography. And that spatial geography analogy is incredibly accurate to how the mass actually functions behind the scenes. And to really understand how a machine plots out that map today, we first kind of have to understand how humans have historically tried to define what meaning even is.

Right, because that's a philosophical question almost. It is. And there is a famous quote from 1957 from a linguist named John Rupert Firth. He said, a word is characterized by the company it keeps. The company it keeps. I love that. So meaning isn't something a word just possesses in a vacuum. No, not at all. It's completely defined by the surrounding words that constantly hang out with it. Exactly. I mean, if you see the word bark next to tree and branches, it implies one physical reality, right? Sure. But if you see it next to dog at least, it implies a totally different one. What's fascinating here is that this foundational thought from cognitive psychology and early linguistics from the 1950s, it still dictates how modern neural networks operate today. That's great. It's literally teaching machines to read by having them study the company that words keep across just massive data sets. Wait, hold on, I have to ask about the timeline here, because Firth's quote is from 1957. Yeah. If they had this whole company a word keeps theory back during the Cold War,

why did this technology just sit in a drawer for 50 years? Like, if researchers knew the secret to mapping meaning, why didn't we have smart search engines and chat bots back in the 80s or 90s? Well, that is the perfect question. And I can tell you, it wasn't for lack of trying. Oh, they tried. Oh, yeah. Early computer scientists did try to build these vector spaces in the 60s and 70s for information retrieval. But they ran headfirst into a massive computational road block, which is known as the curse of dimensionality. The curse of dimensionality, that sounds terrifying. What were scientists actually dealing with there? Well, imagine you were trying to map the company a word keeps by building a giant spreadsheet. OK, I'm picturing a spreadsheet. You want to count every single time a word appears next to every other word in the English language. That's a lot of words. Exactly. In those early models, every single unique word in your vocabulary required its own distinct column, its own dimension. So if you had a working vocabulary of 100,000 words, for your mathematical space had 100,000 dimensions.

Right. And the result was a wildly sparse space, meaning if you look at the row for the word cat, you might have a few numbers in the columns for meow or milk. But the other 99,998 columns in that spreadsheet were just zeros. Oh, wow. So they were trying to brute force a spreadsheet where almost every single cell was totally empty. Pretty much. And the computers of the 70s and 80s just choked on the sheer volume of nothingness. The calculations became impossibly heavy. There was just too much empty space to process efficiently. So the obvious question was, how do we compress this data without losing the underlying meaning? Right. In the late 1980s, scientists introduced something called latent semantic analysis. They used a mathematical technique, specifically singular value decomposition, to squash that massive sparse space into a much smaller, denser number of dimensions. Singular value decomposition. That is a heavy string of words. What is actually happening to our map of towns when they do that? Think of it like shining a very bright flashlight

on a highly complex three-dimensional object and looking at its two-dimensional shadow on the wall. OK, I can picture that. You are technically losing a dimension of data, right? But you still keep the core shape and the vital structural relationships of the original object. Well, I see. It forced the data to compress, finding hidden or latent semantic relationships. Then around the year 2000, a researcher named Yoshua Benjyo pushed this even further. He and his colleagues used early neural networks to learn what they called a distributed representation. OK. This drastically reduced that high dimensionality into a tight dense vector of maybe just a few hundred dimensions. And getting it down to a few hundred dimensions is where the magic really starts to take off, right? Absolutely. Because then we hit 2013, which is widely considered a massive pivotal milestone in this field. A team at Google, led by Tomas Michalof, created a toolkit called Word to Vec. Yes, Word to Vec. It seems like this was the moment the dam finally broke. Because Word to Vec trained these dense vector maps

significantly faster than anything before it, it really moved Word embeddings from being this new, expensive academic experiment into something you could actually use in practical software. It democratized the map making process for language, absolutely. Word to Vec proved that you could train high quality Word embeddings on huge data sets very efficiently. Right. And the way it trains is brilliant. It essentially plays a massive game of fill in the blank. The algorithm looks at a sentence, hides one word, and then looks at the local context window, which is the few words immediately before and after the blank and tries to guess the missing word. Oh, wow. So it's actively testing itself on the company of the word keeps. Yes. And every time it guesses wrong, it reaches into its own internal wiring and slightly adjust the mathematical dials, the weights, to make a better guess next time. That's so cool. Over millions and millions of iterations reading massive chunks of the internet, those internal dial settings actually become the coordinate numbers. The network's attempt to predict neighbors generates the map.

OK, so Word to Vec is incredibly fast, and it builds this beautiful dense map. But human language isn't just vast. It's incredibly messy. It's ambiguous. It really is. Here's where it gets really interesting. Think about a word like club. OK. If you say to me, the club I tried yesterday was great. What are you actually talking about? Right. It could be anything. Are you talking about a turkey club sandwich? Are you talking about a loud night club downtown? Or are you talking about a nine iron golf club? This brings us to a major hurdle. In linguistics, this is known as polysomy words, with multiple related meanings and homonomy words that sound the same but have totally different meanings. Yeah. And this was the fatal flaw of early static embeddings like Word to Vec. They conflated all of those distinct meanings into a single point on our map. That's a huge geographical trap. If the algorithm tries to map club, and it averages out the sandwich, the golf club, and the night club, it's going to drop the pins somewhere in the middle of the ocean.

Exactly. It forces a single word to be a single vector, which means it is geographically lost because it's trying to be in three places at once. Which means it fails to understand the specific sentence you are typing. So the field had to adapt. Researchers began developing multi-sense embeddings. Approaches like multi-sense skip gram or MSG tried to handle this. But a really robust solution was MSSA or most suitable sense annotation. OK, how did that work? It brought in outside knowledge. MSSA relies on pre-existing lexical databases, things like WordNet. These are essentially massive structured digital dictionaries that already know the different definitions of a word exist. Wait, if club has three different meanings, how does the computer know which one I mean in real time before it does the map? By looking at a predefined sliding window of context around the word, MSSA uses a knowledge-based approach to label the specific sense of the word before it calculates the math. Oh, I get it. Once the computer knows you are talking about the sandwich because words like turkey and bacon are nearby,

then it produces a distinct multi-sense embedding just for the sandwich definition. So essentially teaching the computer to ask, wait, which definition are we using right now before it plots the point? Exactly. That makes total sense. But even that isn't the final form of this technology. The research moves into the late 2010s, introducing the modern marvels that power the math of AI tools we use today like Elmo and Bert. Yeah, this was the true paradigm shift. Because even with WordNet, you are still relying on a fixed dictionary. Great. Models like Bert, which stands for bi-directional encoder representations from transformers, they create the embedding on the fly at the token level. Every single occurrence of a word gets its own unique mathematical coordinate based entirely on the specific sentence it appears in at that exact moment. So if I write a whole paragraph about playing a round of golf, and I use the word club five different times in slightly different ways, Bert isn't just pulling a prepackaged golf club vector from a dictionary file.

Not at all. It actively reads the entire sentence surrounding that specific instance of the word club. It weighs the importance of every other word in that sentence simultaneously and generates a brand new coordinate right then and there. That is amazing. It places occurrences of the same exact word in totally different regions of its embedding space, depending on the nuance of the surrounding sentence. It finally mastered Firth's rule from 1957. It truly judges the word entirely by the company it is keeping at that exact second. That dynamic on the fly mapping is why chatbots suddenly got incredibly good at understanding context over the last few years. Yeah, it was a game changer. Okay, so we've mapped out how these vectors decode the grammar and the meaning of human language. But if language is really just a sequence of symbols, what other sequential data can we translate into a vector space? Because the research takes a while turn here. It really does. It introduces something called biovectors, developed by researchers as Gary and MoFrad. They realized we can treat biological sequences

like DNA, RNA, and amino acids as words. It is a brilliant unconventional application of the technology. They created ProteVec for proteins, which are amino acid sequences, and Genevec for gene sequences. They essentially treat chunks of these biological building blocks as N-grams. And just to clarify, an N-gram is basically just a continuous sequence or chunk of items from a given sample, right? Like taking a long string of letters and breaking it into three letter words. Precisely. They took massive databases of genetic code, broke them down into these N-gram words, and asked the algorithm to read the code of life using the exact same mathematical tools we use to read emails. But how does reading a sequence of amino acids like text actually translate to physical biological reality? Think back to the text analogy. Just as the word bark constantly appearing next to tree implies a semantic relationship. A certain sequence of amino acids constantly appearing next to another sequence implies a physical relationship. Oh, interesting. It implies they are highly likely to fold together physically

in a three-dimensional protein structure. The neural network maps out the towns of genetic code and discovers the biophysical rules of how they interact without a human biologist having to explicitly program those rules into the machine. That is staggering. It learns the physical language of biology just by looking at the company of the genes keep. And biology isn't the only non-human language we are translating. There are also incredible applications in game design. Oh, yes. Researchers, Rybee and Cook proposed a way to discover emergent gameplay by using logs Yes, another fascinating translation. To do this, you transcribe the actions that occur during a game like a chess match into a formal language. Every move, like moving upon from E2 to E4, is logged as a word in a sequence. A full game is a sentence. Millions of games become a massive book. Once you have that resulting text, you run it through the same word to vet embedding process. If we connect this to the bigger picture, the resulting vectors capture expert knowledge

about games that are not explicitly stated in the games official rulebook. Right, because the rulebook of chess just tells you how a knight is allowed to move. It doesn't tell you the strategic value of controlling the center of the board. Exactly. But the algorithm learns that advanced strategy, simply by observing millions of sequences of moves, it learns the unspoken strategies, the emergent dynamics, just by mapping out the coordinates of the game logs and seeing which moves constantly hang out together in winning games. Which brings us to a critical realization. If word embeddings are incredibly efficient at capturing the unwritten emergent rules of chess and the unspoken biophysical laws of proteins, we have to confront how they capture the unwritten rules of human society. So what does this all mean? If they learn exactly what they observe, what happens when we look at the data we are feeding these massive models? Well, this is a major area of study within natural language processing. The researchers point out that word embeddings inevitably contain the biases and stereotypes

contained within their training data sets. A landmark 2016 paper by Bolit Bosse and several colleagues analyzed a publicly available word-to-vec embedding that had been trained on a massive corpus of Google News Texts. And it is worth pointing out that Google News Text is written by professional journalists. This wasn't trained on random toxic internet forum comments. It was considered a highly professional, clean data set. It was. But when the researchers used this trained embedding to extract word analogies using that exact same king-minus-man plus woman math, we talked about at the beginning, they found deeply disproportionate word associations. According to the paper, the model generated analogies like man is to computer programmer as woman is to homemaker. The math literally calculated a stereotype. It met the town of man physically closer to programmer and the town of woman physically closer to homemaker. And it did this purely because of the historical structural biases present in the millions of news articles it read. It absorbed societal bias and encoded it

as a factual geographical rule of language. This raises an important question regarding how these tools are deployed. Research by G.U. Zoo and others demonstrates that if these trained word embeddings are applied to real-world tasks without careful oversight, the algorithms do not just passively reflect the bias found in that unaltered training data. They actively perpetuated. Yes, the researchers concluded that word embeddings can even amplify these existing biases when they are integrated into automated systems, like resume screening tools or search engine rankings. So according to these studies, we are feeding them our history. They're encoding that history into rigid mathematical coordinates and then applying those rules to our future. It really reframes how we look at the data we provide them. It really does. The researchers who publish these papers stress that it requires a massive ongoing effort in the field to actively devise these embeddings. I'm going to even do that. They essentially have to mathematically adjust the map after it is drawn, shifting the coordinates

to ensure that the vector for a neutral profession like programmer is equidistant to both man and woman. It is a complex technical challenge that the industry is still grappling with today. We have covered a massive amount of ground today. You've taken quite a journey with us. We started all the way back with John Rupert Firth's 1957 linguistic theories, the foundational idea that words are completely defined by their neighbors. A long time ago. Yeah. We watched the computer scientists of the 80s and 90s battle the terrifying curse of dimensionality, finally shrinking those massive spreadsheets using singular value decomposition and breaking through with Google's word-to-vec predicting hidden words. We navigated the tricky, multi-meaning puzzle of the word club and saw how modern tools like Bert use dynamic on the fly context windows to solve it. We even explored how this exact same math is decoding the physical folding of DNA, capturing the hidden strategies of chess, and reflecting the biases

hidden in our own professional data sets. It represents a profound shift in how we interact with technology. Word embeddings are the hidden geography of meaning. They are the underlying infrastructure, shaping the digital tools you interact with every single day, turning abstract human thought into computable math. Absolutely. But I want to leave you with a final lingering question to explore on your own. Think about the digital footprint you create every single day. The sequence of apps you open, the locations your GPS tracks, the transactions you make. OK. If word embeddings are powerful enough to capture the unwritten rules of chess, the folding of proteins, and the invisible patterns of human bias, just by looking at the company your word keeps. What unspoken rules about your life and habits might be mapped out? If the daily digital trails you leave behind, we're embedded into a vector space. What a thought. Your whole life, your routines and secrets mapped out in a multidimensional space. Maybe the town of morning coffee is mathematically right next door to the town of checking emails.

That giant multidimensional map of towns isn't just for language anymore, it's mapping us. Thank you for joining us on this deep dive. Stay curious, and we will catch you next time.

More episodes

More from pplpod

View all episodes →