Click blackberry, find apple
Twenty-one sentences, no pretrained model. Drag to rotate, click a word to see its nearest neighbors, then click one of them to measure how far apart they are.
You shall know a word by the company it keeps
This quote and theory published by the Philological Society in a special book volume "Studies in Linguistic Analysis" in 1957 is what this demo puts into practice.
If you click blackberry you'll see the nearest neighbor is apple. They do not appear together in any of the 21 sentences and yet are the closest neighbors to one another. Within these 21 sentences both are used as technology-company words and as fruit, depending on context. But how are those two words or any words in this demo placed to create this context, giving the words meaning by surrounding them with similar ones?
Counting the company
To place the words in the demo we use PPMI. For each word we look at the five words to its left and the five to its right within the same sentence, and count every pairing that falls inside that range. That's the whole counting step without grammar or parsing.
Watch the row totals build up. They're what the next step needs.
PPMI stands for Positive Pointwise Mutual Information. Raw counts on their own are misleading. Common words bump into everything just by being frequent. So instead of asking how often two words appear together, PPMI asks whether they appear together more often than chance would predict.
Chance here has a precise meaning. Sliding the window across all 21 sentences produces 890 word pairs in total, counting each co-occurrence from both sides. phone turns up in 72 of them, company in 42. If the two were unrelated, you'd expect them to land together 72 × 42 / 890 ≈ 3.4 times. They actually land together 12 times, about three and a half times what chance predicts.
Now compare healthy and orange. Each appears in only 10 pairs. You'd expect them together 10 × 10 / 890 ≈ 0.11 times. They appear together in 2 sentences, which the window counts as 4 pairs.
| pair | count | expected | PPMI |
|---|---|---|---|
healthy × orange |
4 | 0.11 | 5.15 |
phone × company |
12 | 3.40 | 1.82 |
phone and company appear together three times as often as healthy and orange, yet their PPMI is about a third as high. The value itself is the base-2 logarithm of observed ÷ expected, so every +1 is a doubling. A PPMI of 5.15 means roughly 36× more often than chance. One last step turns PMI into PPMI, and it's where the extra P for "positive" comes from. When a pair appears together less often than chance predicts, the value goes negative, and PPMI sets it to zero instead. In 21 sentences, two words rarely meeting doesn't mean they're unrelated. apple and blackberry never meet at all, and they end up nearest neighbors.
From 49 numbers to three
In the demo click on the table icon in the top right to see the fully calculated PPMI table based on all words within the 21 sentences excluding common function words like "the" and "a". There you'll see we have 49 words, and therefore 49 dimensions, which is impossible to display directly. The demo shows a three-dimensional space, so we use SVD / LSA (Singular Value Decomposition, known as Latent Semantic Analysis when applied to text) to find the three strongest dimensions in the PPMI table.
Each row of that table holds one word's full position as 49 numbers, one per vocabulary word. That's already the vector. The problem is that 49 numbers can't be turned into a dot on a screen. You need three.
The obvious move would be to pick three columns and drop the rest. That throws away almost everything. Instead, SVD finds three new directions through the 49-dimensional space, picked so that the words spread out along them as widely as possible. Spread is what distinguishes words from each other. The three directions with the most spread are the ones that lose the least when you discard the other 46.
A direction here isn't a column. It's a weight for every column at once, with 49 numbers of its own, positive on some columns and negative on others. Nothing gets picked; everything gets a vote. To find a word's coordinate you multiply its row against those weights and add the results up. One number. Do it three times, with three sets of weights, and you have the x, y and z that place the dot.
The second direction comes from the same search, run on what's left after the first is subtracted, which is why sweet, fruit and green sit at PC1's positive pole and PC2's negative one. The third works the same way.
Keeping the top three isn't a rough approximation someone settled on. Eckart–Young–Mirsky proves that taking the strongest k directions gets you closer to the original than any other k-dimensional version, and three is simply the k that fits on a screen.
Apple and blackberry sit 89° apart in the full 49 dimensions. A direct comparison of their rows shows they're unrelated. After projection they're 9° apart. In the raw table they share almost no context words. Their food and tech cells sit in different columns. But those columns pull in the same direction, and the projection makes that visible.
| PC1 | PC2 | PC3 | |
|---|---|---|---|
| fruit | +6.92 | -2.11 | -0.37 |
| healthy | +4.24 | -1.33 | -0.36 |
| orange | +4.24 | -1.33 | -0.36 |
| apple | +2.65 | -0.52 | +1.44 |
| blackberry | +2.96 | -0.04 | +1.70 |
| phone | +0.39 | +0.68 | +3.15 |
Nobody names the directions. They come out of the factorization and you read afterwards what ended up at each end. I expected the first one to separate food from technology but it actually separates food from everything else, because the food words are the tightest cluster in this corpus. That's the difference between reading a result and choosing one. Switch the projection to Concept Seed Vectors in the demo and you can name the axes yourself. You pick a few words for each axis, and every other word is placed by how often it keeps company with them. You gain labels you can read but lose any chance of being surprised.
Where it breaks
For words that fall clearly into a single category, the result is recognizable clusters of hardware and food, separated by distance.
apple and fruit appear together in 2 sentences, which the window counts as 4 pairs. They still sit 1.85 apart on the demo's scale. apple and blackberry never appear together at all and sit 0.24 apart. Apple gets exactly one vector. Its tech sentences drag it away from the fruit cluster, its fruit sentences drag it away from tech. It ends up between both and belongs fully to neither. More data won't fix that. Twenty-one sentences is also nowhere near enough to place every word well.
What you've just watched is the older of two approaches. Counting and factorizing goes back to Latent Semantic Analysis in 1990, two decades before word2vec.
word2vec, which landed in 2013, works the other way round. It never builds a matrix of counts at all. It streams through the corpus and nudges vectors, closer for words that appeared together, further apart for random pairs, starting from random positions. Learning the vectors directly with a small neural network, instead of building the full count table first, is a big part of why it caught on. For a 100,000-word vocabulary that table has 10 billion cells, about 40 GB if stored in full, and it grows with the square of the vocabulary. word2vec only ever holds a few hundred numbers per word.
It was fed roughly 100 billion words from Google News articles. That's 5 to 6.7 billion sentences, or 1 to 1.25 million books at 80,000 to 100,000 words each. Those vectors have 300 dimensions rather than three, room enough to separate clusters this corpus can only blur.
The two turned out to be closer than they looked. In 2014 Omer Levy and Yoav Goldberg showed word2vec is implicitly factorizing a shifted version of the PMI matrix, which is the table above before its negative values are set to zero. A year later, with Ido Dagan, they showed that once you port word2vec's tuning tricks back into the count-based method, PPMI and SVD perform comparably.
Current frontier models are large neural networks trained on far more data, with several thousand dimensions depending on the model. Their embeddings are contextual, computed from the text they're reading, so they can pull apple toward tech or food depending on the surrounding words. That solves the limitation above.