Embeddings -- Words as Coordinates
Knowledge
How does an LLM "understand" that "dog" and "cat" are similar, but "dog" and "tax return" have little to do with each other? The answer is embeddings. An embedding turns every word (or every token) into a list of numbers -- a so-called vector. These numbers describe the meaning of the word in a mathematical space.
iImagine...
Every word has an address in a huge city. Words with similar meanings live in the same neighborhood. "Dog," "cat," and "hamster" all live on "Pets Street." "Car," "bus," and "bicycle" live in the "Vehicles" district. The more similar two words are, the closer their houses stand.
Click on a point to see similarities — or enable vector math
Positions and similarity values are simplified. Real embedding spaces have hundreds of dimensions — this is a 3D projection.
Understanding
In reality, this "city" does not have just three dimensions (length, width, height) like our world, but 768 to 4,096 dimensions. That is unimaginable for us humans, but mathematically no problem. A computer can easily calculate how close two points are in a space with thousands of dimensions.
The beautiful thing is: you can even do math with embeddings. The most famous example is:
King - Man + Woman = Queen
The meaning of words thus becomes something computers can work with.
Apply
You encounter embeddings in everyday life more often than you think. When Spotify suggests similar songs, when Netflix recommends movies, or when you search in an app for "affordable hotel on the beach" and also get results for "budget room on the coast" -- embeddings are often behind it. They enable semantic search: searching by meaning rather than exact words.
In the AI world, embeddings are also the foundation for RAG (Retrieval-Augmented Generation) -- a technique where LLMs access your own documents. More on that later.
Reflect
What best describes an embedding?