Embeddings in Detail
Knowledge
In the Beginner course, you learned that embeddings turn words into numbers. Now let's look at how this works mathematically and why it's so powerful.
An embedding is a vector -- an ordered list of numbers. Modern embedding models like text-embedding-3-large from OpenAI produce vectors with 3,072 dimensions. That means every word, sentence, or document is described by 3,072 numbers. These numbers encode semantic properties -- dimensions of meaning that the model learned during training.
iDimensions are not categories
A single dimension in an embedding doesn't correspond to a human-readable concept like "animal" or "color". They are abstract, mathematically learned features. Only the interplay of all dimensions together produces meaning.
Understand
Vector spaces and semantic proximity
Imagine a three-dimensional space -- like a room with length, width, and height. Every word has a position in it. Words with similar meanings are close together. Now expand that room to 3,072 dimensions. The principle stays the same: Semantically similar concepts have similar vectors.
Click on a point to see similarities — or enable vector math
Positions and similarity values are simplified. Real embedding spaces have hundreds of dimensions — this is a 3D projection.
Cosine Similarity -- The measure of similarity
How do you measure whether two vectors are "similar"? The most common method is cosine similarity. It measures the angle between two vectors:
- 1.0 = identical direction (maximum similarity)
- 0.0 = perpendicular (no relationship)
- -1.0 = opposite direction (antonym)
The formula:
cosine_similarity(A, B) = (A . B) / (|A| * |B|)
The advantage over Euclidean distance: cosine similarity is independent of vector length. A long sentence and a short word can still have high similarity if they point in the same "direction".
Two embedding vectors have a cosine similarity of 0.92. What does that mean?
Apply
Where embeddings are used in practice
- Semantic Search: Instead of searching for exact keywords, you compare embedding vectors. The query "How do I cancel my subscription?" also finds documents about "end membership" or "terminate plan".
- Clustering: Automatically grouping thousands of customer feedback entries by topic -- without predefined categories.
- Anomaly Detection: Finding texts that are semantically different from all others.
- RAG (Retrieval-Augmented Generation): Finding relevant documents for an LLM -- the foundation for the next section.
Dimensions in practice
| Model | Dimensions | Typical Use |
|---|---|---|
| text-embedding-3-small | 1,536 | Cost-effective, general search |
| text-embedding-3-large | 3,072 | Highest quality, complex semantics |
| Cohere embed-v4 | 1,536 | Multilingual, compact |
| Voyage 3 | 1,024 | General-purpose, multilingual (voyage-code-3 is for code) |
Reflect
A company wants to make its internal knowledge base searchable. Employees should be able to ask questions in natural language. Which technology is the best foundation?