Zum Inhalt springen

Embeddings in Detail

Knowledge

In the Beginner course, you learned that embeddings turn words into numbers. Now let's look at how this works mathematically and why it's so powerful.

An embedding is a vector -- an ordered list of numbers. Modern embedding models like text-embedding-3-large from OpenAI produce vectors with 3,072 dimensions. That means every word, sentence, or document is described by 3,072 numbers. These numbers encode semantic properties -- dimensions of meaning that the model learned during training.

iDimensions are not categories

A single dimension in an embedding doesn't correspond to a human-readable concept like "animal" or "color". They are abstract, mathematically learned features. Only the interplay of all dimensions together produces meaning.

Understand

Vector spaces and semantic proximity

Imagine a three-dimensional space -- like a room with length, width, and height. Every word has a position in it. Words with similar meanings are close together. Now expand that room to 3,072 dimensions. The principle stays the same: Semantically similar concepts have similar vectors.

Animals
Colors
Programming
Food

Click on a point to see similarities — or enable vector math

Positions and similarity values are simplified. Real embedding spaces have hundreds of dimensions — this is a 3D projection.

Cosine Similarity -- The measure of similarity

How do you measure whether two vectors are "similar"? The most common method is cosine similarity. It measures the angle between two vectors:

  • 1.0 = identical direction (maximum similarity)
  • 0.0 = perpendicular (no relationship)
  • -1.0 = opposite direction (antonym)

The formula:

cosine_similarity(A, B) = (A . B) / (|A| * |B|)

The advantage over Euclidean distance: cosine similarity is independent of vector length. A long sentence and a short word can still have high similarity if they point in the same "direction".

Two embedding vectors have a cosine similarity of 0.92. What does that mean?

Apply

Where embeddings are used in practice

  • Semantic Search: Instead of searching for exact keywords, you compare embedding vectors. The query "How do I cancel my subscription?" also finds documents about "end membership" or "terminate plan".
  • Clustering: Automatically grouping thousands of customer feedback entries by topic -- without predefined categories.
  • Anomaly Detection: Finding texts that are semantically different from all others.
  • RAG (Retrieval-Augmented Generation): Finding relevant documents for an LLM -- the foundation for the next section.

Dimensions in practice

ModelDimensionsTypical Use
text-embedding-3-small1,536Cost-effective, general search
text-embedding-3-large3,072Highest quality, complex semantics
Cohere embed-v41,536Multilingual, compact
Voyage 31,024General-purpose, multilingual (voyage-code-3 is for code)

Reflect

A company wants to make its internal knowledge base searchable. Employees should be able to ask questions in natural language. Which technology is the best foundation?