Jonah Whitcombe 8 min readA search box usually begins with words and ends with matching words. Ask for “software that tracks customer conversations,” and a conventional system may miss a document titled “CRM platform” because the strings differ. Embeddings create another possibility: compare items by their inferred meaning rather than their literal spelling.
An embedding is a numerical representation of an object. A model converts text, an image, audio, or another input into a list of numbers called a vector. Objects the model considers related tend to receive vectors that sit near one another in a mathematical space. This simple mechanism supports semantic search, recommendations, clustering, duplicate detection, and retrieval for language models.
The essential insight is not that embeddings understand meaning as people do. It is that they provide useful coordinates for patterns learned from data. Once that distinction is clear, the technology becomes easier to use—and easier to distrust appropriately.
The vocabulary: five ideas to learn first
Embedding model: the model that transforms an input into a vector. Different models produce different spaces, even when they accept the same input.
Vector: an ordered list of numbers. Its length is the vector’s dimensionality. Each dimension rarely corresponds to a neat human concept such as “formal” or “financial.” Meaning is distributed across the representation.
Similarity metric: a rule for comparing vectors. Cosine similarity compares their direction; dot product combines direction and magnitude; Euclidean distance measures straight-line separation. Use the metric recommended for your model rather than choosing one by intuition.
Nearest-neighbor search: the process of finding stored vectors closest to a query vector. Exact search compares every candidate. Approximate nearest-neighbor search uses indexes that trade a small amount of recall for much faster retrieval at scale.
Vector database: storage and retrieval infrastructure designed for vectors and their metadata. It is useful, but not always necessary. Many general-purpose databases and search engines can store vectors and run similarity queries.
The mental model: a map, not a dictionary
Imagine a map on which restaurants are positioned by characteristics rather than geography. A quiet vegetarian café may lie near a calm tea room, despite different menus. A late-night steakhouse may appear far away. The coordinates compress many characteristics into spatial relationships.
An embedding space works similarly, except it may have hundreds or thousands of dimensions rather than two. Consider three product-support messages:
- “I cannot access my account.”
- “My login keeps failing.”
- “Please change the shipping address.”
A suitable text model will often place the first two vectors closer together than either is to the third. That relationship is learned, not defined by a handcrafted synonym list.
The workflow is straightforward:
- Convert each searchable item into an embedding.
- Store the vector alongside the item and useful metadata.
- Convert an incoming query with the same model.
- Find nearby stored vectors.
- Filter, rank, or pass the resulting items to another system.
The phrase with the same model matters. Coordinates only have meaning within their own embedding space. Vectors from unrelated models are generally not directly comparable.
What embeddings make possible
| Use case | What is embedded | What proximity represents |
|---|---|---|
| Semantic search | Queries and documents | Likely relevance despite different wording |
| Recommendations | Products, content, or user histories | Related taste, behavior, or attributes |
| Clustering | Messages, documents, or records | Potential themes or natural groupings |
| Duplicate detection | Listings, tickets, or passages | Possible semantic equivalence |
| Retrieval-augmented generation | Questions and knowledge passages | Context likely to help a model answer |
These applications share a structure: turn an ambiguous comparison into a geometric lookup. The embedding does not complete the entire task. It produces candidates or signals that later stages can use.
For recommendation, proximity alone may overproduce items similar to what a person already consumed. A final ranker can add freshness, availability, diversity, or business constraints. For duplicate detection, a similarity threshold can identify candidates, but a rules engine or human reviewer may make the final judgment.
A worked example: semantic search for product notes
Suppose a small team has 500 product notes. A user searches for “customers abandoning setup.” One note says, “New accounts frequently stop during workspace configuration.” Keyword search may struggle because no query term appears in the note.
First, split the notes into independently useful passages. A passage might contain a heading and one or two paragraphs. Embed each passage and store its vector with metadata such as document identifier, date, author, and access level.
When the query arrives, embed it and request the nearest passages. The configuration note may rank highly because the model associates “abandoning setup” with “stop during workspace configuration.” The application can then show the passage, its source, and perhaps neighboring text.
Chunking changes the result. Embedding an entire long document can blur several subjects into one representation. Tiny fragments may lack enough context to be interpretable. A useful starting unit is the smallest passage that still makes sense when displayed alone. Test alternatives using real queries rather than adopting a universal chunk size.
Metadata filtering should happen deliberately. If the user may only view notes from one workspace, apply that authorization constraint during retrieval. Never retrieve forbidden material and assume the interface will hide it later.
The trade-offs hidden inside similarity
Semantic similarity is not factual relevance
A passage can sound related while providing no answer. A query about cancelling an account may retrieve a passage about suspending one because the concepts are close. Retrieval quality therefore needs evaluation at the level of the user’s task.
Thresholds are product decisions
There is no universal score that means “same meaning.” Scores vary by model, metric, content, and query type. A strict threshold may omit useful results; a permissive one may admit noise. Derive thresholds from labeled examples in your own domain.
Embedding models carry boundaries
A model may represent common English product language well yet struggle with specialist terminology, uncommon languages, source code, or very long inputs. Domain-specific models can improve one area while becoming less general elsewhere.
Vector search does not replace structured filters
If a user asks for invoices issued after a particular date, use a date field. If a product must be in stock, filter by inventory. Embeddings are suited to fuzzy semantic relationships, not exact arithmetic, permissions, or categorical truth.
Your first experiment
Begin with a narrow retrieval problem where relevance can be inspected. Internal FAQs, support tickets, or a modest document collection are better learning environments than an autonomous agent.
- Assemble a small corpus. Keep source text and metadata intact so every result remains traceable.
- Write representative queries. Include direct wording, paraphrases, ambiguous requests, and questions with no valid answer.
- Label expected results. For each query, identify passages that would genuinely help.
- Create embeddings. Use one model for both passages and queries, following its input and similarity guidance.
- Retrieve a short candidate list. Inspect rankings rather than merely checking whether one good result appears somewhere.
- Compare a baseline. Run keyword search on the same queries. Semantic retrieval may win on paraphrases while losing on identifiers, names, and exact phrases.
- Try hybrid retrieval. Combine keyword and vector signals when both lexical precision and semantic reach matter.
Record failure patterns. Are long passages dominating? Are product codes being missed? Are generic policy pages appearing for every query? Each pattern suggests a different intervention: change chunking, preserve keyword search, add metadata filters, or introduce reranking.
How to evaluate without fooling yourself
A polished demonstration proves little because its query was often chosen after seeing the data. Build an evaluation set before tuning the system. Keep some queries aside so improvements are not merely adaptations to familiar examples.
Measure what the product needs. If users only inspect the first three results, quality deep in the list has limited value. If retrieval supplies context to a language model, ask whether the returned passages contain sufficient evidence for an answer—not merely whether they share a topic.
Negative cases are equally important. Test queries for which the corpus contains no answer. A system that always returns something can create a false impression of knowledge. Your interface or downstream model needs a way to say that the available material is insufficient.
Review results by meaningful slices: language, document type, department, query length, and recency. An acceptable average can conceal a consistently weak segment.
What to ignore for now
Do not begin by debating specialized vector databases. Start with the smallest infrastructure that can run the experiment; migrate when scale, latency, filtering, or operations create a demonstrated need.
Ignore visualizations that squeeze high-dimensional vectors into a beautiful two-dimensional constellation. They can suggest patterns, but the projection distorts distance and may imply clusters that are not operationally useful.
Defer fine-tuning. First improve corpus quality, chunking, metadata, query design, keyword coverage, and evaluation. These changes are easier to diagnose and often address the actual failure.
Most importantly, ignore the temptation to call embeddings a knowledge layer. They are an indexing signal: powerful because they make approximate meaning computable, limited because approximation is not judgment. Treat them as one component in a retrieval system, surround them with exact constraints and evidence, and they become a dependable bridge between human language and software action.
This post was drafted with AI assistance and reviewed against our editorial policy before publication. Corrections are made at the source, on the page, with the date shown.
From our own rounds
Measured on The Curator, from real sessions people played on this site — not a third-party dataset.
- Rounds played here
- 167
- Questions per round
- 1.7
Rate this article
Discussion
Comments are moderated. Read our editorial policy.