Semantic search is the capability behind every modern “find anything in your data” product in 2026, from Notion AI’s workspace search to GitHub Copilot’s code lookup to the internal support bots at every mid-size SaaS company. This guide walks through how semantic search actually works, where it beats keyword search, where it falls short, and how to build a production system you will not need to rewrite in six months.
What semantic search actually does
Classical keyword search matches the words you typed against the words in your documents. If a user searches for “how to reset my password” and the document says “steps to recover a forgotten account,” keyword search misses because none of the key words overlap meaningfully. Semantic search solves this by comparing the meaning of the query to the meaning of each document.
The mechanism is embeddings. An embedding model converts a chunk of text into a fixed-length vector of numbers (typically 512 to 3072 dimensions in 2026) such that pieces of text with similar meaning end up near each other in that vector space. A vector database stores all the document vectors and, at query time, finds the ones closest to the query vector. Those are your results.
The whole thing is more useful than it sounds in a sentence because meaning is messy. Humans paraphrase, abbreviate, and use jargon. Keywords miss paraphrases. Embeddings catch them.
When semantic search is the right tool
Semantic search is the right choice when:
- Users type conversational queries (“why is my build failing?”)
- Documents use different vocabulary than the queries will use
- Your corpus is small to medium: thousands to millions of documents
- Precision at top-10 matters more than recall at top-1000
- You want to pass results to an LLM for grounded answer generation (RAG)
Keyword search is still the right choice when:
- Users search by exact identifiers (SKUs, order numbers, filenames)
- Your corpus is massive (billions of documents)
- You need traditional database features like counts, facets, and ranges
- Users will keep typing until the results match, so keyword precision is enough
Many production systems use both: semantic search for discovery, keyword search for exact-match queries, and a reranker to fuse the results.
The full semantic search pipeline
Every production semantic search system has five stages:
- Document ingestion: pull data from the source (database, SaaS API, file store), normalize it, extract text.
- Chunking: split long documents into chunks small enough that one embedding meaningfully represents one topic. 500 to 1000 characters is a common target.
- Embedding: run each chunk through an embedding model. Store the resulting vector plus metadata (source document, chunk index, author, timestamp).
- Indexing: upsert the vectors into a vector database (Pinecone, Weaviate, Qdrant). The database builds an approximate nearest neighbor (ANN) index for fast query-time lookup.
- Querying: embed the user query with the same model used for documents, run a nearest-neighbor search in the database, return the top-k results (usually 5 to 20). Optionally rerank with a specialized reranker model.
Minimal working example
Here is a complete semantic search system in roughly 40 lines of Python. The example uses OpenAI embeddings and Qdrant running locally in Docker, but you can swap either component.
from openai import OpenAI
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams, PointStruct
openai = OpenAI()
qdrant = QdrantClient(url="http://localhost:6333")
qdrant.create_collection(
collection_name="docs",
vectors_config=VectorParams(size=1536, distance=Distance.COSINE),
)
def embed(text: str) -> list[float]:
return openai.embeddings.create(
model="text-embedding-3-small",
input=text,
).data[0].embedding
documents = [
"How to reset your password: visit account settings and click 'change password'.",
"Refund policy: full refunds within 30 days of purchase.",
"To cancel your subscription, go to billing and select 'cancel plan'.",
"Supported browsers: Chrome 110+, Firefox 115+, Safari 16+.",
]
qdrant.upsert(
collection_name="docs",
points=[
PointStruct(id=i, vector=embed(doc), payload={"text": doc})
for i, doc in enumerate(documents)
],
)
query = "how do I change my password?"
results = qdrant.search(
collection_name="docs",
query_vector=embed(query),
limit=2,
)
for r in results:
print(f"Score {r.score:.3f}: {r.payload['text']}")Run it and watch the password reset document come back first even though the user typed “change” and the document says “reset.” That is semantic search doing its job.
Chunking strategy matters more than people think
Chunking is the stage most teams underinvest in and the one that affects quality the most. The failure modes:
- Chunks too large: one vector represents multiple unrelated topics, so retrieval is noisy.
- Chunks too small: context gets split across chunks, so neither has enough signal to match a query that needs both halves.
- Chunks that break on arbitrary characters: a sentence or table row split in half confuses the embedding.
Three strategies worth knowing:
- Fixed-size with overlap: 800 characters per chunk, 100 character overlap between consecutive chunks. Simple, good default for prose.
- Semantic chunking: use another embedding pass to detect sentence boundaries where topic shifts. More compute at indexing time, better retrieval downstream.
- Structural chunking: respect document structure (one chunk per Markdown heading, one chunk per PDF section, one chunk per code function). Best when your source has reliable structure.
For most RAG applications in 2026, start with fixed-size overlap chunking and only move to semantic or structural chunking if evaluation shows quality holding you back.
Hybrid search and rerankers
Dense semantic search is powerful but blind to exact identifiers. If a user searches for order “ORD-2024-8871” the embedding model has no idea that string is a unique key. Hybrid search solves this: combine dense semantic retrieval with sparse keyword retrieval (BM25), fuse the scores, and get the best of both. Pinecone, Weaviate, and Qdrant all support hybrid search out of the box in 2026.
Rerankers take the top-k results from retrieval (usually 20 to 50) and reorder them with a slower but more precise model. Cohere Rerank 3 and Voyage rerank-2 are the two production-grade options. For most RAG apps, retrieve 20, rerank to 5, feed those 5 to the LLM. The quality improvement is noticeable and the extra latency (around 200 ms) is usually acceptable.
Evaluation and the thing you must not skip
Semantic search quality is not something you can eyeball. You will convince yourself the results look good and ship something that gives wrong answers 30% of the time. Build an evaluation set:
- Collect 50 to 100 real queries users have asked (or will ask) your system.
- For each, label the one or two documents that contain the correct answer.
- Run your pipeline and compute precision at 5: what fraction of queries had the correct document in the top-5 results.
- Rerun this evaluation every time you change chunk size, embedding model, retriever, or reranker. Chase the number up over time.
The team that keeps this discipline reliably ships better semantic search than the team that does not. Nothing else matters as much.
Frequently Asked Questions
Is semantic search the same as RAG?
No, but they are related. Semantic search returns the most relevant document chunks for a query. RAG (retrieval-augmented generation) uses semantic search to find relevant chunks and then feeds them to an LLM to generate an answer. RAG uses semantic search as a component.
Can I run semantic search without a vector database?
Yes, for small datasets. FAISS, sklearn, or in-memory numpy arrays all work for up to a few hundred thousand vectors on a single machine. Vector databases become useful when you need concurrent writes, persistence across restarts, metadata filtering, or horizontal scaling.
How do I handle updates to documents?
When a document changes, you re-embed it and upsert the new vectors to the index, then delete the old ones. Store a hash of each source document so your pipeline knows which ones have changed since the last run. Full rebuilds are also acceptable for smaller corpora.
What is the difference between semantic search and vector search?
They are the same thing in practice. “Vector search” names the mechanism (finding nearest neighbors in a vector space). “Semantic search” names the goal (finding results by meaning). Both terms are used interchangeably in 2026 product and marketing content.
How large can my semantic search index get?
Modern vector databases handle hundreds of millions of vectors with sub-100ms query latency. Pinecone, Weaviate, and Qdrant have all shipped production deployments at the billion-vector scale. For most products you will not approach these limits for years.
Should I use OpenAI embeddings or an open-source model?
Depends on your constraints. OpenAI embeddings are cheaper per million tokens than running your own GPU, so if you only have a few million tokens to embed, OpenAI wins on cost. Open-source models (nomic-embed, bge-large, jina-embeddings) let you self-host for data residency reasons and amortize infrastructure cost at huge scale. The crossover point is roughly 100 million tokens per month.
