
Embedding — what it means to turn sentence meaning into coordinates
Contents
“If you use embeddings, AI will automatically understand sentence meaning” is only half right. Embeddings don’t “understand” meaning; they convert meaning into numeric coordinates to distinguish what’s close or far. If you miss this difference, you’ll look in the wrong place when RAG (retrieval-augmented generation) search results look odd.
Background / why this term now
Behind RAG systems, internal document search, and chatbot “related question suggestions,” you’ll almost always find embeddings. Since you can’t just shove all your company documents into an LLM (large language model), embeddings are used to first filter for documents related to the query.
The issue is that embeddings are the most frequently misunderstood component in RAG tooling. As vector database providers like Pinecone explain, a RAG pipeline stores the knowledge base as vector embeddings and, at query time, retrieves relevant context to ground the LLM’s response (https://cyclr.com/resources/ai/understanding-vector-databases-a-deep-dive-with-pinecone). In other words, embeddings play the librarian that selects which materials to show before the LLM “thinks.” If this librarian fetches the wrong book, even a very smart LLM will produce an inaccurate answer.
Key data and current landscape
Embeddings are the result of converting data such as text or images into vectors (coordinates) made up of floating-point numbers. According to OpenAI’s official docs, embeddings are used to numerically measure relatedness between text strings and are applied to search, clustering, recommendations, anomaly detection, and classification (https://developers.openai.com/api/docs/guides/embeddings).
| Model (provider) | Vector dimensions | Notes | Source/date |
|---|---|---|---|
| text-embedding-3-small (OpenAI) | 1536 | Low cost, max input 8,192 tokens | OpenAI API docs |
| text-embedding-3-large (OpenAI) | 3072 | MTEB 64.6, no score changes since release | premai.io comparison, 2026 |
| Gemini Embedding 2 (Google) | Not disclosed | Supports 5 modalities—text, image, video, audio, PDF; top performance with MTEB multilingual task average 68.32 | Google, announced March 10, 2026 |
| Voyage 4 Large (Voyage AI) | Not disclosed | Top nDCG@10 on search benchmarks; strong for code and technical document search | premai.io comparison, 2026 |
Interpretation: Even though they’re all “embeddings,” vector dimensionality and training data differ by model, so results vary. More dimensions aren’t always more accurate, and rankings flip depending on language and document type. In practice, if you have a lot of multilingual documents, Cohere or the Qwen3-Embedding family have been reported to score higher than OpenAI models (https://www.premai.io/blog/best-embedding-models-for-rag-2026-ranked-by-mteb-score-cost-and-self-hosting/).
Deep dive
1) What exactly is an embedding?
An embedding converts a sentence into, say, a 1536- or 3072-length array of numbers. You can’t read meaning directly from this array. Instead, the model is trained so that sentences with similar meaning end up close to each other in a high-dimensional space.
Closeness between two vectors is usually measured with cosine similarity. It uses the cosine of the angle between vectors; the more similar the directions, the closer to 1, and the more unrelated, the closer to 0 (https://devocean.sk.com/blog/techBoardDetail.do?ID=165293). The vectors for “puppy” and “pet dog” are close, while “puppy” and “stock market” are far apart.
2) Common misconceptions — why they happen
The most common misconception is thinking an embedding model “reads and understands” sentences and produces answers like an LLM. In reality, an embedding model never generates answers. It only computes coordinates from input. Because conversational AI like ChatGPT and embedding models are often used within the same API flow, it’s easy to think “one AI handles everything,” but the model that retrieves and the model that generates are entirely different and do not correct each other’s mistakes.
A second misconception is that “embedding search is always more accurate than keyword search.” Embeddings excel at semantic similarity, but they can be weaker than traditional keyword search for negation (“is not”), numeric comparisons (“less than last year”), or exact matches of proper nouns. In vector space, “close” does not necessarily mean “close to the correct answer.”
3) What to actually evaluate
A commonly cited metric when choosing an embedding model is MTEB (Massive Text Embedding Benchmark). However, because it averages many tasks (search, classification, clustering, etc.), it’s better to check the specific subtask score that matches your use case (e.g., RAG retrieval). Also, higher vector dimensionality increases storage cost and search latency, so you should test with your real data to find the right trade-off among accuracy, cost, and speed.
4) When it’s useful and when it’s pointless
Embeddings are efficient when you have too many documents to fit into an LLM context window at once, or when recommending similar products or posts. Conversely, if you only have a few dozen documents and clear-cut answers, it’s often enough to feed the full documents directly to the LLM or use simple keyword search rather than building vector search from scratch. Adopting a vector database adds infrastructure cost and ongoing maintenance.
In practice — four rules for reading embeddings correctly
- Embedding models and LLMs are different models. If search results look off, first suspect the embedding/retrieval stage, then the LLM’s answer generation.
- MTEB is only an average for reference. Check detailed benchmarks that match your target language(s) and document types.
- High cosine similarity isn’t necessarily the correct answer. If negation, numeric comparisons, or exact proper-noun matches matter, consider a hybrid approach that includes keyword search.
- If your corpus is small or fits within the context window, verify whether a simpler method suffices before standing up a vector database.
References
Contents
Related posts

RAG — How to keep models up to date without fine-tuning
RAG retrieves relevant documents at query time and injects them into the prompt instead of retraining the model. It shines when knowledge changes often or citations are required, but it does not eliminate hallucinations.

RAG — how to know the latest information without fine-tuning
RAG retrieves relevant documents at query time and inserts them into the prompt instead of retraining the model. It’s often better than fine-tuning for fast-changing knowledge or citation needs, but it doesn’t eliminate hallucinations.

Fine-tuning — the misconception that you can teach your company to AI
If you mistake fine-tuning for "teaching AI your company’s entire knowledge," you’ll pay the cost and end up with a model that can’t reflect fresh info. Here’s how LoRA works and when to try RAG or prompt engineering first.

What is AI hallucination?
Hallucination is not the result of AI 'breaking'. It is closer to a structural side effect: today's training and evaluation methods award more points for a plausible guess than for saying 'I don't know'. Once you understand the mechanism, the way you deal with it changes too.