An embedding is a list of numbers that represents the meaning of something: a sentence, a paragraph, an image, a product description. A model reads the content and outputs a few hundred or a few thousand numbers, and the useful property is that content with similar meaning produces similar numbers. “How do I cancel my subscription” and “ending my plan” land close together despite sharing almost no words.

If a vector database is the warehouse, embeddings are what it stores. This entry is about the representation itself: what it captures, what it misses, and why that determines how well any AI system understands your content.

The idea, without the maths

Imagine plotting every piece of text you own in a space where position encodes meaning. Articles about page speed cluster in one region, articles about branding in another, and a piece about how site speed affects brand perception sits between them. Nothing about that positioning depends on shared vocabulary. It depends on what the text is about.

An embedding is a set of coordinates in that space. The space has far more than three dimensions, typically several hundred to a few thousand, which is why it cannot be drawn, but the intuition survives: closeness means similarity of meaning.

Similarity is usually measured with cosine similarity, which compares the direction two vectors point rather than how long they are. The practical upshot is that a short paragraph and a long article on the same subject can still register as similar.

What embeddings are used for

Semantic retrieval. Finding the passages that answer a question, which is the core of retrieval-augmented generation and of any assistant that answers from your own content.

Site and internal search that understands intent rather than requiring the exact term.

Recommendations, surfacing related articles or products by similarity rather than by manually maintained tags.

Clustering and deduplication, grouping similar content, which is genuinely useful during a content audit for finding pages that overlap.

Classification, routing support tickets or tagging content by comparing against labelled examples.

Embeddings also underpin semantic search generally, which is why a page can rank for a question it never literally contains.

The part that determines quality: chunking

Content is not embedded as a whole document. It is split into chunks, and each chunk gets its own vector. This is where most retrieval quality is won or lost, and it is largely invisible to anyone who has not built one of these systems.

A chunk that cuts mid-argument loses the context that made it meaningful. A chunk that begins “This approach has three drawbacks” is useless in isolation, because nothing in it says what approach. A chunk covering four unrelated points produces a vector that averages them and sits close to nothing in particular.

Splitting on headings, keeping sections self-contained, and allowing slight overlap between chunks works far better than cutting at a fixed character count. Which leads to the point that matters if you own a website rather than build retrieval systems.

Why this matters for your content

Your pages are being chunked and embedded by systems you do not control. Every AI assistant that retrieves and cites web content does this, including ChatGPT Search and Perplexity. You have no say in their chunking strategy. You have complete say over whether your content chunks well.

Content that chunks well has sections that stand alone. A descriptive heading, then a passage that makes sense to someone who has read nothing above it. Content that chunks badly meanders, depends on three paragraphs of earlier setup, or buries its answer in the middle of a long undifferentiated block.

This is the same advice as answer-first structure and semantic HTML, arrived at from a different direction. Writing that is easy to quote is writing that gets quoted, and the mechanism is that a self-contained section becomes a clean, meaningful chunk.

Choosing a model, and the cost of changing it

Different models produce different embeddings, and vectors from different models are not comparable. All content in a collection must be embedded by the same model.

That makes model choice a commitment. Switching means re-embedding everything, which costs compute and time proportional to your corpus. Worth understanding before you build rather than after, because the temptation to upgrade to a newer model arrives regularly.

Beyond that, models differ in dimension count, maximum input length, language coverage, and whether they are tuned for retrieval, similarity or classification. For most business applications the practical criteria are language support, input length that fits your chunks, and cost per million tokens.

What embeddings cannot do

They do not know what is true. An embedding captures what text means, not whether it is correct. A confidently wrong passage embeds just as cleanly as an accurate one, which is why retrieval quality depends on the quality of what you put in.

They are poor at exact matching. Part numbers, SKUs, names and error codes are precisely where semantic similarity is least useful, because near-identical strings can differ in exactly the way that matters. This is why serious systems run hybrid search, combining keyword matching with embedding similarity. Each covers the other’s blind spot.

They inherit the biases of their training data, including uneven quality across languages and domains.

They lose nuance. Compressing a paragraph into a few hundred numbers is lossy by definition. Negation and subtle qualification are the usual casualties: a passage saying something is not recommended can sit uncomfortably close to one recommending it.

Where this fits

Embeddings sit underneath the retrieval layer of any assistant that answers from your own content, which is what our AI chatbot development work builds, and they explain why structure and clarity carry so much weight in our AI search optimization services. If you want your content to be retrieved and cited accurately by systems you do not control, book a discovery call.