Your application already runs on DynamoDB, and now the product team wants semantic search or retrieval-augmented generation. The reflexive answer is a dedicated vector database plus a synchronisation job, which brings overhead, stale results, and another database to secure.
The decision is where the vectors live, and this article explains why that matters.
What is native vector search in a general-purpose database?
Native vector search stores vector embeddings as an attribute of items in the database your application already runs on, indexing and querying them in place. One table serves as both operational datastore and vector store. The shift is placement: the index moves into the database you already run.
Embeddings are lists of floats that encode the semantic meaning of text, images or audio.
The alternative is a dedicated vector database, a separate system such as Pinecone, Weaviate, Milvus or Qdrant that needs a sync pipeline. A standalone vector store adds synchronisation overhead, operational complexity and query latency. Native search skips that, because the vectors live where the data already is.
The ceiling is 4,096 dimensions per vector. The motivating workloads are retrieval-augmented generation and semantic search, and the thing it fixes is data sprawl: no second copy of your dataset. DynamoDB is the running example throughout, and this is part of vector search going native.
What is the copy-and-synchronize pipeline that native vector search removes?
A copy-and-synchronise pipeline is the ETL or sync job that streams vectors from your operational database into a separate vector store and keeps the two in step. It costs operational overhead, risks stale results, forces you to run two databases, and adds extra failure modes. Native vector search removes it entirely.
Before the DynamoDB update, teams typically copied data into OpenSearch or another vector database using DynamoDB Streams or custom pipelines, running two data layers plus the backfill, retry and sync work. One analyst told InfoWorld that a second database meant paying at least $700 a month regardless of usage, and that a synced-five-minutes-ago copy can mean an agent acting on stale data.
AWS’s earlier zero-ETL path went DynamoDB to S3 to OpenSearch. It sounds hands-off, but you still run a separate search cluster, as ScyllaDB points out.
Native search removes the pipeline: an embedding goes into the table as a plain list of numbers, and a search returns ranked matches. We compare the two in native versus a dedicated vector database and work through the cost of the pipeline native vector search removes.
That pipeline is a symptom of a larger force: data gravity.
What is “data gravity”, and why does it tax AI data architectures?
Data gravity is the tendency for data to accumulate weight and cost as it grows, making it harder and more expensive to move or duplicate. It taxes AI architectures because each extra copy of a dataset multiplies cost, sync work and staleness.
The term describes the invisible pull that draws applications, services and more data to large datasets. Data does not move nearly as easily as compute, and every system that makes you copy data first levies a gravity tax, as VAST Data argues.
The failure mode is three copies of the same dataset: operational, vector and object storage. Your data engineers spend more time on pipelines than AI features, and the copies, siloed, do not agree, as Zilliz’s lakebase analysis describes.
That is where the vector lake and vector lakebase split matters. A vector lake is raw embeddings parked cheaply in object storage such as S3 Vectors. A vector lakebase is the queryable layer over that storage, so you search parked vectors without copying them into a live database. Native search is one response to data gravity: query the vectors where they sit. When vector search earns its place at all is the companion question.
Why does storing vectors alongside operational data reduce latency and cost?
Storing vectors beside operational data lets one similarity query return item data with the matches rather than just IDs, removing the second round trip a dedicated store requires. Approximate nearest neighbour indexes trade a little exactness for large speed and cost wins. One database means lower infrastructure, sync and operational cost.
In the two-database pattern you find the nearest neighbours, then fetch the records separately. With colocation, the search returns item data alongside the results, as AWS’s build guide explains.
The index beneath this is usually HNSW, though IVF and DiskANN are common too. That trade-off is why recall is a first-class metric, as Moor Insights & Strategy notes.
Embedding model churn is the part teams miss. Because the index fixes dimensions and distance function, swapping embedding models changes the coordinate system, and the index returns neighbours from the wrong neighbourhood, as Tian Pan explains. A changed model forces a full rebuild, coupling your model choice to the store.
Embeddings are generated upstream, typically through Amazon Bedrock with a model like Titan Text Embeddings V2. AWS publishes single-digit-millisecond latency and 99%+ recall, figures worth treating as vendor claims rather than independent benchmarks.
AWS read the same trade-off and shipped it as a first-class feature.
Why did AWS make vector search generally available inside DynamoDB?
AWS made DynamoDB vector search generally available in August 2026, the seventh AWS service with vector capability after S3 Vectors launched in December 2025. It answered DynamoDB applications that wanted semantic search without a second database, collapsing a common two-database architecture into one operational data layer.
AWS leaned into horizontal scaling, working with existing tables, and pay-per-use billing. It shipped straight to general availability in all commercial regions and GovCloud with no public preview, so existing teams could enable it without a migration. Vector indexes run on on-demand capacity mode, and adding an index to an existing table backfills it at no extra charge.
Go to the primary sources: the AWS What’s New announcement and the DynamoDB vector search documentation carry the current limits, distance functions and billing detail.
Futurum’s Brad Shimmin compared the prior state to plumbing held together with solder and duct tape in his analysis. The move matches vector capability rivals like MongoDB and Couchbase had already shipped, and sits within the wider shift to native vector search.
Wrapping it all up
Native vector search is a placement shift. The similarity index moves into the operational database your app already runs on, and the question to ask is where the vectors live.
The copy-and-synchronise pipeline is the signature of vectors living in the wrong place, and data gravity is the tax that mistake keeps charging as your data grows.
AWS’s August 2026 general availability is the industry reading the same signal: semantic search becoming a first-class attribute of the operational datastore. Next time you weigh vector search for your business, ask where the vectors live and what copies it creates.
Frequently Asked Questions
How do ANN indexes like HNSW actually work?
An approximate nearest neighbour (ANN) index such as HNSW organises vectors into a navigable graph, so a query hops toward the closest matches instead of comparing it against every vector in the table. It trades a little exactness for large speed and cost wins, which is why recall becomes a first-class metric. Published figures put DynamoDB at single-digit-millisecond latency with 99%+ recall, though AWS authored those numbers.
Why does changing my embedding model force a full rebuild?
Because the vector index fixes both the number of dimensions and the distance function, vectors from a different model no longer line up with the ones already indexed. Swapping from one embedding model to another produces vectors of a different shape or meaning, so every embedding must be regenerated and the index rebuilt from scratch. That coupling ties your choice of embedding model directly to the store.
What is a vector lakebase, and how is it different from a vector lake?
A vector lake is raw embeddings parked cheaply in object storage such as Amazon S3 Vectors, where idle capacity costs nothing. A vector lakebase is the queryable layer built over that storage, letting you search the parked vectors without copying them into a live database. The split matters because it keeps the bulk of your vectors cheap while still supporting on-demand search.
Where can I find the official AWS documentation for DynamoDB vector search?
AWS announced the general availability milestone in its What’s New post and documents the capability in the DynamoDB vector search pages of the AWS documentation. Both are primary sources worth reading alongside this explainer, particularly for current limits, supported distance functions and billing details. This article is a reference for the architectural shift, not a step-by-step walkthrough.
Do I need to migrate my data to use native vector search?
No. Native vector search works with the DynamoDB tables you already run, because embeddings are stored as an attribute of the items already there. There is no second copy of the dataset to build, no export step and no pipeline to maintain. You add vectors to the items you already hold, and DynamoDB backfills new indexes on existing tables at no extra charge.
Where do the embeddings I store in DynamoDB come from?
Embeddings are generated upstream, typically through Amazon Bedrock using a model such as Titan Text Embeddings V2, then written into your items as an attribute. The constraint that matters is compatibility: the output must match your index configuration, meaning the same model, dimensions and distance function, so query vectors and indexed vectors line up.
Is native vector search more expensive than running a dedicated vector database?
Usually less, because you are not paying for a second database alongside a sync pipeline. DynamoDB’s serverless, pay-per-use model charges for the capacity you consume rather than idle infrastructure, and colocation removes duplicated storage and the operational overhead of keeping two stores in step. Actual savings still depend on your workload shape, so model it against your own usage.
What happens if my dataset grows very large? Does native search still scale?
Native vector search scales with DynamoDB itself, which is horizontally scalable and fully managed, so growth does not force you to re-architect around a second system. That is the point of the placement reframe: the vectors grow exactly where the operational data already grows. Exact performance at large scale depends on index configuration and query patterns, so test against your own workload.
Does native vector search replace keyword search, or can I use both?
Vector search complements keyword search rather than replacing it. Keyword search matches literal terms and stays strong for exact strings, identifiers and precise filters. Vector search matches meaning, so it handles paraphrases and fuzzy intent that keyword search misses. Many applications get the best result from both, combining exact filters with semantic similarity against the same table.
Can native vector search handle images and audio, or only text?
Embeddings encode the semantic meaning of text, images or audio, so the technique is not limited to text. Most RAG and semantic search pipelines today use text embeddings, but the same in-place index can store vectors produced from images or audio, provided the dimensions and distance function match the index configuration. The retrieval mechanics are identical either way.
Is the 99% recall figure reliable?
Treat it with caution. The single-digit-millisecond latency and 99%+ recall figures come from AWS, which authored them, so they are vendor claims rather than independently verified benchmarks. They are plausible for a well-tuned approximate index, but your own recall and latency will depend on your data, query patterns and index settings, so measure before relying on them.
What is the difference between vector search and retrieval-augmented generation?
Vector search is the retrieval mechanism; retrieval-augmented generation is the application built on top of it. In a RAG pipeline, vector search finds the most relevant passages, and those passages are fed to a language model so it can answer using your data rather than only its training. Semantic search is the simpler cousin that surfaces the matches directly, without the generation step.