Approximate nearest neighbour search has been a solved problem for a while. HNSW, IVF, DiskANN: the maths hasn’t changed much in years. What shifted across 2025 and 2026 is where the index lives. It moved out of the standalone vector database and into the database you already run, sitting beside the operational data it describes.
In August 2026, AWS made native vector search generally available inside DynamoDB. Embeddings can now live in the same table as your operational data, at up to 4,096 dimensions, with single-digit-millisecond latency and 99%+ recall, and with no second store to synchronise. It was the seventh AWS service to gain vector capability, after S3 Vectors in December 2025.
“Vectors beside operational data” is becoming the default, though the default is not right for every workload. This guide maps the shift, the trade-offs, the economics and the counter-current, so you can judge whether it matters to you.
In This Series
- What Native Vector Search Is and Why AWS Put It in DynamoDB: the mechanism and the milestone.
- When Native Vector Search Beats a Dedicated Vector Database: the three-tier comparison.
- The Real Cost of Native Vector Search Versus a Dedicated Vector Database: the total cost of ownership.
- When Vector Search Is Worth It and When Files and Agents Are Enough: the scale threshold and the counter-current.
What is native vector search, and why did it land inside the database?
Native vector search is an approximate nearest neighbour (ANN) index embedded in a general-purpose database, so embeddings live in the same table and transaction as your operational data. One table serves as both your operational datastore and your vector store. It landed inside the database because the index now sits beside the data it describes: ANN retrieval is mature, and the workloads that want it — RAG and semantic search — already run there.
The shift is architectural. It retires the bolt-on second database and the copy-and-synchronise pipeline that streamed vectors out to it. That pipeline produced data sprawl: duplicated copies of the same dataset, plus the risk of an index lagging behind the record it describes.
It landed now because the demand was already there. Applications already on DynamoDB wanted semantic search without a second system, and AWS leaned into serverless design that works with existing tables at pay-per-use pricing. Underneath sits data gravity, the pull that makes copying data progressively more expensive, so the index coming to the data is the cheaper direction.
For the full mechanism, see what native vector search actually is.
When does native vector search beat a dedicated vector database?
It is a three-tier choice. A native index (DynamoDB) keeps vectors in your existing table. A vector-capable extension (pgvector on PostgreSQL) bolts the capability onto a database you run. A standalone store (Milvus/Zilliz, Qdrant, Weaviate, Chroma) is purpose-built. Oracle 26ai’s native VECTOR type is the oldest tier, a SQL-native type that predates the 2026 wave. Choose by whether your operational data is already in place, whether the index fits in memory, and how you weigh recall against latency.
The comparisons people actually search for are DynamoDB versus pgvector, a dedicated store versus an extension, and a vector lakebase versus a standalone store. The question most people are asking is “is my existing database good enough?”, which reframes the choice as a question of fit.
Regardless of tier, every architecture faces the same dial: recall against latency. You fix the index configuration once it is set, check the memory fit against a rough sizing calculation, and tune ef_search or nprobe against quantization. Independent, apples-to-apples benchmarks are scarce, so VectorDBBench is a reference point to weight sceptically. That is the question when native beats dedicated answers.
What do the different approaches actually cost to run?
Cost flips with workload shape rather than dataset size. Object-storage tiers such as S3 Vectors cost nothing while idle, whereas managed stores carry minimums. Pinecone’s Standard tier starts around $50 a month plus roughly $16 per million read units, and Turbopuffer still yields sub-$10 bills at modest scale. Native options ride a database you already pay for, metered per vector write and search. Model the total cost of ownership: storage, writes, searches, idle minimums and engineering time.
The economics come down to idle versus steady state: object storage scales to zero, while provisioned stores charge for capacity you don’t touch. Build (self-host pgvector or Milvus), buy (managed Pinecone, Turbopuffer, Weaviate), or go native (the database you already pay for). Non-price factors can override the cost: lock-in, reversibility, and on-premises or regulated cases where a managed store is off the table.
Native index usage is metered per use. Vector writes consume write capacity, searches are metered separately, and both show up in consumed-capacity reporting. Add the line item that doesn’t appear on a vector-store invoice: the token and compute spend of alternative retrieval. The full walk-through lives in what each approach costs to run.
At what scale does vector search earn its place — and do files and agents change that?
Below a few million vectors, pgvector or a native index is generally good enough. A single tuned HNSW index holds roughly 50 to 100 million vectors on one machine before sharding enters. The counter-current: the LlamaIndex filesystem-versus-vector benchmark showed a file-exploring agent winning on quality at small scale, then inverting as the corpus grew. Filesystem recall is fixed, while vector recall is tunable. Judge by fit and cost.
The better framing of “do I need a vector database?” is threshold-finding. Count your documents and chunks, measure your latency budget, and work out whether the index fits in memory. A rough memory estimate is vector count times dimensions times bytes per element, adjusted by quantization.
On the demand side, filesystem-plus-agent retrieval and open-source local-first tooling such as zvec-grep (ripgrep plus BM25 plus vectors) argue that many teams reach for a vector database before they need one. Agent memory is write-heavy and grows without bound, retrieval is the bottleneck in RAG pipelines, and a deterministic tunable shortlist raises the reliability floor. That is at what scale vector search earns its place.
Resource Hub: Vector Search Deep Dives
Understand the Shift
- What Native Vector Search Is and Why AWS Put It in DynamoDB: the conceptual grounding, for when you want the shift explained first.
Decide the Approach
- When Native Vector Search Beats a Dedicated Vector Database: the three-tier comparison and the recall-versus-latency criteria.
Weigh the Economics and the Counter-Current
- The Real Cost of Native Vector Search Versus a Dedicated Vector Database: total cost of ownership and metering, for when the decision is on budget.
- When Vector Search Is Worth It and When Files and Agents Are Enough: the scale threshold and the files-and-agents calculus.
Where to start
You don’t need to read all four in order — pick the question you’re trying to answer. Start with the explainer on what native vector search is for the conceptual grounding, the native-versus-dedicated comparison for choosing an architecture.
When the decision rests on budget, or you suspect you’re over-engineering, go to the total-cost breakdown or the scale-threshold analysis.
Frequently Asked Questions
What is data gravity, and why does it push vectors into the database?
Data gravity is the pull that makes a dataset harder and more expensive to move as it grows, because everything that touches it wants to sit nearby. Rather than copy vectors to a second store and keep the two in sync, you run the index where the operational data already lives. That placement argument, not a new algorithm, is what drove the 2026 shift.
Is vector search the same thing as semantic search?
No, though they are often conflated. Vector search is the retrieval mechanism: an approximate nearest neighbour index over embeddings. Semantic search is the outcome it enables, matching on meaning rather than keywords. You can build semantic search on vectors, and vectors are not the only route to it. The distinction matters when vendors advertise meaning while shipping an index.
What does “approximate” in approximate nearest neighbour actually mean?
Approximate means the index returns a near-perfect match set rather than the guaranteed closest vectors. It trades a sliver of recall for a large speed gain, which is why DynamoDB reports 99%+ recall rather than 100%. You tune that trade with knobs such as HNSW’s ef_search and IVF’s nprobe. Higher recall costs latency; lower recall buys speed.
Now that 4,096 dimensions is possible, should I use the biggest embedding available?
No. Bigger embeddings cost more to store, write and search, and past a point they add little accuracy for a given task. Match the dimension count to your retrieval quality target and your memory budget, not to the largest number a service supports. Many workloads perform perfectly well at 768 or 1,536 dimensions.
What is quantization, and should I turn it on?
Quantization compresses each vector to fewer bytes per element, shrinking the memory footprint and often the cost. It buys headroom when your index would otherwise spill out of RAM, at the price of some recall. Turn it on when memory is the binding constraint and you can tolerate a small accuracy loss, then measure recall against your own queries before committing.
How do I work out whether my vector index will fit in memory?
Multiply your vector count by the dimensions by the bytes per element, then adjust for quantization. One million 1,536-dimension float32 vectors, for example, need roughly six gigabytes before overhead. Compare that with the RAM you can dedicate and remember HNSW holds about 50 to 100 million vectors on one machine. If the sum exceeds your memory, quantize, shard, or reconsider a standalone store.
Can I filter on my normal database attributes and run vector search in the same query?
Yes, and that is the quiet advantage of native search. Because vectors sit beside operational data in one table, you can constrain a query by attributes such as tenant, status or date and rank the survivors by similarity, without syncing a separate store. That hybrid filtering keeps results consistent with your current records rather than a stale replica.
Do I still need BM25 or a re-ranker if I have vector search?
Sometimes. Vectors capture meaning but can miss exact tokens such as part numbers, names or rare terms that keyword search nails. Hybrid retrieval pairs both and often beats either alone, and a re-ranker sharpens the shortlist before it reaches the model. Treat vectors as one strong signal, not a replacement for the rest of your retrieval stack.
What happens to my vector index when the source record is updated or deleted?
If you update the record, you rewrite its embedding in the same write, and a delete removes both together. That is the consistency the copy-and-synchronise pipeline struggled to guarantee, where a separate index could lag behind the record it described. Native placement puts the record and its vector in one transaction, so they cannot drift apart.
Do I have to re-embed everything if I switch embedding models?
Yes. Embeddings from one model are not comparable with another, so a model change means re-embedding the whole corpus and rebuilding the index. That is a genuine switching cost, and it applies to native and standalone stores alike. Choose an embedding model deliberately, because changing it later is a full reindex, not a config tweak.
What is a vector lakebase, and how is it different from a standalone store?
A vector lakebase keeps embeddings in object storage, such as S3 Vectors, and queries them from there rather than from a provisioned cluster. It scales to zero when idle, so it suits spiky or long-tail workloads. A standalone store runs dedicated infrastructure you pay for continuously. The difference is the same idle-versus-steady-state split that separates the cost tiers.
Can I isolate one tenant’s vectors from another’s in a native index?
Access control follows the database that hosts the index, so tenant isolation depends on how you partition and permission the table rather than on the vector capability itself. That is an advantage when you already trust the database’s security model. If you need hard isolation between many tenants, a dedicated store or separate tables may be simpler to reason about.