You can be told two contradictory things in the same afternoon: agents killed vector search, or every serious team needs a vector database. Both camps are confident, both are partly right, and neither gives you a crossover point. That crossover has a name: the scale threshold, where corpus size, latency budget and concurrency decide whether vector search earns its keep. It sits inside a broader vector search going native shift, where nearest-neighbour search became a column type in databases you already run.
What is the best vector search approach for a team under 10 million vectors?
Under a few million vectors, pgvector is good enough and you rarely need a dedicated vector database. Most teams sit below that line, and the “do I need a new database” question mostly answers itself: no.
The ceiling is higher than most people think. A single tuned HNSW index holds roughly 50 to 100 million vectors on one machine before sharding enters the picture, well past where most teams actually live. The deciding question is memory fit. The sizing calculation is simple: vectors times dimensions times bytes per element. One million vectors at 1024 dimensions and four bytes each is about 4 GB raw, before HNSW graph overhead adds another two to three times that. That arithmetic sets your hardware tier; vendor marketing does not.
Quantisation moves the threshold. Int8 scalar quantisation cuts the footprint by about 75 percent, and binary quantisation collapses each dimension to a single bit, a 32x reduction that took one 100-million-vector index from 300 GB to under 10 GB. If the index fits, HNSW gives the best recall and latency. If it doesn’t, IVF and DiskANN keep centroids in memory and vectors on disk.
The stores then slot onto the curve. Chroma is a lightweight option for small teams prototyping under 10 million vectors. Pinecone is a managed reference point with a $50 a month floor. Milvus and Zilliz Cloud sit at the upper end, built for billions of vectors. The question for your team is fit and cost, covered in what a vector database actually costs, and part of the broader native vector search shift.
Fit and cost assume you need an index at all. The sharper question is whether files and an agent can do the job without one.
Files and agents versus vector search at scale: which wins?
The headline was that filesystem tools beat vector search, and it describes a small scale, not the whole curve. In the LlamaIndex benchmark, a filesystem-exploring agent scored correctness 8.4 versus 6.4 and relevance 9.6 versus 8.0 against vector RAG. It also took 11.17 seconds versus RAG’s 7.36. Then the corpus grew and the result inverted: RAG pulled ahead on speed and slightly on correctness, because a pre-built index holds its latency while file search degrades.
The counter-current is real. The argument that keeps showing up, including on HackerNoon, is that teams reach for a vector database before they need one. Qwen‘s zvec-grep (zg) makes that concrete: a local-first tool that unifies ripgrep, BM25 and vector search behind one CLI, with an on-device index and no standalone service to deploy. For a small self-hosted team, that is a lot of retrieval power for zero infrastructure.
The nuance is what the headline misses. Filesystem recall is not tunable; grep returns all matches with no intrinsic ranking. Vector recall is tunable, through knobs like HNSW’s ef_search, and that matters when you have to weigh recall against latency. Which is why production systems don’t choose: they run hybrid search, fusing lexical and vector ranked lists, with reciprocal rank fusion. It is the consensus default across the field.
Hybrid search handles static corpora. The workload that truly forces vectors back is agent memory.
Why do agent workloads still need vector search at scale?
Agent memory is a write-heavy stream of conversation summaries, learned facts and past decisions, stored as embeddings and appended, consolidated and rewritten across sessions, rather than a static corpus you index once and query forever. That write path is what separates it from RAG’s read-heavy retrieval, and it grows without bound. Filesystem exploration cannot serve that pattern.
Retrieval is the bottleneck. Roughly 41 percent of end-to-end RAG latency sits in the retrieval step, and every turn pays it again. There is a reliability problem too. Agent reliability fell from about 60 percent on a single run to roughly 25 percent across eight runs, a drop a deterministic first-stage retriever fixes. Agent accuracy problems are almost always retrieval problems.
That is the case for a tunable vector shortlist. Filesystem recall has no knobs and cannot hold sub-200ms retrieval across millions of items under concurrency. A vector index narrows millions of candidates to a shortlist in milliseconds, then the agent reasons over it. At the upper end of the curve, Milvus and Zilliz Cloud earn their keep, and the recurring write cost of unbounded memory is why what a vector database actually costs matters as much as recall.
So when is vector search worth it?
When your workload crosses the scale threshold, and not before. Below a few million vectors, pgvector and native search cover you, and a single tuned HNSW index reaches 50 to 100 million vectors on one machine before sharding.
The production default is hybrid. “Files won” and “vectors won” are both true, just at different points on the same curve. The teams that get this right are not the ones with the strongest opinion; they counted their documents, measured their latency budget and modelled their token cost before picking a side.
The decision is a sum: does the index fit in memory, and what does the store cost to run? Answer those two questions and vector search in general-purpose databases stops being a debate and becomes a tool you reach for when you need it.
Frequently Asked Questions
Do I need a vector database if I already run PostgreSQL?
No, not by default. If your corpus sits under a few million vectors, pgvector runs inside the Postgres you already operate, so you skip a second system to secure, back up and monitor. Reach for a dedicated store only when the index stops fitting in memory or concurrency pushes latency past what one tuned HNSW index can hold.
Is pgvector good enough for production, or is it just for prototypes?
pgvector is production-grade well below its ceiling, not a toy. A single tuned HNSW index holds roughly 50 to 100 million vectors on one machine before sharding, which covers most teams outright. The limit is memory and write throughput, not credibility: once the index spills or concurrent writes bite, a dedicated store such as Milvus earns its cost.
What is hybrid search, and when should I use it?
Hybrid search runs a lexical query and a vector query over the same corpus, then fuses the two ranked lists, usually with reciprocal rank fusion. It matters because neither half dominates: keywords catch exact identifiers and rare terms, vectors catch paraphrase and meaning. For most production systems hybrid is the sensible default, and it is what zvec-grep bundles into one CLI.
Is grep really better than embeddings now?
No, that headline is a small-scale snapshot, not a verdict. In the LlamaIndex benchmark the filesystem agent won on quality at tiny scale (correctness 8.4 vs 6.4) but lost on latency (11.17s vs 7.36s), and the result inverted as the corpus grew. File recall is not tunable; vector recall is.
Do I need embeddings at all if keyword search already works?
Often you do not, and leaning on BM25 or full-text search first is a reasonable starting point. Embeddings earn their place when queries are conversational, when wording varies from your documents, or when the corpus is too large to scan. If exact terms already surface the right results, adding vectors is cost without benefit.
How much does a vector database actually cost to run?
Managed stores such as Pinecone bill for storage and query volume, so the minimums are low but the bill scales with writes and index size. The real cost driver is usually re-embedding and reindexing as your data changes, not the queries. Compare that recurring figure against the operating cost of a pgvector instance you already pay for.
How does embedding dimension size affect memory and cost?
Memory scales linearly with dimension, so it matters more than most teams expect. The sizing calculation is vectors times dimensions times bytes per element: one million 1024-dimension vectors at four bytes each is about 4 GB raw, before HNSW graph overhead. Dropping to 512 dimensions, or applying int8 quantisation, roughly halves that footprint.
What retrieval latency should I target for a production agent?
Aim for sub-200ms retrieval, because the retrieval step carries roughly 41% of end-to-end RAG latency and every turn pays it again. Filesystem exploration cannot hold that target across millions of items under concurrency. A pre-built vector index keeps latency flat as the corpus grows, which is what makes the number sustainable.
Can a filesystem agent serve many concurrent users at scale?
No, not reliably. Grep, glob and read are fine for a single agent exploring a bounded repository, but they are neither tunable nor viable for sub-200ms retrieval across millions of items when many users query at once. Vector search turns that unbounded scan into a deterministic shortlist, which is what holds up under concurrency.
Do I have to re-embed all my documents when I switch embedding models?
Yes, in practice. Vectors from two models live in different spaces, so mixing them quietly degrades recall rather than failing loudly. Switching models means re-embedding the corpus and rebuilding the index, which is why the choice is worth getting right early and why write-heavy agent memory makes re-embedding a recurring cost.