Skip to main content
Favicon of pgvector

pgvector

What is pgvector?

Pgvector is an open-source Postgres extension for developers who want vector similarity search inside an existing database. It supports exact and approximate nearest-neighbor search, multiple vector types, and distance functions including L2, inner product, and cosine distance, plus quantization for tighter storage and faster retrieval. Teams use it with GitHub APIs and webhooks, and customers include Shopify and Spotify. It is software you run inside your own Postgres, and the project publishes no paid plans.

Last verifiedHow we evaluate

Screenshot of pgvector website

At a glance

Best for
Pgvector is best for developers who want vector search inside Postgres without adding a separate database.
Pricing
Free and open source — a self-hosted Postgres extension, no paid plans
API
Yes — GitHub advertises APIs and webhooks for getting data and events, and for automating workflows within GitHub.

What it actually is

pgvector is a Postgres extension (C, built with the standard Postgres extension build system) that adds a vector column type plus supporting index types (HNSW, IVFFlat) and distance operators (L2, inner product, cosine, L1, and — for binary vectors — Hamming and Jaccard). It's maintained by Andrew Kane (creator of many similar 'ann' extensions across languages) as a single-purpose extension, not a company or hosted product. Everything runs inside your existing Postgres instance: same backups, same replication, same transactions, same JOINs against your relational data.

Technical capabilities, verified against the current release (v0.8.6)

Vector storage: 4 * dimensions + 8 bytes per vector, up to 2,000 dimensions for indexed vector columns, 4,000 for halfvec (half-precision), and 64,000 for binary-quantized bit vectors; raw (unindexed) vector columns can go up to 16,000 dimensions. Supports exact search out of the box and two approximate index types (HNSW and IVFFlat), plus binary quantization, sparse vectors, iterative index scans (to recover recall lost to post-filtering), and partial/partitioned indexing for multi-tenant isolation. Requires Postgres 13+; official Docker images are published for Postgres 13 through 18. Client libraries exist for roughly 30 languages, maintained in separate pgvector/pgvector-<lang> repos.

Maturity and maintenance

22,700+ GitHub stars, created April 2021, actively maintained: the latest tag is v0.8.6 (2026-07-29), with bug-fix PRs merged as recently as 2026-08-04 (a race condition between HNSW insert and vacuum), and commits as recent as 2026-08-15. Release cadence has been roughly monthly through 2026 (0.8.2 through 0.8.6, plus an 0.8.7 unreleased branch already in the changelog). No security advisories are filed against the repo, and no entries appear in the OSV.dev vulnerability database. One caveat: contribution is heavily concentrated — Andrew Kane accounts for 1,915 of the commits in the top-10 contributor list against 18 for the next-most-active contributor — so this is a single-maintainer project in practice, even though it's widely deployed and gets outside code review from contributors at AWS and other Postgres vendors on individual PRs.

License

The PostgreSQL License — a permissive, OSI-approved license structurally similar to MIT/BSD (no copyleft, no field-of-use restriction). GitHub's own license detector reports it as "Other / NOASSERTION" because the license text is PostgreSQL's own non-templated wording rather than a byte-for-byte match to a known template, not because the terms are unusual or restrictive — the actual LICENSE file text is standard permissive language used by PostgreSQL itself since the 1990s.

Where it stops being the right tool

Independent community reporting (via Qdrant's own comparison post, so read as an interested party) and pgvector's own README agree on the shape of the ceiling: index build time and recall both degrade as the HNSW graph grows past what fits in maintenance_work_mem, and the graph needs to fit in memory for good query performance. Using the README's own storage formula, 10 million vectors at a common embedding size (1536 dimensions, e.g. OpenAI's text-embedding-3-small) works out to roughly 61.5 GB of raw vector storage before HNSW graph overhead — more RAM than most single Postgres instances carry by default. Timescale's Postgres extension pgvectorscale (a separate, third-party project, not part of pgvector itself) adds a disk-based DiskANN index aimed specifically at this gap; Timescale's own benchmark (their product, their numbers) claims 16x higher throughput and 28x lower p95 latency than Pinecone's storage-optimized tier at 50M vectors — useful context on the ceiling, but a vendor benchmark of their own extension, not independent.

Where it's actually deployed

pgvector ships as an installable/preinstalled extension on Amazon RDS and Aurora PostgreSQL, Neon, Supabase, Google Cloud SQL for PostgreSQL, and Azure Database for PostgreSQL, among others — confirmed directly on AWS's and Neon's own docs. Its own README declines to maintain a hosted-provider list and instead points to a community-tracked GitHub issue, which is a reasonable choice for a project this widely adopted but means you should verify support and version on your specific provider before assuming it's current.

Frequently asked questions

What is pgvector?

Pgvector is an open-source Postgres extension for developers who want vector similarity search inside an existing database. It supports exact and approximate nearest-neighbor search, multiple vector types, and distance functions including L2, inner product, and cosine distance, plus quantization for tighter storage and faster retrieval. It is software you run inside your own Postgres, and the project publishes no paid plans.

How much does pgvector cost? Is it free?

Pgvector is free: it is an open-source Postgres extension you install in a database you already run, and the project publishes no paid plans and no trial.

What is pgvector used for? Who is it for?

Pgvector is used for exact and approximate nearest neighbor search, single-precision, half-precision, binary, and sparse vectors, and L2 distance. It's built for Backend engineers, Data platform teams, and Product engineers building recommendation or similarity features from embeddings.

Does pgvector have an API and what does it integrate with?

GitHub advertises APIs and webhooks for getting data and events, and for automating workflows within GitHub. It integrates with Docker, Homebrew, PGXN, APT, Yum, and 25 more.

Editor's read

Check whether your embedding workload needs the Enterprise tier's data residency, SCIM provisioning, or SAML single sign-on. Those controls are only listed on Enterprise, so teams with compliance or identity requirements should verify the upgrade path before standardizing on lower tiers.

Share:

Sponsored
Favicon

 

  
 

Explore other Agent Tools & Integrations

Favicon

 

  
  
Favicon

 

  
  
Favicon