Chroma

Open-source embedding database for storing, indexing, and querying vectors and full-text/metadata search for AI apps.

Works with AI agents:llms.txt
Ask your AI about this, with this page as the source:ChatGPT ↗Claude ↗Perplexity ↗

Chroma is a database for the retrieval step of AI apps: you store documents with their embeddings and metadata, then look up the closest matches to a query, optionally filtered by metadata or matched by keyword or regex. It handles text, images and other modalities.

The model is collections of records (an ID, a document, an embedding, metadata). You can hand Chroma raw text and let a configured embedding function call a model such as OpenAI, Cohere or a local sentence-transformers model, or supply your own vectors. In Python it can run in-process, in memory or persisted to a local folder; the JavaScript/TypeScript and Rust clients talk to a Chroma server, which you start with its CLI or Docker image. LangChain and LlamaIndex integrate with it.

For a small team the useful parts are: the same API locally, self-hosted and in Chroma Cloud; a CLI that copies collections between a local server and Cloud in either direction; and full-text, regex and metadata filtering next to vector search. Chroma Cloud is serverless and runs in AWS us-east-1 or GCP europe-west1, with data kept in the chosen region; single-tenant and bring-your-own-cloud deployments are available on request.

The main limit when self-hosting is memory: a single node keeps each HNSW index in RAM, and once collections outgrow it performance collapses, so you size the machine for your total embeddings. Chroma Cloud has its own per-collection caps, such as 5 million records and 300 records per write.

Where it fits

How Chroma itself is built

11 tools, from its own code, website and Product Hunt page.

Who uses it

13 makers' products, each linked to the source that shows it, and 69 open-source projects that declare it in their code.

The maker says so 10Customer story 3Declared in code 69How evidence is collected →
VoltAgentBuild TS AI agents with n8n-style observability

“Chroma makes it super easy to manage embeddings for AI apps. We love the open-source focus and how quickly it integrates into RAG pipelines.”

Vector DatabaseMaker says so · source ↗
ClarmTurn visitors into pipeline, automatically

“Jeff (the founder) is incredible - super knowledgeable and I'm super bullish on the direction of the product. Let's go!”

Vector DatabaseMaker says so · source ↗
ReefFrom chat to downloadable Excel, with every formula built in

“One of the best vectorDB out there. Easy and opensource”

Vector DatabaseMaker says so · source ↗
AllysonYour AI Executive Assistant

“ChromaDB is the perfect solution for our vector database needs, allowing us to store and manage all our embeddings efficiently. As an open-source AI application database, ChromaDB offers the flexibility and scalability required for our various AI projects. Its robust features ensure that our data infrastructure can grow alongside our applications, providing a solid foundation for our AI-driven initiatives.”

Vector DatabaseMaker says so · source ↗
ConvoMemory & observability for LLM apps

“Powered memory storage with a dead-simple, blazing-fast open-source vector DB. Far easier to self-host than alternatives.”

Vector DatabaseMaker says so · source ↗
DocuSparkChat with Documents. Get Answers. Take Action. Fully Private

“Blazing fast Vector Database, easy to implement, and stack it up especially for Docuspark.”

Vector DatabaseMaker says so · source ↗
Aix-DBAix-DB 基于 LangChain/LangGraph 框架,结合 MCP Skills 多智能体协作架构,实现自然语言到数据洞察的端到端转换。Vector DatabaseIn its code · source ↗

Open source: a project that declares Chroma as a dependency in its public code — verifiable, but not necessarily a live product.

What makers say

5 makers on why they use Chroma, in their own words on Product Hunt.

Chroma makes it super easy to manage embeddings for AI apps. We love the open-source focus and how quickly it integrates into RAG pipelines.
VoltAgentSep 2026 ↗
Jeff (the founder) is incredible - super knowledgeable and I'm super bullish on the direction of the product. Let's go!
ClarmSep 2026 ↗
ChromaDB is the perfect solution for our vector database needs, allowing us to store and manage all our embeddings efficiently. As an open-source AI application database, ChromaDB offers the flexibility and scalability required for our various AI projects. Its robust features ensure that our data infrastructure can grow alongside our applications, providing a solid foundation for our AI-driven initiatives.
AllysonSep 2026 ↗
Powered memory storage with a dead-simple, blazing-fast open-source vector DB. Far easier to self-host than alternatives.
ConvoSep 2026 ↗
Blazing fast Vector Database, easy to implement, and stack it up especially for Docuspark.
DocuSparkSep 2026 ↗

Loved and watch-outs

Themes that recur in makers' words and Hacker News comments, each linked to what it summarises, with how Product Hunt tags its reviews.

Most loved
  • Open source and easy to self-host, with a simple API that gets embeddings stored quickly. PHHN
  • Full-text and regex search sit alongside vector search, and collection forking suits changing code. trychroma.comHNHN 2HN 3
  • Plugs quickly into RAG pipelines and local-first tools built on LangChain or Ollama. PHHNHN 2HN 3
  • The serverless cloud version has no knobs to tune and a usage-based tier. HNHN 2
Watch-outs
  • Its feature set is narrower than Milvus or Weaviate, lacking vector quantization and some index options. HNHN 2

Reliability and open issues

Most wanted on GitHubOpen on 2026-10-04 in chroma-core/chroma, with activity in the last year — issues and feature requests by 👍.

Who switches

Public pull requests on GitHub since Oct 2024 whose title says "X to Y" — real code changes moving a project from one tool to another, by developers in general. Open a row to see the pull requests.

Chroma → pgvector7 PRs
Chroma → Pinecone6 PRs
Chroma → Qdrant4 PRs

Alternatives to Chroma

All alternatives by situation →

Questions makers ask about Chroma

Can I run Chroma inside my app without a server?

In Python, yes. The in-memory client runs Chroma in your process, and the persistent client saves to a local folder (.chroma by default). The JavaScript/TypeScript client always needs a Chroma server. source ↗

How much RAM does a self-hosted Chroma need?

The vector index must fit in memory. Chroma's own tests put the capacity at roughly 0.245 million 1024-dimension embeddings per GB of RAM, plus about a gigabyte for the system, and they advise against machines under 2 GB. source ↗

Can I move collections between my own server and Chroma Cloud?

Yes. The chroma copy CLI command copies chosen collections, or all of them, from a local server to Cloud and back. source ↗

Where is Chroma Cloud data stored, and is it SOC 2 certified?

Each database stays in the region you pick, currently AWS us-east-1 or GCP europe-west1. Chroma Cloud is SOC 2 Type II certified, and single-tenant or in-your-own-VPC deployments are available by contacting the team. source ↗

What limits apply on Chroma Cloud?

Defaults include 4,096 embedding dimensions, 5 million records per collection, 300 records per write and 300 results per query, plus 10 concurrent reads and 10 concurrent writes per collection. Most can be raised on request. source ↗

Does it work with LangChain and LlamaIndex?

Yes. Both have Chroma integrations, as do Haystack and other frameworks listed in the docs. source ↗

Does self-hosted Chroma send telemetry?

No. Since version 1.5.4 Chroma no longer collects product telemetry; you can export your own OpenTelemetry data, which is never sent to Chroma. source ↗

Is Chroma free?

Yes — there is a free tier a small product can run on; paid use starts at Usage-based. source ↗

Is Chroma open source or self-hostable?

Open source, and you can self-host it. source ↗

Can AI coding agents work with Chroma?

It serves an llms.txt docs index.

Who uses Chroma?

13 makers' products we track, each with a source, and 69 open-source projects declare it in their code. source ↗