Chroma MCP Server
OfficialChroma MCP サーバー
モデル コンテキスト プロトコル (MCP)は、LLM アプリケーションと外部データ ソースまたはツール間の簡単な統合のために設計されたオープン プロトコルであり、LLM に必要なコンテキストをシームレスに提供するための標準化されたフレームワークを提供します。
このサーバーは、Chroma を利用したデータ取得機能を提供し、AI モデルが生成されたデータとユーザー入力に基づいてコレクションを作成し、ベクター検索、全文検索、メタデータ フィルタリングなどを使用してそのデータを取得できるようにします。
特徴
柔軟なクライアントタイプ
テストと開発用の一時的(メモリ内)
ファイルベースのストレージの永続性
セルフホスト型 Chroma インスタンス用の HTTP クライアント
Chroma Cloud 統合用のクラウド クライアント (api.trychroma.com に自動的に接続します)
コレクション管理
コレクションの作成、変更、削除
ページ区切りをサポートするすべてのコレクションを一覧表示する
コレクション情報と統計を取得する
最適化されたベクトル検索のためのHNSWパラメータを設定する
コレクションを作成するときに埋め込み関数を選択する
ドキュメント操作
オプションのメタデータとカスタムIDを使用してドキュメントを追加する
セマンティック検索を使用してドキュメントをクエリする
メタデータとドキュメントコンテンツを使用した高度なフィルタリング
IDまたはフィルターでドキュメントを取得する
全文検索機能
サポートされているツール
chroma_list_collections- ページネーションをサポートするすべてのコレクションを一覧表示しますchroma_create_collection- オプションのHNSW設定で新しいコレクションを作成するchroma_peek_collection- コレクション内のドキュメントのサンプルを表示するchroma_get_collection_info- コレクションの詳細情報を取得するchroma_get_collection_count- コレクション内のドキュメント数を取得するchroma_modify_collection- コレクションの名前またはメタデータを更新するchroma_delete_collection- コレクションを削除するchroma_add_documents- オプションのメタデータとカスタムIDを使用してドキュメントを追加するchroma_query_documents- 高度なフィルタリングを備えたセマンティック検索を使用してドキュメントをクエリしますchroma_get_documents- IDまたはページ区切りのフィルターでドキュメントを取得しますchroma_update_documents- 既存のドキュメントのコンテンツ、メタデータ、または埋め込みを更新するchroma_delete_documents- コレクションから特定のドキュメントを削除する
埋め込み関数
Chroma MCP は、 default 、 cohere 、 openai 、 jina 、 voyageai 、 roboflowなどのいくつかの埋め込み関数をサポートしています。
埋め込み関数はChromaのコレクション設定を利用します。この設定は、コレクションの埋め込み関数を永続化して取得できるようにします。コレクション設定を使用してコレクションを作成すると、将来のクエリや挿入で取得する際に同じ埋め込み関数が使用され、埋め込み関数を再度指定する必要はありません。埋め込み関数の永続化はChromaのバージョン1.0.0で追加されたため、バージョン0.6.3以下でコレクションを作成した場合、この機能はサポートされません。
外部APIを利用する埋め込み関数にアクセスする場合は、埋め込み関数の環境変数に記載されている正しい形式でAPIキーの環境変数を追加してください。
Related MCP server: PDF Knowledgebase MCP Server
Claude Desktopでの使用
一時クライアントを追加するには、
claude_desktop_config.jsonファイルに次の内容を追加します。
"chroma": {
"command": "uvx",
"args": [
"chroma-mcp"
]
}永続クライアントを追加するには、
claude_desktop_config.jsonファイルに以下を追加します。
"chroma": {
"command": "uvx",
"args": [
"chroma-mcp",
"--client-type",
"persistent",
"--data-dir",
"/full/path/to/your/data/directory"
]
}これにより、指定されたデータ ディレクトリを使用する永続クライアントが作成されます。
Chroma Cloud に接続するには、
claude_desktop_config.jsonファイルに以下を追加します。
"chroma": {
"command": "uvx",
"args": [
"chroma-mcp",
"--client-type",
"cloud",
"--tenant",
"your-tenant-id",
"--database",
"your-database-name",
"--api-key",
"your-api-key"
]
}これにより、SSL を使用して api.trychroma.com に自動的に接続するクラウド クライアントが作成されます。
**注:**引数に API キーを追加することはローカル デバイスでは問題ありませんが、安全のため、 argsリスト内の--dotenv-path引数を使用して環境構成ファイルのカスタム パスを指定することもできます。例: "args": ["chroma-mcp", "--dotenv-path", "/custom/path/.env"] 。
[独自のクラウド プロバイダー上のセルフホスト型 Chroma インスタンス]( https://docs.trychroma.com/production/deployment )に接続するには、
claude_desktop_config.jsonファイルに次のコードを追加します。
"chroma": {
"command": "uvx",
"args": [
"chroma-mcp",
"--client-type",
"http",
"--host",
"your-host",
"--port",
"your-port",
"--custom-auth-credentials",
"your-custom-auth-credentials",
"--ssl",
"true"
]
}これにより、セルフホストされた Chroma インスタンスに接続する HTTP クライアントが作成されます。
デモ
共有ナレッジベースやコンテキストウィンドウへのメモリの追加などの参考使用法については、Chroma MCP ドキュメントをご覧ください。
環境変数の使用
環境変数を使ってクライアントを設定することもできます。サーバーは、 --dotenv-pathで指定されたパス(デフォルトは作業ディレクトリの.chroma_env )にある.envファイル、またはシステム環境変数から変数を自動的に読み込みます。コマンドライン引数は環境変数よりも優先されます。
# Common variables
export CHROMA_CLIENT_TYPE="http" # or "cloud", "persistent", "ephemeral"
# For persistent client
export CHROMA_DATA_DIR="/full/path/to/your/data/directory"
# For cloud client (Chroma Cloud)
export CHROMA_TENANT="your-tenant-id"
export CHROMA_DATABASE="your-database-name"
export CHROMA_API_KEY="your-api-key"
# For HTTP client (self-hosted)
export CHROMA_HOST="your-host"
export CHROMA_PORT="your-port"
export CHROMA_CUSTOM_AUTH_CREDENTIALS="your-custom-auth-credentials"
export CHROMA_SSL="true"
# Optional: Specify path to .env file (defaults to .chroma_env)
export CHROMA_DOTENV_PATH="/path/to/your/.env" 埋め込み関数の環境変数
APIキーにアクセスする外部埋め込み関数を使用する場合は、 CHROMA_<>_API_KEY="<key>"命名規則に従ってください。Cohere APIキーを設定するには、環境変数CHROMA_COHERE_API_KEY=""を設定します。この環境変数を.envファイルに追加し、 CHROMA_DOTENV_PATH環境変数または--dotenv-pathフラグを使用して、その場所を安全に保管できるように設定することをお勧めします。
Available Tools
13 toolschroma_add_documentsC
Add documents to a Chroma collection.
Args:
collection_name: Name of the collection to add documents to
documents: List of text documents to add
ids: List of IDs for the documents (required)
metadatas: Optional list of metadata dictionaries for each document
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes | ||
| documents | Yes | ||
| ids | Yes | ||
| metadatas | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description lacks any behavioral disclosure beyond the basic action. It does not mention whether duplicate IDs cause errors or overwrite, whether documents are validated for size or format, or what the return value indicates. With no annotations provided, the description carries the full burden and fails to address key behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is structured with an Args block, which is clear, but it largely restates the schema information. While not excessively long, it could be more concise by omitting redundant parameter descriptions and focusing on unique behavioral details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool mutates data and has no output schema or annotations, the description is incomplete. It does not explain the return value, constraints (e.g., document count limits), or side effects of adding documents to an existing collection. The agent lacks enough context for reliable use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description provides basic semantics for each parameter (e.g., 'ids: List of IDs for the documents (required)'). This adds moderate value beyond the parameter names but lacks depth, such as uniqueness constraints for 'ids' or format expectations for 'documents'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool adds documents to a Chroma collection, using the verb 'add' and specifying the resource 'documents to a Chroma collection'. However, it does not differentiate from sibling tools like 'chroma_update_documents' or 'chroma_delete_documents', which share similar contexts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. For instance, it does not state that the collection must exist before adding, nor does it explain when to use 'add' over 'update' or 'delete'. The agent is left without context for decision-making.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_create_collectionB
Create a new Chroma collection with configurable HNSW parameters.
Args:
collection_name: Name of the collection to create
embedding_function_name: Name of the embedding function to use. Options: 'default', 'cohere', 'openai', 'jina', 'voyageai', 'ollama', 'roboflow'
metadata: Optional metadata dict to add to the collection
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes | ||
| embedding_function_name | No | default | |
| metadata | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It mentions 'configurable HNSW parameters' but doesn't explain what these are, their defaults, or behavioral traits like error handling, permissions needed, or what happens on duplicate collection names. This leaves significant gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and front-loaded with the main purpose, followed by a structured Args section. Each sentence adds value, though the HNSW reference is vague and could be more precise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and 3 parameters with 0% schema coverage, the description is incomplete. It explains parameters well but lacks behavioral context (e.g., mutation effects, error cases) and doesn't address the HNSW configuration mentioned, leaving gaps for a creation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning by explaining each parameter's purpose: collection_name for naming, embedding_function_name with specific options, and metadata as an optional dict. This covers all 3 parameters adequately, though it doesn't detail HNSW parameters mentioned in the opening.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Create a new Chroma collection') and resource ('Chroma collection'), distinguishing it from siblings like chroma_delete_collection or chroma_modify_collection. However, it doesn't fully specify what 'configurable HNSW parameters' means, which slightly reduces specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for creating collections, but provides no explicit guidance on when to use this tool versus alternatives like chroma_fork_collection or chroma_modify_collection. It lists embedding function options, which hints at context, but lacks clear when/when-not instructions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_delete_collectionB
Delete a Chroma collection.
Args:
collection_name: Name of the collection to delete
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description bears full responsibility for behavioral disclosure. However, it merely says "Delete a Chroma collection" without stating that the operation is irreversible, requires the collection to exist, or any side effects (e.g., data loss). This is a critical gap for a deletion tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise, consisting of one sentence for the tool and one line for the parameter. It is front-loaded with the core action and resource, with no extraneous information. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite the tool having only one parameter, no output schema, and no annotations, the description is too sparse. It omits crucial context for a deletion operation, such as whether the action is reversible, what happens to associated data, and what the response looks like. A more complete description would include warnings about irreversibility and conditions for success.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must compensate. It adds a brief explanation for the single parameter ("Name of the collection to delete"), which clarifies its role but does not provide additional details like format, validation rules, or examples. This is minimal but sufficient for a simple string parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ("Delete") and the resource ("a Chroma collection"), which distinguishes it from sibling tools like chroma_create_collection or chroma_list_collections. The verb and resource are specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives, such as chroma_modify_collection or chroma_fork_collection. There is no mention of prerequisites (e.g., collection must exist) or when not to use it, leaving the agent with limited context for decision-making.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_delete_documentsA
Delete documents from a Chroma collection.
Args:
collection_name: Name of the collection to delete documents from
ids: List of document IDs to delete
Returns:
A confirmation message indicating the number of documents deleted.
Raises:
ValueError: If 'ids' is empty
Exception: If the collection does not exist or if the delete operation fails.
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes | ||
| ids | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description lists error conditions (empty ids, missing collection) and mentions it returns a confirmation message, but it does not state that the operation is irreversible or discuss side effects. Given no annotations, more behavioral details could be added.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured with a clear purpose statement followed by parameter, return, and error sections. Every sentence adds value without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple deletion tool with two required parameters, the description covers purpose, parameters, returns, and errors. However, it omits details on permanence, handling of non-existent IDs, and usage context, leaving some gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates by explaining the purpose of each parameter: 'collection_name' is the collection's name, 'ids' are document IDs. While basic, it adds necessary meaning beyond the schema's titles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with 'Delete documents from a Chroma collection,' which is a specific verb-resource combination that clearly differentiates from sibling tools like adding, querying, or listing documents.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. For example, it does not mention prerequisites (e.g., obtaining IDs via get_documents) or when deletion is appropriate, which is critical given 12 siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_fork_collectionC
Fork a Chroma collection.
Args:
collection_name: Name of the collection to fork
new_collection_name: Name of the new collection to create
metadata: Optional metadata dict to add to the new collection
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes | ||
| new_collection_name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description lacks behavioral details. It does not disclose whether the fork is deep or shallow, if metadata is copied, or if the operation is reversible. The brief description leaves significant ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short but structured as a docstring with repeated 'Args:' lines. It could be more concise, e.g., by combining the purpose and parameter descriptions. Every sentence is functional but not optimally efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and the complexity of a mutation tool, the description is insufficient. It fails to explain what happens to existing data, whether embeddings are copied, or how success is indicated. The tool's behavior remains opaque.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds minimal meaning beyond parameter names: 'Name of the collection to fork' and 'Name of the new collection to create' are largely redundant. It also mentions an optional 'metadata' parameter that is not present in the input schema, creating inconsistency. Schema coverage is 0%, so the description should compensate more.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'Fork a Chroma collection', which is a specific verb and resource indicating duplication. However, it does not differentiate from sibling tools like chroma_create_collection or chroma_delete_collection, as 'fork' could imply partial copying or aliasing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. It does not mention prerequisites, such as the original collection must exist, or contrast with creating a new collection from scratch.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_get_collection_countC
Get the number of documents in a Chroma collection.
Args:
collection_name: Name of the collection to count
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears full burden. It merely restates the obvious action without disclosing return format, error handling (e.g., missing collection), or side effects. This is a serious gap for a tool that returns a computed value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise and front-loaded with the purpose sentence. However, it is overly minimal and could include more details in the same space (e.g., return type). It is acceptable but not exemplary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite low complexity (1 param, no output schema), the description fails to specify the return format (e.g., integer, JSON object) or behavior on errors. The agent lacks sufficient information to confidently invoke and parse results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description should add meaning to the lone parameter. It says 'Name of the collection to count,' which slightly clarifies purpose but repeats the schema title. No additional constraints, formats, or examples are provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Get the number of documents in a Chroma collection' with a specific verb and resource. It uniquely identifies the tool's function and distinguishes it from siblings like chroma_list_collections (list collections) and chroma_get_collection_info (get info).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, such as chroma_get_collection_info which might also provide count information. There are no prerequisites, exclusions, or context for optimal usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_get_collection_infoC
Get information about a Chroma collection.
Args:
collection_name: Name of the collection to get info about
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose all behavioral traits. It only states the basic operation without explaining error handling (e.g., if collection doesn't exist), read-only nature, or return format. This leaves significant ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with a single line and structured docstring. No redundant information, though it could be more informative without being verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description should explain what data is returned (e.g., metadata fields). It does not. The simple parameter and operation suggest a straightforward tool, but missing return type and error behavior leave gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%. The description's Args section merely repeats the parameter name and a trivial description ('Name of the collection'), adding little beyond the schema's title. For a single parameter, the baseline expectation is higher.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action and resource ('Get information about a Chroma collection'), which distinguishes it from siblings like chroma_list_collections (list all) and chroma_get_collection_count (count). However, it lacks specifics on what 'information' entails.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, such as chroma_list_collections or chroma_get_collection_count. No prerequisites or typical scenarios are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_get_documentsA
Get documents from a Chroma collection with optional filtering.
Args:
collection_name: Name of the collection to get documents from
ids: Optional list of document IDs to retrieve
where: Optional metadata filters using Chroma's query operators
Examples:
- Simple equality: {"metadata_field": "value"}
- Comparison: {"metadata_field": {"$gt": 5}}
- Logical AND: {"$and": [{"field1": {"$eq": "value1"}}, {"field2": {"$gt": 5}}]}
- Logical OR: {"$or": [{"field1": {"$eq": "value1"}}, {"field1": {"$eq": "value2"}}]}
where_document: Optional document content filters
Examples:
- Contains: {"$contains": "value"}
- Not contains: {"$not_contains": "value"}
- Regex: {"$regex": "[a-z]+"}
- Not regex: {"$not_regex": "[a-z]+"}
- Logical AND: {"$and": [{"$contains": "value1"}, {"$not_regex": "[a-z]+"}]}
- Logical OR: {"$or": [{"$regex": "[a-z]+"}, {"$not_contains": "value2"}]}
include: List of what to include in response. By default, this will include documents, and metadatas.
limit: Optional maximum number of documents to return
offset: Optional number of documents to skip before returning results
Returns:
Dictionary containing the matching documents, their IDs, and requested includes
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes | ||
| ids | No | ||
| include | No | ||
| limit | No | ||
| offset | No | ||
| where | No | ||
| where_document | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It describes retrieval behavior with filtering, pagination, and includes, but does not discuss error conditions, authentication needs, rate limits, or whether the operation is idempotent. The read-only nature is implied but not explicitly stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a brief summary followed by a bulleted parameter list. It is moderately long but every sentence adds value. Could be slightly more concise by removing redundant wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 7 parameters and no output schema, the description covers most aspects: parameter details, examples, and a general return description. Missing information includes possible error responses and collection existence requirements, but overall sufficiently complete for effective tool usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the description extensively explains each parameter with clear examples for complex objects like 'where' and 'where_document', including supported operators (comparison, logical, regex). Default values for 'include' are noted. This adds significant value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Get documents from a Chroma collection with optional filtering,' which is a specific verb+resource. It naturally distinguishes from sibling tools like chroma_query_documents (which uses vector similarity) and chroma_add_documents.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for retrieving documents with metadata/content filtering, but does not explicitly guide when to use this tool over alternatives like chroma_query_documents (which uses embedding similarity). No exclusions or when-not-to-use guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_list_collectionsA
List all collection names in the Chroma database with pagination support.
Args:
limit: Optional maximum number of collections to return
offset: Optional number of collections to skip before returning results
Returns:
List of collection names or ["__NO_COLLECTIONS_FOUND__"] if database is empty
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the return format (list of names or a sentinel value) and mentions pagination. However, it does not clarify behavior when limit/offset exceed bounds (e.g., does it return an error or empty list?). The phrase 'list all' conflicts with pagination, causing slight ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise: a single-line purpose, followed by Args and Returns sections. Every sentence adds value, and the structure is clean and easy to parse. No unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and 2 optional parameters, the description reasonably covers purpose and return format. However, it lacks details on pagination defaults (e.g., default limit when omitted) and error handling for invalid parameters. The slight inconsistency 'list all' vs pagination reduces completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% (parameters lack descriptions), so the description must compensate. It explains 'limit: Optional maximum number of collections to return' and 'offset: Optional number of collections to skip before returning results.' This adds meaningful semantics beyond the raw schema, clarifying defaults (None) and optionality.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'List all collection names in the Chroma database with pagination support.' This specific verb ('list') and resource ('collection names') distinguishes it from siblings like chroma_get_collection_info (which retrieves details of a single collection) and chroma_create_collection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly guide when to use this tool vs alternatives. It implies usage for listing collection names, but lacks cues like 'use this to get an overview of collections; for details on a specific collection, use chroma_get_collection_info.' The pagination support is mentioned but no guidance on setting limit/offset.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_modify_collectionB
Modify a Chroma collection's name or metadata.
Args:
collection_name: Name of the collection to modify
new_name: Optional new name for the collection
new_metadata: Optional new metadata for the collection
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes | ||
| new_metadata | No | ||
| new_name | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description must carry full burden. It only says 'modify' implying mutation, but lacks details on side effects, revertibility, required permissions, or whether other collection properties are affected. Minimal disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Description is short and front-loaded with purpose. The docstring lists parameters efficiently with no extraneous text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and minimal schema coverage; description should provide more context about return values, errors, or behavioral outcomes of modification. It lacks completeness for a tool with no annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, but description explains all three parameters (collection_name, new_name, new_metadata) in a docstring format. It adds basic meaning beyond property titles, though no detailed constraints or formats.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Modify a Chroma collection's name or metadata' with a specific verb and resource. It distinguishes from sibling tools like chroma_list_collections, chroma_create_collection, etc., which focus on different operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, no prerequisites, and no conditions for safe usage. The description only lists parameters without context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_peek_collectionB
Peek at documents in a Chroma collection.
Args:
collection_name: Name of the collection to peek into
limit: Number of documents to peek at
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes | ||
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, and the description does not disclose any behavioral traits beyond 'peek at documents'. It does not mention side effects, rate limits, authentication requirements, or what happens if the collection is empty or doesn't exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with a clear purpose sentence followed by structured parameter descriptions. Every sentence serves a purpose; no unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with only 2 parameters and no output schema, the description covers the basic functionality and parameters. However, it lacks details about the return format or behavior in edge cases, which would be helpful but not critical for such a straightforward operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaningful information for both parameters: collection_name is described as 'Name of the collection to peek into' and limit as 'Number of documents to peek at'. Since schema description coverage is 0%, this provides necessary clarity beyond the schema's titles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'Peek at documents in a Chroma collection', which clearly indicates a read operation on a collection. However, it does not distinguish from sibling tools like chroma_get_documents or chroma_query_documents, which also retrieve documents.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool over alternatives. There is no mention of use cases, prerequisites, or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_query_documentsA
Query documents from a Chroma collection with advanced filtering.
Args:
collection_name: Name of the collection to query
query_texts: List of query texts to search for
n_results: Number of results to return per query
where: Optional metadata filters using Chroma's query operators
Examples:
- Simple equality: {"metadata_field": "value"}
- Comparison: {"metadata_field": {"$gt": 5}}
- Logical AND: {"$and": [{"field1": {"$eq": "value1"}}, {"field2": {"$gt": 5}}]}
- Logical OR: {"$or": [{"field1": {"$eq": "value1"}}, {"field1": {"$eq": "value2"}}]}
where_document: Optional document content filters
Examples:
- Contains: {"$contains": "value"}
- Not contains: {"$not_contains": "value"}
- Regex: {"$regex": "[a-z]+"}
- Not regex: {"$not_regex": "[a-z]+"}
- Logical AND: {"$and": [{"$contains": "value1"}, {"$not_regex": "[a-z]+"}]}
- Logical OR: {"$or": [{"$regex": "[a-z]+"}, {"$not_contains": "value2"}]}
include: List of what to include in response. By default, this will include documents, metadatas, and distances.
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes | ||
| include | No | ||
| n_results | No | ||
| query_texts | Yes | ||
| where | No | ||
| where_document | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears the full burden. It discloses default response fields (documents, metadatas, distances), supported filter operators with examples, and use of n_results for pagination. However, it lacks details on rate limits, maximum result limits, or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a well-structured docstring that front-loads the purpose and then systematically describes each parameter. While it includes many examples that increase length, these are necessary for complex parameters and are not redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 6 parameters, no output schema, and no annotations, the description covers all parameters with examples. It mentions the default include list but does not detail the response structure beyond field names. The tool's purpose is clear, but some behavioral details are omitted.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description adds substantial meaning to all 6 parameters, especially for complex fields where and where_document, providing detailed examples of supported operators and logical combinations. This goes far beyond the schema's type and default information.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action 'Query documents from a Chroma collection with advanced filtering,' with a specific verb and resource. It distinguishes from siblings like add, get, update, delete documents by focusing on querying with filtering.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for querying with advanced filtering but provides no explicit guidance on when to use this tool versus alternatives like get_documents or peek_collection. No when-not-to-use or alternative naming is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_update_documentsA
Update documents in a Chroma collection.
Args:
collection_name: Name of the collection to update documents in
ids: List of document IDs to update (required)
embeddings: Optional list of new embeddings for the documents.
Must match length of ids if provided.
metadatas: Optional list of new metadata dictionaries for the documents.
Must match length of ids if provided.
documents: Optional list of new text documents.
Must match length of ids if provided.
Returns:
A confirmation message indicating the number of documents updated.
Raises:
ValueError: If 'ids' is empty or if none of 'embeddings', 'metadatas',
or 'documents' are provided, or if the length of provided
update lists does not match the length of 'ids'.
Exception: If the collection does not exist or if the update operation fails.
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes | ||
| documents | No | ||
| embeddings | No | ||
| ids | Yes | ||
| metadatas | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It states that the tool updates documents and raises ValueError if inputs are mismatched, but it doesn't disclose behavioral traits such as whether updates are full replacements or incremental, whether the operation is idempotent, or any permission requirements. It adequately describes error conditions but lacks broader behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with bullet points for args, returns, and raises. It is front-loaded with the core purpose and is fairly concise, though the 'Args:' header and bullet format could be slightly tighter without losing clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 5 parameters, no output schema, and no annotations, the description covers the input semantics, return type, and error conditions reasonably well. However, it lacks details on overall behavior like incremental vs full updates, idempotency, or prerequisites such as collection existence (though implied by exception). It is mostly complete for an update tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the description provides detailed explanations for each parameter, including constraints like 'Must match length of ids if provided.' This adds significant meaning beyond the schema's property titles and types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'Update documents in a Chroma collection' which is a specific verb+resource. It clearly distinguishes from sibling tools like add_documents (add new), delete_documents (delete), query_documents (query), etc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains what can be updated (embeddings, metadatas, documents) but does not provide explicit guidance on when to use this tool vs alternatives like add_documents (for new documents) or modify_collection (for collection settings). Usage is implied but not clearly contrasted.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
13 tool updates
v1.0.0- First observed
chroma_add_documents - First observed
chroma_create_collection - First observed
chroma_delete_collection - First observed
chroma_delete_documents - First observed
chroma_fork_collection - First observed
chroma_get_collection_count - First observed
chroma_get_collection_info - First observed
chroma_get_documents - First observed
chroma_list_collections - First observed
chroma_modify_collection - First observed
chroma_peek_collection - First observed
chroma_query_documents - First observed
chroma_update_documents
TDQS
Scored across 13 tools
Each tool has a clearly distinct purpose targeting specific Chroma operations. For example, chroma_get_documents retrieves documents with filtering, while chroma_query_documents performs semantic search; chroma_peek_collection provides a quick preview, and chroma_get_collection_info returns metadata. No tools appear to overlap in functionality.
All tools follow a consistent 'chroma_verb_noun' pattern with snake_case throughout. The naming convention is perfectly uniform, making it easy to predict tool names and understand their functions at a glance.
With 13 tools, this server provides comprehensive coverage for Chroma vector database operations without being overwhelming. The count aligns well with the domain scope, offering complete collection management and document CRUD operations.
The toolset provides complete coverage for Chroma operations: collection lifecycle (create, list, get info, modify, fork, delete), document lifecycle (add, get, update, delete, query), and utility functions (count, peek). No obvious gaps exist for typical vector database workflows.
Maintenance
Related MCP Connectors
The Needle MCP server enables semantic search on documents stored in files like PDFs, DOCX, and XLSX by connecting AI applications to external data sources. It provides capabilities to create and manage document collections, perform natural language searches on stored content, and retrieve relevant information without requiring exact keyword matches.
Ingest, manage, and retrieve documents for RAG-powered AI applications
The CustomGPT.ai MCP server is a fully managed, RAG-powered endpoint that connects large language models with private knowledge bases and external data sources. It provides tools for retrieval-augmented generation queries (send_message), data ingestion (upload_file), and source listing, enabling AI agents to query private documents like PDFs with high accuracy and real-time citations.
Remote ChromaDB vector database MCP server with streamable HTTP transport
Related MCP Servers
- AlicenseAqualityDmaintenanceA Model Context Protocol server providing vector database capabilities through Chroma, enabling semantic document search, metadata filtering, and document management with persistent storage.641MIT
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol server that enables intelligent document search and retrieval from PDF collections, providing semantic search capabilities powered by OpenAI embeddings and ChromaDB vector storage.13MIT
- AlicenseNot gradedqualityDmaintenanceAn MCP server that exposes ChromaDB vector database operations, enabling AI assistants to perform collection management and semantic document searches. It supports HTTP, persistent, and in-memory connection modes along with various embedding providers including OpenAI and HuggingFace.MIT
- FlicenseNot gradedqualityDmaintenanceA fully offline local RAG server that utilizes ChromaDB and Ollama to index and query PDF, text, and Markdown documents. It allows users to manage local knowledge bases and perform semantic searches with AI-generated responses.-