Chroma MCP Server
OfficialChroma MCP 服务器
模型上下文协议 (MCP)是一种开放协议,旨在实现 LLM 应用程序与外部数据源或工具之间的轻松集成,提供标准化框架,无缝地为 LLM 提供所需的上下文。
该服务器提供由 Chroma 提供支持的数据检索功能,使 AI 模型能够基于生成的数据和用户输入创建集合,并使用矢量搜索、全文搜索、元数据过滤等检索该数据。
特征
灵活的客户类型
用于测试和开发的临时(内存中)
基于文件的持久存储
用于自托管 Chroma 实例的 HTTP 客户端
用于 Chroma Cloud 集成的云客户端(自动连接到 api.trychroma.com)
收藏管理
创建、修改和删除集合
列出所有支持分页的集合
获取收集信息和统计数据
配置 HNSW 参数以优化矢量搜索
创建集合时选择嵌入函数
文档操作
添加具有可选元数据和自定义 ID 的文档
使用语义搜索查询文档
使用元数据和文档内容进行高级过滤
通过 ID 或过滤器检索文档
全文搜索功能
支持的工具
chroma_list_collections- 列出所有支持分页的集合chroma_create_collection- 使用可选的 HNSW 配置创建新集合chroma_peek_collection- 查看集合中的文档样本chroma_get_collection_info- 获取有关集合的详细信息chroma_get_collection_count- 获取集合中的文档数量chroma_modify_collection- 更新集合的名称或元数据chroma_delete_collection- 删除收藏集chroma_add_documents- 添加带有可选元数据和自定义 ID 的文档chroma_query_documents- 使用带有高级过滤功能的语义搜索查询文档chroma_get_documents- 通过 ID 或分页过滤器检索文档chroma_update_documents- 更新现有文档的内容、元数据或嵌入chroma_delete_documents- 从集合中删除特定文档
嵌入函数
Chroma MCP 支持多种嵌入功能: default 、 cohere 、 openai 、 jina 、 voyageai和roboflow 。
嵌入函数利用 Chroma 的集合配置,该配置会持久化集合中选定的嵌入函数以供检索。使用该集合配置创建集合后,在以后的查询和插入操作中,将使用相同的嵌入函数,而无需再次指定嵌入函数。嵌入函数持久化功能是在 Chroma v1.0.0 版本中添加的,因此如果您使用低于 0.6.3 的版本创建集合,则不支持此功能。
访问使用外部 API 的嵌入函数时,请务必添加格式正确的 API 密钥环境变量,该环境变量可在嵌入函数环境变量中找到
Related MCP server: PDF Knowledgebase MCP Server
与 Claude Desktop 一起使用
要添加临时客户端,请将以下内容添加到您的
claude_desktop_config.json文件中:
"chroma": {
"command": "uvx",
"args": [
"chroma-mcp"
]
}要添加持久客户端,请将以下内容添加到
claude_desktop_config.json文件中:
"chroma": {
"command": "uvx",
"args": [
"chroma-mcp",
"--client-type",
"persistent",
"--data-dir",
"/full/path/to/your/data/directory"
]
}这将创建一个使用指定数据目录的持久客户端。
要连接到 Chroma Cloud,请将以下内容添加到您的
claude_desktop_config.json文件中:
"chroma": {
"command": "uvx",
"args": [
"chroma-mcp",
"--client-type",
"cloud",
"--tenant",
"your-tenant-id",
"--database",
"your-database-name",
"--api-key",
"your-api-key"
]
}这将创建一个使用 SSL 自动连接到 api.trychroma.com 的云客户端。
**注意:**在本地设备上,在参数中添加 API 密钥是可以的,但为了安全起见,您还可以使用args列表中的--dotenv-path参数为环境配置文件指定自定义路径,例如: "args": ["chroma-mcp", "--dotenv-path", "/custom/path/.env"] 。
要连接到[您自己的云提供商上的自托管 Chroma 实例]( https://docs.trychroma.com/production/deployment ),请将以下内容添加到您的
claude_desktop_config.json文件中:
"chroma": {
"command": "uvx",
"args": [
"chroma-mcp",
"--client-type",
"http",
"--host",
"your-host",
"--port",
"your-port",
"--custom-auth-credentials",
"your-custom-auth-credentials",
"--ssl",
"true"
]
}这将创建一个连接到您自托管的 Chroma 实例的 HTTP 客户端。
演示
在Chroma MCP 文档中查找参考用法,例如共享知识库和向上下文窗口添加内存
使用环境变量
您还可以使用环境变量来配置客户端。服务器将自动从--dotenv-path指定路径下的.env文件(默认为工作目录中的.chroma_env文件)或系统环境变量中加载变量。命令行参数的优先级高于环境变量。
# Common variables
export CHROMA_CLIENT_TYPE="http" # or "cloud", "persistent", "ephemeral"
# For persistent client
export CHROMA_DATA_DIR="/full/path/to/your/data/directory"
# For cloud client (Chroma Cloud)
export CHROMA_TENANT="your-tenant-id"
export CHROMA_DATABASE="your-database-name"
export CHROMA_API_KEY="your-api-key"
# For HTTP client (self-hosted)
export CHROMA_HOST="your-host"
export CHROMA_PORT="your-port"
export CHROMA_CUSTOM_AUTH_CREDENTIALS="your-custom-auth-credentials"
export CHROMA_SSL="true"
# Optional: Specify path to .env file (defaults to .chroma_env)
export CHROMA_DOTENV_PATH="/path/to/your/.env" 嵌入函数环境变量
使用访问 API 密钥的外部嵌入函数时,请遵循命名约定CHROMA_<>_API_KEY="<key>" 。因此,要设置 Cohere API 密钥,请设置环境变量CHROMA_COHERE_API_KEY="" 。我们建议将其添加到某个 .env 文件中,并使用CHROMA_DOTENV_PATH环境变量或--dotenv-path标志设置该位置以便妥善保管。
Available Tools
13 toolschroma_add_documentsC
Add documents to a Chroma collection.
Args:
collection_name: Name of the collection to add documents to
documents: List of text documents to add
ids: List of IDs for the documents (required)
metadatas: Optional list of metadata dictionaries for each document
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes | ||
| documents | Yes | ||
| ids | Yes | ||
| metadatas | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description lacks any behavioral disclosure beyond the basic action. It does not mention whether duplicate IDs cause errors or overwrite, whether documents are validated for size or format, or what the return value indicates. With no annotations provided, the description carries the full burden and fails to address key behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is structured with an Args block, which is clear, but it largely restates the schema information. While not excessively long, it could be more concise by omitting redundant parameter descriptions and focusing on unique behavioral details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool mutates data and has no output schema or annotations, the description is incomplete. It does not explain the return value, constraints (e.g., document count limits), or side effects of adding documents to an existing collection. The agent lacks enough context for reliable use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description provides basic semantics for each parameter (e.g., 'ids: List of IDs for the documents (required)'). This adds moderate value beyond the parameter names but lacks depth, such as uniqueness constraints for 'ids' or format expectations for 'documents'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool adds documents to a Chroma collection, using the verb 'add' and specifying the resource 'documents to a Chroma collection'. However, it does not differentiate from sibling tools like 'chroma_update_documents' or 'chroma_delete_documents', which share similar contexts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. For instance, it does not state that the collection must exist before adding, nor does it explain when to use 'add' over 'update' or 'delete'. The agent is left without context for decision-making.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_create_collectionB
Create a new Chroma collection with configurable HNSW parameters.
Args:
collection_name: Name of the collection to create
embedding_function_name: Name of the embedding function to use. Options: 'default', 'cohere', 'openai', 'jina', 'voyageai', 'ollama', 'roboflow'
metadata: Optional metadata dict to add to the collection
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes | ||
| embedding_function_name | No | default | |
| metadata | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It mentions 'configurable HNSW parameters' but doesn't explain what these are, their defaults, or behavioral traits like error handling, permissions needed, or what happens on duplicate collection names. This leaves significant gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and front-loaded with the main purpose, followed by a structured Args section. Each sentence adds value, though the HNSW reference is vague and could be more precise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and 3 parameters with 0% schema coverage, the description is incomplete. It explains parameters well but lacks behavioral context (e.g., mutation effects, error cases) and doesn't address the HNSW configuration mentioned, leaving gaps for a creation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning by explaining each parameter's purpose: collection_name for naming, embedding_function_name with specific options, and metadata as an optional dict. This covers all 3 parameters adequately, though it doesn't detail HNSW parameters mentioned in the opening.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Create a new Chroma collection') and resource ('Chroma collection'), distinguishing it from siblings like chroma_delete_collection or chroma_modify_collection. However, it doesn't fully specify what 'configurable HNSW parameters' means, which slightly reduces specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for creating collections, but provides no explicit guidance on when to use this tool versus alternatives like chroma_fork_collection or chroma_modify_collection. It lists embedding function options, which hints at context, but lacks clear when/when-not instructions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_delete_collectionB
Delete a Chroma collection.
Args:
collection_name: Name of the collection to delete
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description bears full responsibility for behavioral disclosure. However, it merely says "Delete a Chroma collection" without stating that the operation is irreversible, requires the collection to exist, or any side effects (e.g., data loss). This is a critical gap for a deletion tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise, consisting of one sentence for the tool and one line for the parameter. It is front-loaded with the core action and resource, with no extraneous information. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite the tool having only one parameter, no output schema, and no annotations, the description is too sparse. It omits crucial context for a deletion operation, such as whether the action is reversible, what happens to associated data, and what the response looks like. A more complete description would include warnings about irreversibility and conditions for success.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must compensate. It adds a brief explanation for the single parameter ("Name of the collection to delete"), which clarifies its role but does not provide additional details like format, validation rules, or examples. This is minimal but sufficient for a simple string parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ("Delete") and the resource ("a Chroma collection"), which distinguishes it from sibling tools like chroma_create_collection or chroma_list_collections. The verb and resource are specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives, such as chroma_modify_collection or chroma_fork_collection. There is no mention of prerequisites (e.g., collection must exist) or when not to use it, leaving the agent with limited context for decision-making.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_delete_documentsA
Delete documents from a Chroma collection.
Args:
collection_name: Name of the collection to delete documents from
ids: List of document IDs to delete
Returns:
A confirmation message indicating the number of documents deleted.
Raises:
ValueError: If 'ids' is empty
Exception: If the collection does not exist or if the delete operation fails.
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes | ||
| ids | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description lists error conditions (empty ids, missing collection) and mentions it returns a confirmation message, but it does not state that the operation is irreversible or discuss side effects. Given no annotations, more behavioral details could be added.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured with a clear purpose statement followed by parameter, return, and error sections. Every sentence adds value without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple deletion tool with two required parameters, the description covers purpose, parameters, returns, and errors. However, it omits details on permanence, handling of non-existent IDs, and usage context, leaving some gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates by explaining the purpose of each parameter: 'collection_name' is the collection's name, 'ids' are document IDs. While basic, it adds necessary meaning beyond the schema's titles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with 'Delete documents from a Chroma collection,' which is a specific verb-resource combination that clearly differentiates from sibling tools like adding, querying, or listing documents.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. For example, it does not mention prerequisites (e.g., obtaining IDs via get_documents) or when deletion is appropriate, which is critical given 12 siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_fork_collectionC
Fork a Chroma collection.
Args:
collection_name: Name of the collection to fork
new_collection_name: Name of the new collection to create
metadata: Optional metadata dict to add to the new collection
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes | ||
| new_collection_name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, and the description lacks behavioral details. It does not disclose whether the fork is deep or shallow, if metadata is copied, or if the operation is reversible. The brief description leaves significant ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short but structured as a docstring with repeated 'Args:' lines. It could be more concise, e.g., by combining the purpose and parameter descriptions. Every sentence is functional but not optimally efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and the complexity of a mutation tool, the description is insufficient. It fails to explain what happens to existing data, whether embeddings are copied, or how success is indicated. The tool's behavior remains opaque.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds minimal meaning beyond parameter names: 'Name of the collection to fork' and 'Name of the new collection to create' are largely redundant. It also mentions an optional 'metadata' parameter that is not present in the input schema, creating inconsistency. Schema coverage is 0%, so the description should compensate more.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'Fork a Chroma collection', which is a specific verb and resource indicating duplication. However, it does not differentiate from sibling tools like chroma_create_collection or chroma_delete_collection, as 'fork' could imply partial copying or aliasing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. It does not mention prerequisites, such as the original collection must exist, or contrast with creating a new collection from scratch.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_get_collection_countC
Get the number of documents in a Chroma collection.
Args:
collection_name: Name of the collection to count
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears full burden. It merely restates the obvious action without disclosing return format, error handling (e.g., missing collection), or side effects. This is a serious gap for a tool that returns a computed value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise and front-loaded with the purpose sentence. However, it is overly minimal and could include more details in the same space (e.g., return type). It is acceptable but not exemplary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite low complexity (1 param, no output schema), the description fails to specify the return format (e.g., integer, JSON object) or behavior on errors. The agent lacks sufficient information to confidently invoke and parse results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description should add meaning to the lone parameter. It says 'Name of the collection to count,' which slightly clarifies purpose but repeats the schema title. No additional constraints, formats, or examples are provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Get the number of documents in a Chroma collection' with a specific verb and resource. It uniquely identifies the tool's function and distinguishes it from siblings like chroma_list_collections (list collections) and chroma_get_collection_info (get info).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, such as chroma_get_collection_info which might also provide count information. There are no prerequisites, exclusions, or context for optimal usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_get_collection_infoC
Get information about a Chroma collection.
Args:
collection_name: Name of the collection to get info about
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose all behavioral traits. It only states the basic operation without explaining error handling (e.g., if collection doesn't exist), read-only nature, or return format. This leaves significant ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with a single line and structured docstring. No redundant information, though it could be more informative without being verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description should explain what data is returned (e.g., metadata fields). It does not. The simple parameter and operation suggest a straightforward tool, but missing return type and error behavior leave gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%. The description's Args section merely repeats the parameter name and a trivial description ('Name of the collection'), adding little beyond the schema's title. For a single parameter, the baseline expectation is higher.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action and resource ('Get information about a Chroma collection'), which distinguishes it from siblings like chroma_list_collections (list all) and chroma_get_collection_count (count). However, it lacks specifics on what 'information' entails.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, such as chroma_list_collections or chroma_get_collection_count. No prerequisites or typical scenarios are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_get_documentsA
Get documents from a Chroma collection with optional filtering.
Args:
collection_name: Name of the collection to get documents from
ids: Optional list of document IDs to retrieve
where: Optional metadata filters using Chroma's query operators
Examples:
- Simple equality: {"metadata_field": "value"}
- Comparison: {"metadata_field": {"$gt": 5}}
- Logical AND: {"$and": [{"field1": {"$eq": "value1"}}, {"field2": {"$gt": 5}}]}
- Logical OR: {"$or": [{"field1": {"$eq": "value1"}}, {"field1": {"$eq": "value2"}}]}
where_document: Optional document content filters
Examples:
- Contains: {"$contains": "value"}
- Not contains: {"$not_contains": "value"}
- Regex: {"$regex": "[a-z]+"}
- Not regex: {"$not_regex": "[a-z]+"}
- Logical AND: {"$and": [{"$contains": "value1"}, {"$not_regex": "[a-z]+"}]}
- Logical OR: {"$or": [{"$regex": "[a-z]+"}, {"$not_contains": "value2"}]}
include: List of what to include in response. By default, this will include documents, and metadatas.
limit: Optional maximum number of documents to return
offset: Optional number of documents to skip before returning results
Returns:
Dictionary containing the matching documents, their IDs, and requested includes
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes | ||
| ids | No | ||
| include | No | ||
| limit | No | ||
| offset | No | ||
| where | No | ||
| where_document | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It describes retrieval behavior with filtering, pagination, and includes, but does not discuss error conditions, authentication needs, rate limits, or whether the operation is idempotent. The read-only nature is implied but not explicitly stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a brief summary followed by a bulleted parameter list. It is moderately long but every sentence adds value. Could be slightly more concise by removing redundant wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 7 parameters and no output schema, the description covers most aspects: parameter details, examples, and a general return description. Missing information includes possible error responses and collection existence requirements, but overall sufficiently complete for effective tool usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the description extensively explains each parameter with clear examples for complex objects like 'where' and 'where_document', including supported operators (comparison, logical, regex). Default values for 'include' are noted. This adds significant value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Get documents from a Chroma collection with optional filtering,' which is a specific verb+resource. It naturally distinguishes from sibling tools like chroma_query_documents (which uses vector similarity) and chroma_add_documents.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for retrieving documents with metadata/content filtering, but does not explicitly guide when to use this tool over alternatives like chroma_query_documents (which uses embedding similarity). No exclusions or when-not-to-use guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_list_collectionsA
List all collection names in the Chroma database with pagination support.
Args:
limit: Optional maximum number of collections to return
offset: Optional number of collections to skip before returning results
Returns:
List of collection names or ["__NO_COLLECTIONS_FOUND__"] if database is empty
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the return format (list of names or a sentinel value) and mentions pagination. However, it does not clarify behavior when limit/offset exceed bounds (e.g., does it return an error or empty list?). The phrase 'list all' conflicts with pagination, causing slight ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise: a single-line purpose, followed by Args and Returns sections. Every sentence adds value, and the structure is clean and easy to parse. No unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and 2 optional parameters, the description reasonably covers purpose and return format. However, it lacks details on pagination defaults (e.g., default limit when omitted) and error handling for invalid parameters. The slight inconsistency 'list all' vs pagination reduces completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% (parameters lack descriptions), so the description must compensate. It explains 'limit: Optional maximum number of collections to return' and 'offset: Optional number of collections to skip before returning results.' This adds meaningful semantics beyond the raw schema, clarifying defaults (None) and optionality.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'List all collection names in the Chroma database with pagination support.' This specific verb ('list') and resource ('collection names') distinguishes it from siblings like chroma_get_collection_info (which retrieves details of a single collection) and chroma_create_collection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly guide when to use this tool vs alternatives. It implies usage for listing collection names, but lacks cues like 'use this to get an overview of collections; for details on a specific collection, use chroma_get_collection_info.' The pagination support is mentioned but no guidance on setting limit/offset.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_modify_collectionB
Modify a Chroma collection's name or metadata.
Args:
collection_name: Name of the collection to modify
new_name: Optional new name for the collection
new_metadata: Optional new metadata for the collection
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes | ||
| new_metadata | No | ||
| new_name | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description must carry full burden. It only says 'modify' implying mutation, but lacks details on side effects, revertibility, required permissions, or whether other collection properties are affected. Minimal disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Description is short and front-loaded with purpose. The docstring lists parameters efficiently with no extraneous text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and minimal schema coverage; description should provide more context about return values, errors, or behavioral outcomes of modification. It lacks completeness for a tool with no annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, but description explains all three parameters (collection_name, new_name, new_metadata) in a docstring format. It adds basic meaning beyond property titles, though no detailed constraints or formats.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Modify a Chroma collection's name or metadata' with a specific verb and resource. It distinguishes from sibling tools like chroma_list_collections, chroma_create_collection, etc., which focus on different operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, no prerequisites, and no conditions for safe usage. The description only lists parameters without context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_peek_collectionB
Peek at documents in a Chroma collection.
Args:
collection_name: Name of the collection to peek into
limit: Number of documents to peek at
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes | ||
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, and the description does not disclose any behavioral traits beyond 'peek at documents'. It does not mention side effects, rate limits, authentication requirements, or what happens if the collection is empty or doesn't exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with a clear purpose sentence followed by structured parameter descriptions. Every sentence serves a purpose; no unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with only 2 parameters and no output schema, the description covers the basic functionality and parameters. However, it lacks details about the return format or behavior in edge cases, which would be helpful but not critical for such a straightforward operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaningful information for both parameters: collection_name is described as 'Name of the collection to peek into' and limit as 'Number of documents to peek at'. Since schema description coverage is 0%, this provides necessary clarity beyond the schema's titles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'Peek at documents in a Chroma collection', which clearly indicates a read operation on a collection. However, it does not distinguish from sibling tools like chroma_get_documents or chroma_query_documents, which also retrieve documents.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool over alternatives. There is no mention of use cases, prerequisites, or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_query_documentsA
Query documents from a Chroma collection with advanced filtering.
Args:
collection_name: Name of the collection to query
query_texts: List of query texts to search for
n_results: Number of results to return per query
where: Optional metadata filters using Chroma's query operators
Examples:
- Simple equality: {"metadata_field": "value"}
- Comparison: {"metadata_field": {"$gt": 5}}
- Logical AND: {"$and": [{"field1": {"$eq": "value1"}}, {"field2": {"$gt": 5}}]}
- Logical OR: {"$or": [{"field1": {"$eq": "value1"}}, {"field1": {"$eq": "value2"}}]}
where_document: Optional document content filters
Examples:
- Contains: {"$contains": "value"}
- Not contains: {"$not_contains": "value"}
- Regex: {"$regex": "[a-z]+"}
- Not regex: {"$not_regex": "[a-z]+"}
- Logical AND: {"$and": [{"$contains": "value1"}, {"$not_regex": "[a-z]+"}]}
- Logical OR: {"$or": [{"$regex": "[a-z]+"}, {"$not_contains": "value2"}]}
include: List of what to include in response. By default, this will include documents, metadatas, and distances.
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes | ||
| include | No | ||
| n_results | No | ||
| query_texts | Yes | ||
| where | No | ||
| where_document | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears the full burden. It discloses default response fields (documents, metadatas, distances), supported filter operators with examples, and use of n_results for pagination. However, it lacks details on rate limits, maximum result limits, or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a well-structured docstring that front-loads the purpose and then systematically describes each parameter. While it includes many examples that increase length, these are necessary for complex parameters and are not redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 6 parameters, no output schema, and no annotations, the description covers all parameters with examples. It mentions the default include list but does not detail the response structure beyond field names. The tool's purpose is clear, but some behavioral details are omitted.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description adds substantial meaning to all 6 parameters, especially for complex fields where and where_document, providing detailed examples of supported operators and logical combinations. This goes far beyond the schema's type and default information.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action 'Query documents from a Chroma collection with advanced filtering,' with a specific verb and resource. It distinguishes from siblings like add, get, update, delete documents by focusing on querying with filtering.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for querying with advanced filtering but provides no explicit guidance on when to use this tool versus alternatives like get_documents or peek_collection. No when-not-to-use or alternative naming is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
chroma_update_documentsA
Update documents in a Chroma collection.
Args:
collection_name: Name of the collection to update documents in
ids: List of document IDs to update (required)
embeddings: Optional list of new embeddings for the documents.
Must match length of ids if provided.
metadatas: Optional list of new metadata dictionaries for the documents.
Must match length of ids if provided.
documents: Optional list of new text documents.
Must match length of ids if provided.
Returns:
A confirmation message indicating the number of documents updated.
Raises:
ValueError: If 'ids' is empty or if none of 'embeddings', 'metadatas',
or 'documents' are provided, or if the length of provided
update lists does not match the length of 'ids'.
Exception: If the collection does not exist or if the update operation fails.
| Name | Required | Description | Default |
|---|---|---|---|
| collection_name | Yes | ||
| documents | No | ||
| embeddings | No | ||
| ids | Yes | ||
| metadatas | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It states that the tool updates documents and raises ValueError if inputs are mismatched, but it doesn't disclose behavioral traits such as whether updates are full replacements or incremental, whether the operation is idempotent, or any permission requirements. It adequately describes error conditions but lacks broader behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with bullet points for args, returns, and raises. It is front-loaded with the core purpose and is fairly concise, though the 'Args:' header and bullet format could be slightly tighter without losing clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 5 parameters, no output schema, and no annotations, the description covers the input semantics, return type, and error conditions reasonably well. However, it lacks details on overall behavior like incremental vs full updates, idempotency, or prerequisites such as collection existence (though implied by exception). It is mostly complete for an update tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the description provides detailed explanations for each parameter, including constraints like 'Must match length of ids if provided.' This adds significant meaning beyond the schema's property titles and types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'Update documents in a Chroma collection' which is a specific verb+resource. It clearly distinguishes from sibling tools like add_documents (add new), delete_documents (delete), query_documents (query), etc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains what can be updated (embeddings, metadatas, documents) but does not provide explicit guidance on when to use this tool vs alternatives like add_documents (for new documents) or modify_collection (for collection settings). Usage is implied but not clearly contrasted.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
13 tool updates
v1.0.0- First observed
chroma_add_documents - First observed
chroma_create_collection - First observed
chroma_delete_collection - First observed
chroma_delete_documents - First observed
chroma_fork_collection - First observed
chroma_get_collection_count - First observed
chroma_get_collection_info - First observed
chroma_get_documents - First observed
chroma_list_collections - First observed
chroma_modify_collection - First observed
chroma_peek_collection - First observed
chroma_query_documents - First observed
chroma_update_documents
TDQS
Scored across 13 tools
Each tool has a clearly distinct purpose targeting specific Chroma operations. For example, chroma_get_documents retrieves documents with filtering, while chroma_query_documents performs semantic search; chroma_peek_collection provides a quick preview, and chroma_get_collection_info returns metadata. No tools appear to overlap in functionality.
All tools follow a consistent 'chroma_verb_noun' pattern with snake_case throughout. The naming convention is perfectly uniform, making it easy to predict tool names and understand their functions at a glance.
With 13 tools, this server provides comprehensive coverage for Chroma vector database operations without being overwhelming. The count aligns well with the domain scope, offering complete collection management and document CRUD operations.
The toolset provides complete coverage for Chroma operations: collection lifecycle (create, list, get info, modify, fork, delete), document lifecycle (add, get, update, delete, query), and utility functions (count, peek). No obvious gaps exist for typical vector database workflows.
Maintenance
Related MCP Connectors
The Needle MCP server enables semantic search on documents stored in files like PDFs, DOCX, and XLSX by connecting AI applications to external data sources. It provides capabilities to create and manage document collections, perform natural language searches on stored content, and retrieve relevant information without requiring exact keyword matches.
Ingest, manage, and retrieve documents for RAG-powered AI applications
The CustomGPT.ai MCP server is a fully managed, RAG-powered endpoint that connects large language models with private knowledge bases and external data sources. It provides tools for retrieval-augmented generation queries (send_message), data ingestion (upload_file), and source listing, enabling AI agents to query private documents like PDFs with high accuracy and real-time citations.
Remote ChromaDB vector database MCP server with streamable HTTP transport
Related MCP Servers
- AlicenseAqualityDmaintenanceA Model Context Protocol server providing vector database capabilities through Chroma, enabling semantic document search, metadata filtering, and document management with persistent storage.641MIT
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol server that enables intelligent document search and retrieval from PDF collections, providing semantic search capabilities powered by OpenAI embeddings and ChromaDB vector storage.13MIT
- AlicenseNot gradedqualityDmaintenanceAn MCP server that exposes ChromaDB vector database operations, enabling AI assistants to perform collection management and semantic document searches. It supports HTTP, persistent, and in-memory connection modes along with various embedding providers including OpenAI and HuggingFace.MIT
- FlicenseNot gradedqualityDmaintenanceA fully offline local RAG server that utilizes ChromaDB and Ollama to index and query PDF, text, and Markdown documents. It allows users to manage local knowledge bases and perform semantic searches with AI-generated responses.-