Skip to main content
Glama
chroma-core

Chroma MCP Server

Official
by chroma-core

chroma_add_documents

Enables adding text documents to a Chroma collection with optional metadata. Specify collection name, document IDs, and content for efficient data storage and retrieval.

Instructions

Add documents to a Chroma collection.

Args:
    collection_name: Name of the collection to add documents to
    documents: List of text documents to add
    ids: List of IDs for the documents (required)
    metadatas: Optional list of metadata dictionaries for each document

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
collection_nameYes
documentsYes
idsYes
metadatasNo

Implementation Reference

  • The chroma_add_documents tool handler. Decorated with @mcp.tool() for automatic registration in FastMCP. Implements document addition to ChromaDB collection with validation for empty inputs, ID uniqueness, and length matching. Uses get_chroma_client() helper and performs add operation.
    @mcp.tool()
    async def chroma_add_documents(
        collection_name: str,
        documents: List[str],
        ids: List[str],
        metadatas: List[Dict] | None = None
    ) -> str:
        """Add documents to a Chroma collection.
        
        Args:
            collection_name: Name of the collection to add documents to
            documents: List of text documents to add
            ids: List of IDs for the documents (required)
            metadatas: Optional list of metadata dictionaries for each document
        """
        if not documents:
            raise ValueError("The 'documents' list cannot be empty.")
        
        if not ids:
            raise ValueError("The 'ids' list is required and cannot be empty.")
        
        # Check if there are empty strings in the ids list
        if any(not id.strip() for id in ids):
            raise ValueError("IDs cannot be empty strings.")
        
        if len(ids) != len(documents):
            raise ValueError(f"Number of ids ({len(ids)}) must match number of documents ({len(documents)}).")
    
        client = get_chroma_client()
        try:
            collection = client.get_or_create_collection(collection_name)
            
            # Check for duplicate IDs
            existing_ids = collection.get(include=[])["ids"]
            duplicate_ids = [id for id in ids if id in existing_ids]
            
            if duplicate_ids:
                raise ValueError(
                    f"The following IDs already exist in collection '{collection_name}': {duplicate_ids}. "
                    f"Use 'chroma_update_documents' to update existing documents."
                )
            
            result = collection.add(
                documents=documents,
                metadatas=metadatas,
                ids=ids
            )
            
            # Check the return value
            if result and isinstance(result, dict):
                # If the return value is a dictionary, it may contain success information
                if 'success' in result and not result['success']:
                    raise Exception(f"Failed to add documents: {result.get('error', 'Unknown error')}")
                
                # If the return value contains the actual number added
                if 'count' in result:
                    return f"Successfully added {result['count']} documents to collection {collection_name}"
            
            # Default return
            return f"Successfully added {len(documents)} documents to collection {collection_name}, result is {result}"
        except Exception as e:
            raise Exception(f"Failed to add documents to collection '{collection_name}': {str(e)}") from e
  • The @mcp.tool() decorator registers the chroma_add_documents function as an MCP tool.
    @mcp.tool()
  • Input schema defined by function type hints: collection_name (str), documents (List[str]), ids (List[str]), metadatas (optional List[Dict]). Returns str confirmation.
    async def chroma_add_documents(
        collection_name: str,
        documents: List[str],
        ids: List[str],
        metadatas: List[Dict] | None = None
    ) -> str:

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.0.0

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description lacks any behavioral disclosure beyond the basic action. It does not mention whether duplicate IDs cause errors or overwrite, whether documents are validated for size or format, or what the return value indicates. With no annotations provided, the description carries the full burden and fails to address key behavioral traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is structured with an Args block, which is clear, but it largely restates the schema information. While not excessively long, it could be more concise by omitting redundant parameter descriptions and focusing on unique behavioral details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool mutates data and has no output schema or annotations, the description is incomplete. It does not explain the return value, constraints (e.g., document count limits), or side effects of adding documents to an existing collection. The agent lacks enough context for reliable use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description provides basic semantics for each parameter (e.g., 'ids: List of IDs for the documents (required)'). This adds moderate value beyond the parameter names but lacks depth, such as uniqueness constraints for 'ids' or format expectations for 'documents'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool adds documents to a Chroma collection, using the verb 'add' and specifying the resource 'documents to a Chroma collection'. However, it does not differentiate from sibling tools like 'chroma_update_documents' or 'chroma_delete_documents', which share similar contexts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. For instance, it does not state that the collection must exist before adding, nor does it explain when to use 'add' over 'update' or 'delete'. The agent is left without context for decision-making.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.