Skip to main content
Glama
Bristlecone2026

Bristlecone Logic Utilities

chunk_text

Read-onlyIdempotent

Split raw text into fixed-size segments with configurable overlap, preparing documents for vector embeddings and RAG pipelines.

Instructions

Partitions raw text documents into uniform sliding-window segments with configurable character overlap. Returns an array of formatted text chunks. Use when preparing unstructured documents for vector database embeddings and RAG retrieval pipelines. Do not use for syntactic token counting or semantic sentence segmentation.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
textYesThe source document text string to segment into discrete chunks.
chunk_sizeNoMaximum character length of each individual chunk segment. Defaults to 500 characters.
chunk_overlapNoNumber of overlapping characters shared between consecutive chunks to maintain semantic context. Defaults to 50 characters.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv0.4.3
    • changedOutput schema / (root)
      Previous value: -{
      -  "properties": {
      -    "chunks": {
      -      "items": {
      -        "type": "string"
      -      },
      -      "type": "array"
      -    },
      -    "total_chunks": {
      -      "type": "integer"
      -    }
      -  },
      -  "required": [
      -    "total_chunks",
      -    "chunks"
      -  ],
      -  "type": "object"
      -}New value: +null
  2. First observedv0.3.0

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds genuine context beyond that: the sliding-window/overlap semantics and the return shape ('array of formatted text chunks'), which matters because no output schema exists. It does not mention edge cases such as behavior when text is shorter than chunk_size, which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences: capability, return value, then usage/anti-usage. It is front-loaded, free of filler, and every sentence contributes distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully states that an array of formatted text chunks is returned, and annotations cover the read-only safety profile. The primary use case and exclusions are complete; only minor operational details (e.g., what 'formatted' means, empty/oversized input handling) are absent, which is acceptable for a three-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents text, chunk_size (default 500, min 50), and chunk_overlap (default 50, min 0) in detail. The description only echoes 'configurable character overlap' and adds no syntax, constraints, or interaction rules (e.g., overlap < chunk_size) beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Partitions raw text documents into uniform sliding-window segments') with the key mechanism (configurable character overlap) named up front. An agent can distinguish it from siblings like repair_json or validate_schema immediately, since none of them concern text segmentation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit positive trigger ('Use when preparing unstructured documents for vector database embeddings and RAG retrieval pipelines') and negative guidance ('Do not use for syntactic token counting or semantic sentence segmentation'). Both when-to-use and when-not-to-use are spelled out, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.