Skip to main content
Glama

ingest_doc

Ingest, crawl, parse, and index documentation from URLs or raw content into DocOrbit; optionally generate an evidence-grounded implementation recipe for a coding task.

Instructions

Ingest, crawl, parse, and index authoritative documentation from any URL or raw content directly into DocOrbit. Tracks documentation sources deterministically in docs.lock. Extracts semantic chunks, OpenAPI endpoints, code examples, and pitfalls. If taskContext is provided, immediately synthesizes and returns an evidence-grounded implementation recipe with exact code and API details.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlNoDocumentation target URL to crawl and ingest (e.g. "https://nextjs.org/docs" or "https://support.atlassian.com/...").
forceNoForce re-fetching and re-crawling documentation even if the source is already tracked (default: false).
titleNoOptional title when ingesting raw content or overriding page title.
formatNoResponse format: "markdown" (default, human/agent-readable documentation) or "json" (structured raw machine data).
contentNoOptional raw markdown/HTML documentation content to index directly without fetching from the web.
refreshNoAlias for force.
maxPagesNoMaximum number of pages to crawl (default: 20, max: 50).
taskContextNoOptional coding task or intent (e.g. "Connect Atlassian Remote MCP" or "Implement Stripe payment element"). If provided, DocOrbit compiles and returns an immediate implementation recipe using the newly ingested docs.
allowLocalhostNoAllow crawling localhost endpoints for testing (default: false).

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed3 schema fields changedv0.2.3
    • addedInput schema / properties / force
      Added value: +{
      +  "description": "Force re-fetching and re-crawling documentation even if the source is already tracked (default: false).",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / format
      Added value: +{
      +  "description": "Response format: \"markdown\" (default, human/agent-readable documentation) or \"json\" (structured raw machine data).",
      +  "enum": [
      +    "markdown",
      +    "json"
      +  ],
      +  "type": "string"
      +}
    • addedInput schema / properties / refresh
      Added value: +{
      +  "description": "Alias for force.",
      +  "type": "boolean"
      +}
  2. First observedv0.1.4

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does meaningfully disclose behavior: sources are tracked deterministically in docs.lock, semantic chunks/OpenAPI endpoints/code examples/pitfalls are extracted, and providing taskContext triggers an immediate evidence-grounded recipe. It still omits network side effects of crawling external sites, permissions, rate limits, and idempotency of re-ingestion.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, front-loaded with the core verb/resource, then behavior, then the conditional return. Dense but each sentence contributes; the taskContext payoff is placed last as a secondary mode rather than buried mid-sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a nine-parameter ingestion tool with no annotations and no output schema, the description adequately conveys what happens (crawl, index, track) and what the taskContext path returns. It is slightly thin on persistent side effects and error/permission behavior, but the core mental model is complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all nine parameters are already documented and the baseline is 3. The description adds only the dual-mode framing (URL fetch vs raw content ingestion) and reinforces the taskContext conditional, without adding format, limit, or aliasing details beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a precise set of verbs (ingest, crawl, parse, index) and a concrete resource (documentation from URL or raw content into DocOrbit). It is unambiguously the lone ingestion/write tool among siblings like search_docs, get_doc, and list_sources, so an agent can route to it without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied by the nature of the tool; the description never states when to prefer this over search_docs/get_doc/find_* siblings, nor any prerequisites or when-not-to-use conditions. It does surface one conditional behavior (supply taskContext to trigger recipe synthesis), which is closer to behavior than routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.