Skip to main content
Glama
vbcherepanov

total-agent-memory

memory_recall_iterative

Read-onlyIdempotent

Break multi-hop queries into sub-questions, retrieve evidence for each, and use a planner LLM to decide if further retrieval is needed. Returns unified evidence with provenance per iteration.

Instructions

v11.0 W1-B: IRCoT-style iterative retrieval. Decomposes the query into sub-questions, retrieves per sub-question, and asks a planner LLM whether more retrieval is needed. Best for multi-hop questions. Returns unified evidence + provenance per iteration.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
queryYes
projectNo
llm_modelNoconfigured
max_itersNo
k_per_iterNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed6 schema fields changedv14.0.0
    • changedInput schema / properties / k_per_iter / default
      Previous value: -5New value: +10
    • addedInput schema / properties / k_per_iter / maximum
      Added value: +50
    • addedInput schema / properties / k_per_iter / minimum
      Added value: +1
    • changedInput schema / properties / llm_model / default
      Previous value: -"haiku"New value: +"configured"
    • addedInput schema / properties / max_iters / maximum
      Added value: +12
    • addedInput schema / properties / max_iters / minimum
      Added value: +1
  2. First observedv0.1.0

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, covering safety. The description adds useful behavioral detail about the internal process (decomposing, retrieving, asking a planner) and the return format ('unified evidence + provenance per iteration'). This goes beyond annotations without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, with the main purpose stated up front. The version prefix 'v11.0 W1-B' is extraneous and could be removed, but it does not impair understanding. The core sentences are efficient and informative, though the parameter details are missing, which slightly reduces the structure's completeness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with 5 parameters and no output schema, the description is incomplete. It explains the high-level algorithm but does not describe how max_iters or k_per_iter affect behavior, what the project or llm_model parameters do, or the exact return structure beyond a vague 'unified evidence + provenance per iteration.' An agent cannot confidently set parameters or parse output without additional information.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain any of the 5 parameters (query, project, llm_model, max_iters, k_per_iter). It mentions retrieval and iteration but does not define how these parameters control behavior. Since the description carries the full burden for parameter semantics and fails to address them, the score is minimal.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: IRCoT-style iterative retrieval that decomposes queries into sub-questions, retrieves per sub-question, and uses a planner LLM. It explicitly notes it is 'Best for multi-hop questions,' which differentiates it from the simpler sibling 'memory_recall'. This gives a specific verb, resource, and scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear usage context by stating 'Best for multi-hop questions,' which signals when to use this tool over alternatives like memory_recall. It implies that for simpler, single-hop queries, other tools are preferable, but it does not explicitly name them or provide exclusions, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.