AgentRoots
AgentRoots MCP server lets you capture, review, retrieve, link, sync, and validate evidence-backed project knowledge as a bounded graph.
research_get_context: build a token-budgeted project continuity packet for a query.research_get_frontier: list unresolved candidate and provisional work at the project frontier.research_query: search current project records with FTS5 and fuzzy typo fallback.research_get_record: fetch one record with evidence, links, and revision history.research_propose: create an untrusted candidate record without executing stored text.research_review: apply lifecycle verdicts; creators cannot self-accept, and acceptance needs evidence.research_link_evidence: attach an external evidence URI, optional hash, summary, and metadata to a record.research_sync: import/export project events and optionally audit packet record usage.research_validate: check SQLite integrity and project governance invariants without mutation.
Provides a read-only adapter for MLflow, enabling agents to reference MLflow runs as evidence in research records.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@AgentRootsPropose a finding that caching cut latency 12%, linking the MLflow run."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
AgentRoots
Different agents. Same roots.
Shared, evidence-backed project knowledge across agents, sessions, and harnesses.
Built by Shanmukha Vellamcheti and Codex, an AI coding collaborator.
Your next agent should inherit the work, not repeat the investigation. AgentRoots keeps what you tried, what you know, why it matters, and what remains to do in one durable project graph. A fresh agent gets compact, relevant pointers and opens the evidence only when needed.
Past attempts. Present understanding. Future intent. One project's shared roots.

AgentRoots is a local-first, open-source MCP server for engineering, research, and long-running agent work. A smaller agent can investigate and propose findings; another agent can review and reuse them later. Optional, read-only conversation backfill also recovers leads that never made it into a handoff or source file. History stays untrusted until conclusions pass review.
No required model API, cloud account, or agent orchestrator. Works with one agent across sessions or several agents sharing a project. It helps reduce repeated exploration, not eliminate every file read or replace verification.
Start here
Requires Python 3.11+ and Git. Use a Python virtual environment if your OS manages its Python installation. Codex proactive hooks also require Node.js.
python -m pip install "agentroots @ git+https://github.com/shanmukha-here/agentroots.git@v0.2.0"
agentroots setupSetup detects supported clients, asks before changing their configuration, and separately asks
whether it may read existing conversations. You can decline history or select particular projects.
Codex requires an additional /hooks review. Downloads and approved backfill continue in the
background; lexical search works while the semantic model warms up. Setup is short, but installation
and large history imports depend on your connection and machine.
Then use your agents normally. On supported hook paths, AgentRoots recalls compact project context at prompt and tool boundaries and queues new findings for review. It stays quiet when there is nothing useful to add. Notification presentation depends on the host; not every client shows toasts.
agentroots status # backends, RAM, storage, and indexing progress
agentroots doctor # configuration and integration checks
agentroots project-current # the project this checkout resolves tov0.2.0 is an alpha, not a universal background-memory service. Codex has a validated MCP and proactive hook path. OpenCode automation is experimental. Claude and other stdio MCP clients can call tools, but do not get automatic hooks from this package. DeepSeek support depends on its harness, not its model name. See the support matrix.
No PyPI publication yet. Existing users: see the upgrade notes. For manual configuration and history selection, see integrations.
Related MCP server: elephantasm-mcp
See the knowledge, not just the conversation

The same versioned state agents query becomes an offline, interactive knowledge map for humans. Search, filter, trace relationships, inspect evidence, and copy record IDs to request corrections. The synthetic overview covers all 14 record types and 12 relationship types. It contains no private project content. Direct editing in the graph is planned; today corrections use the CLI or MCP.
agentroots graph PROJECT_ID /path/outside/repo/project-map.htmlWhy I built AgentRoots
My research involves exploring many hypotheses and experimental paths. Coding agents let me try more of them, faster. But as experiments, conversations, and subagents multiplied, keeping track of what we had already learned became its own problem.
An orchestrator might read a codebase to plan two tasks, then both subagents read the same files to get up to speed. That is the same understanding reconstructed three times. Weeks later, after compaction or a switch to another harness, an agent can suggest an experiment we already discussed and ruled out. The reason may exist only in an old conversation, not in the code or latest handoff.
I wanted that work to become durable without filling every new prompt with the entire past. AgentRoots connects the project's origin and intent, previous attempts, current evidence, and open questions. An agent can inherit that understanding, check what matters for its task, and add to it. Research motivated it, but the same problem appears in software projects and other long-running work.
This is not a claim that other memory tools only remember the past. Memory, planning, and provenance systems overlap. AgentRoots focuses on connecting them in a compact, reviewable project graph, without owning the agents themselves. See the landscape and boundaries.
How the roots grow
Capture leads. Agents propose records, or approved conversation history is indexed read-only. Background extraction creates candidates, never accepted facts.
Ground and review. Link evidence, inspect contradictions, and accept or reject a proposal. A creator cannot accept its own proposal by default. Changes append revisions and events.
Recall when relevant. MCP provides bounded context. Supported hooks also check emerging tool intent and results, not just the initial user prompt. They offer advice, not execution gates.
Keep state current. Revalidate Git and tracker evidence, mark changed findings stale, and preserve failed approaches. Accepted results can explicitly resolve goals or questions.
The durable boundary is the project, not the agent, model, harness, or conversation. Routing uses explicit metadata, Git remotes, and registered aliases. Conversation text cannot choose a different project. Bind a preferred project name once when needed:
agentroots project-bind PROJECT_ID /path/to/repositoryMatching project identities does not automatically synchronize separate machines. Share approved state through export/import or backup/restore; hosted team sync is not shipped.
Small context, inspectable evidence
The default MCP context response offers up to five matches within a conservative 200 estimated token budget, with short record IDs for opening details. Hook injections are capped at 180 estimated tokens. These are payload budgets, not guarantees about a host's wrapper tokens, total turn cost, or provider cache hits. Changing hook context is appended at supported event boundaries, not rewritten into a stable system-prompt prefix.
{"tool":"research_get_context","arguments":{"project":"my-project","query":"previous cache failures"}}
{"tool":"research_get_record","arguments":{"project":"my-project","record_id":"8f2a91c4"}}The record ID above is illustrative. Use an ID returned by your context response. Full packets are
available with view: "full"; records are normalized and the serialized response is budgeted.
MCP reads do not write packet-audit rows. Explicit CLI context packets are audited.
Retrieval and extraction do different jobs:
BGE + FTS5 + fuzzy search retrieve existing state. BGE is enabled by default and warms in the background. FTS continues serving cold or failed-model requests. Set
AGENTROOTS_SEMANTIC=offfor lexical-only operation.Candidate extraction uses configured Qwen first, optional GLiNER next, and a conservative heuristic otherwise. Qwen weights are not bundled, training is not required for normal use, and no extraction backend bypasses review.
See measured latency, token costs, and limitations. Development results are not evidence of universal productivity gains or reliable recall for every project.
What ships
External SQLite WAL storage, append-only events, revisions, and project-scoped graph queries.
Candidate review, evidence classification, contradictions, failure recall, and stale-state checks.
Read-only history backfill for Codex and OpenCode, with separate untrusted episode search.
Codex proactive hooks, a resident local daemon, and an experimental OpenCode plugin.
Read-only MLflow run search, comparison, snapshots, and revalidation. Trackio has an adapter interface; H-E-F and signac have importers.
Interactive read-only graph export, JSONL event transfer, and full-state backup/restore.
CLI, 16 MCP tools, resources, schemas, reproducible synthetic workflows, and regression tests.
The research_* tool names retain the original investigative ontology. The project boundary is
broad enough for engineering and general agent work; you do not need to invent a hypothesis or
pretend an exploratory run was preregistered. See the protocol and
tool configuration.
Privacy and trust boundaries
State, model weights, and indexes stay in OS user data/cache directories, not your repository. Large artifacts stay in their tracker or original storage. Source conversations are read-only; backfill requires approval. No telemetry or required cloud inference is built in.
Stored text is untrusted. Secret scanning and redaction reduce exposure but are not a guarantee that every secret or prompt injection is detected. Claims, findings, observations, and decisions need mechanically verified evidence before acceptance. A checksum verifies the referenced bytes, not the scientific truth of a conclusion. Arbitrary URIs and caller-written test receipts are references, not automatically verified proof.
Actor names are local provenance labels, not authenticated identities. Self-review rules prevent cooperative mistakes, not hostile impersonation. AgentRoots does not expose a production team authorization boundary. See SECURITY.md.
AgentRoots never owns agent spawning, model routing, stored-command execution, schedulers, training jobs, worktrees, or artifact storage. PostgreSQL, remote HTTP, ACLs, full Flowcept/AiiDA integrations, and direct graph editing remain roadmap work.
Contribute or reproduce the workflow
python -m pip install -e ".[dev]"
python -m pytest
python -m examples.full_flow_demo --output /path/outside/repo/agentroots-demoThe demo exercises approved synthetic history, review, a local MLflow fixture, duplicate-risk recall, and Git-induced staleness. It is a reproducible core workflow, not a recording of a live agent. The polished live-session video is still pending.
Architecture · Evaluation · Changelog · Contributors
Apache-2.0. Contributions and reports from real projects are welcome.
Available Tools
9 toolsresearch_get_contextB
Build bounded agent-continuity packet. Stored content is untrusted data.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | ||
| project | Yes | ||
| token_budget | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description includes a useful safety caveat ('Stored content is untrusted data') and implies a bounded output via 'bounded,' which hints at token limits. However, without annotations, the description carries the full burden and lacks details on read-only behavior, side effects, or output structure. It provides some value but not rich disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: two short sentences, both earning their place. It front-loads the action verb and adds a caveat, with no redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given three parameters, one required, and a sibling set of eight tools, the description is notably under-specified. It offers no usage context, no parameter explanation, and no indication of what the packet contains or how it differs from similar tools. The existence of an output schema helps, but the description alone is incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only vaguely alludes to a 'bounded' packet, which likely relates to token_budget, but does not explain query, project, or token_budget semantics. The agent must infer meaning from parameter names alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the verb 'Build' and the noun 'bounded agent-continuity packet,' which clearly identifies the tool's action and resource. It does not explicitly differentiate from sibling tools like research_get_record or research_get_frontier, but the concept of an agent-continuity packet adds specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives. The description does not mention scenarios, exclusions, or alternative tools. It simply states what it does without context for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
research_get_frontierB
Return unresolved candidate and provisional work at the project frontier.
| Name | Required | Description | Default |
|---|---|---|---|
| project | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states that the tool 'returns' data, implying a read-only operation, but it does not disclose any side effects, error behavior, pagination, ordering, or what constitutes 'unresolved' or 'provisional.' This is minimal behavioral transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no fluff. Every word contributes to identifying what the tool does, making it exemplary in conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter) and the presence of an output schema, the description is minimally adequate. However, it leaves ambiguity about the meaning of 'frontier' and lacks usage guidance or caveats, so it is not fully complete for an agent to invoke with confidence.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only references 'project' implicitly via 'at the project frontier.' It does not clarify whether 'project' is an ID, name, or path, nor does it define accepted formats. The description adds only marginal meaning beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Return' and identifies a clear resource: 'unresolved candidate and provisional work at the project frontier.' This distinguishes it from sibling tools like research_get_context and research_get_record, though it does not explicitly name alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'at the project frontier' implies when to use this tool (retrieving pending work), but it provides no explicit guidance on when not to use it or which alternative might be better for other scenarios. Thus, usage is implied but not explicitly directed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
research_get_recordA
Return one current record with evidence, links, and revision history.
| Name | Required | Description | Default |
|---|---|---|---|
| record_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It adds context by specifying 'current record' (implying latest version) and enumerates return contents, but it does not disclose behavior for missing IDs, error cases, or side effects. Some transparency is present, but it is limited.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that directly states the tool's purpose and scope. It contains no fluff and is appropriately front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with one parameter and an output schema. The description adequately conveys the core behavior (returning one current record with certain contents) and benefits from the existence of an output schema. It omits edge cases or usage context, but given the low complexity, this is acceptable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention the record_id parameter or explain how to identify the record. The parameter name is self-explanatory, but the description adds no additional semantics or usage guidance, failing to compensate for the lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Return') and resource ('one current record') and lists included content (evidence, links, revision history). This clearly distinguishes it from sibling tools like research_get_context or research_query.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. There is no mention of when to choose this over research_query or research_get_context, nor any exclusions or alternative tool names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
research_link_evidenceB
Attach an external evidence URI, optional hash, and compact summary to a record.
| Name | Required | Description | Default |
|---|---|---|---|
| uri | Yes | ||
| kind | Yes | ||
| actor | Yes | ||
| summary | No | ||
| metadata | No | ||
| record_id | Yes | ||
| content_hash | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description must disclose side effects and constraints. It only states that evidence is attached, but doesn't say whether the record must already exist, whether attachments are appended or replaced, or what response/errors to expect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler or repetition. It communicates the core action efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with seven parameters and no annotations, one sentence is insufficient. It lacks parameter semantics for kind/actor, usage context, and behavioral details like idempotency or validation, even though an output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description needed to explain all seven parameters. It adds meaning for uri, content_hash ('optional hash'), and summary ('compact summary'), but leaves required fields kind and actor undefined, and record_id is only implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Attach') and names the resource ('external evidence URI...to a record'), which clearly distinguishes it from the read/query/sync siblings. It is unambiguous about the tool's core action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives such as research_validate or research_propose. There is no mention of prerequisites, exclusions, or typical scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
research_proposeA
Create an untrusted candidate record. This never executes stored text.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | ||
| mode | No | exploratory | |
| title | Yes | ||
| creator | Yes | ||
| project | Yes | ||
| metadata | No | ||
| record_type | Yes | ||
| idempotency_key | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses a critical safety behavior: 'This never executes stored text.' This is valuable beyond the schema and addresses a common security concern for storing untrusted content. However, it doesn't mention other behaviors like idempotency, permissions, or return details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences, front-loading the core purpose and adding the safety note. It is concise and appropriately sized for a simple creation tool, with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 8 parameters, 5 required, and no parameter explanations, the description is incomplete for a tool of this complexity. It gives the core purpose and a safety guarantee, but lacks context on how to use the parameters, what 'untrusted' implies for processing, or when to set mode. The presence of an output schema helps slightly, but the description alone is not enough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description provides no explanation of any of the 8 parameters (e.g., mode, metadata, idempotency_key). The parameter names are self-explanatory to some degree, but the description adds no value for understanding how to fill them correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Create') and the object ('an untrusted candidate record'), which is a specific resource type. This distinguishes it from sibling tools like get_record, query, and sync, which are read/validate operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for creating a new candidate record, but it provides no explicit guidance on when to use it versus alternatives, nor any exclusions. The distinction from siblings is implicit based on the verb 'create'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
research_queryB
Search current project records with FTS5 and fuzzy typo fallback.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | No | ||
| project | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full burden for behavioral disclosure. It discloses the use of FTS5 and fuzzy typo fallback, which is helpful context, but it does not mention read-only nature, pagination, or error behavior, which are relevant for a query tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that immediately states the core action and key differentiator (FTS5/fuzzy fallback). It is concise and front-loaded with the most important information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The presence of an output schema removes the need to describe return values, but the description is still thin for a 3-parameter search tool. It lacks detail on query syntax, fuzzy fallback behavior, or result limiting, though the schema covers parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has no parameter descriptions (0% coverage), and the description only indirectly references 'query' and 'project' through 'Search current project records'. The 'limit' parameter is completely unexplained, so the description does not compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'search' and identifies the resource as 'current project records'. It differentiates from sibling tools like research_get_record by mentioning FTS5 and fuzzy typo fallback, making its purpose clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. The description only states what it does, with no mention of exclusions, prerequisites, or alternative tools, leaving the AI to infer usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
research_reviewC
Apply lifecycle review. Creators cannot self-accept; acceptance needs evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| actor | Yes | ||
| comment | No | ||
| verdict | Yes | ||
| record_id | Yes | ||
| expected_revision | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility for disclosing behavior. It reveals a key constraint (creators cannot self-accept; acceptance needs evidence) but does not state whether the tool mutates state, requires specific permissions, or what side effects occur. This leaves critical behavioral traits undisclosed for a review action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no redundant wording. The first sentence states the primary action, and the second adds a valuable constraint. It is concise and front-loaded, though the first sentence could have been more specific without losing brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 5 parameters, an output schema, and is part of a family of research tools with distinct workflow roles. The minimal description does not explain how this review fits into the lifecycle, what the parameters mean, or what the output schema represents, leaving significant gaps for an agent to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining the 5 parameters. It only implicitly touches on actor and verdict via the self-accept rule, but record_id, comment, and expected_revision are not explained. The description fails to add meaning beyond the raw schema for most parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the action ('Apply lifecycle review') and a resource (review), and adds a specific rule about self-acceptance and evidence. However, it does not explicitly distinguish from sibling tools like research_validate, and 'lifecycle review' is somewhat jargon-heavy without further context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no explicit guidance on when to use this tool versus alternatives such as research_validate or research_propose. The rule about creators and evidence is a policy constraint, not a usage direction, and there is no mention of when not to use the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
research_syncC
Import events, export project events, and optionally audit packet records used.
| Name | Required | Description | Default |
|---|---|---|---|
| events | No | ||
| project | Yes | ||
| packet_id | No | ||
| used_record_ids | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of disclosing side effects. It states actions like 'import' and 'export' but does not clarify whether these modify data, require specific permissions, or have other behavioral implications. The phrase 'optionally audit' is undefined, leaving the tool's behavior largely opaque.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single short sentence, which is concise, but its structure is disjointed—listing three actions without clear separation or emphasis. The phrase 'and optionally' adds confusion about whether audit is a separate mode or a modifier. It is not well-organized given the multi-part functionality.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists, the description does not explain how the tool's multiple modes are invoked or what inputs are required for each. For a tool with two primary directions (import/export) and an audit option, the description omits critical context about parameter combinations and expected outcomes, leaving it incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the four undocumented parameters. It loosely maps to parameters (events, project, packet_id, used_record_ids) but does not explain their formats, defaults, or how they control the import/export/audit modes. This is insufficient for a 4-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description lists three actions: import events, export project events, and optionally audit packet records. While each action has a verb and resource, the overall purpose is vague because it's unclear how these actions relate or when each is intended to be performed. It does not clearly distinguish itself from sibling tools beyond the name 'sync'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives. The description does not mention any prerequisites, exclusions, or comparisons to sibling tools like research_query or research_validate. Users are left to infer usage context on their own.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
research_validateA
Check SQLite integrity and project governance invariants without mutation.
| Name | Required | Description | Default |
|---|---|---|---|
| project | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description explicitly states 'without mutation', which is a key behavioral trait, but with no annotations provided, it carries full burden. It does not disclose error behavior, prerequisites, or potential side effects like locking or resource usage. This provides some value but lacks depth.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, succinct sentence that is front-loaded with the core action. Every word contributes to the meaning, with no filler or redundant content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has a simple interface with one parameter and an output schema, so return values are covered. However, the description lacks guidance on the 'project' parameter and when to use validation, making it minimally adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema only defines a single 'project' string parameter with no description, and the tool description does not explain what 'project' refers to or how it should be formatted. With 0% schema description coverage, the description should compensate but fails to add meaning to the parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks SQLite integrity and project governance invariants, using the specific verb 'Check' and identifying a distinct resource. This distinguishes it from sibling tools like research_query or research_get_record, which are for retrieval or mutation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used for validation tasks, but it does not explicitly state when to use it over alternatives or provide exclusion criteria. There is no mention of related tools or conditions, leaving usage guidance only implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.1.0- First observed
research_get_context - First observed
research_get_frontier - First observed
research_get_record - First observed
research_link_evidence - First observed
research_propose - First observed
research_query - First observed
research_review - First observed
research_sync - First observed
research_validate
TDQS
Scored across 9 tools
Each tool has a distinct purpose: context retrieval, frontier scanning, querying, record fetching, validation, proposal creation, lifecycle review, evidence linking, and synchronization. There is no overlap or ambiguity between tool functions.
All tools consistently use the 'research_' prefix with a clear verb_noun pattern (e.g., get_context, link_evidence, validate). The naming is uniform, predictable, and follows a standard convention.
Nine tools is well-scoped for a research project management server, covering the core lifecycle without redundancies. The number falls comfortably within the ideal range and each tool serves a necessary function.
The toolset covers creation (propose), lifecycle management (review), evidence linking, querying, validation, and sync, which addresses the main workflows. However, there is no explicit update or delete tool, so direct modification of record content is not exposed; the review and propose tools may handle this indirectly via versioning.
Maintenance
Related MCP Connectors
- WauldoOAuthcom.wauldo
Stateless agentic tools over MCP: concept extraction, long-context, knowledge graph, planning.
Graph-native persistent memory for AI agents — 33 MCP tools, zero-LLM writes.
Persistent memory for AI agents to retain, retrieve, and recall conversation context through MCP.
Long-term memory for AI coding agents: durable project facts, recalled by every MCP client.
Related MCP Servers
- AlicenseBqualityBmaintenanceMCP server providing persistent engineering memory and spec-driven development workflows for AI coding agents, preserving learnings across sessions.413,058 PyPIBusiness Source 1.1
- AlicenseAqualityCmaintenanceMCP server for long-term agent memory, providing persistent memory, searchable knowledge, and evolving identity for AI agents.53Apache 2.0
- FlicenseNot gradedqualityDmaintenanceShared memory and orchestration for coding agents, enabling persistent knowledge, multi-agent coordination, and a canonical workflow across MCP-compatible AI clients.19 npm110-
- AlicenseNot gradedqualityCmaintenanceMCP server that provides AI agents with persistent memory, cross-agent sharing, and context management, enabling them to remember conversations, track complex tasks, and evolve skills across tools.2MIT