mcp-metsuke-crunchtools
OfficialMetsuke is a stateful MCP server for defining, scheduling, gathering, storing, and retrieving cited report outputs.
Create and manage report definitions: list them, fetch a gather spec, and upsert name, prompt, owner agent, cron schedule, timezone, and source config.
Trigger gathers on demand or let the built-in cron scheduler fire them, opening durable runs with run IDs and per-report concurrency locks.
Save gathered outputs as structured findings, each with summary, optional source URL, section/theme, actors, dates, and tracker references.
Read the freshest output or a specific gathered date, and browse run history with metadata and finding counts.
Delete one output by ID or bulk-prune outputs by keeping the newest N or deleting before a date.
Read deterministic pre-gathered sweep results by run and section, paging through compact records from Slack, Gmail, calendar, feeds, or Jira collectors.
Store everything locally in SQLite (WAL), with no external services, and serve over stdio, SSE, or streamable HTTP.
Gathers inbox threads via the [gmail](/mcp/servers/integrations/gmail)_waiting sweep collector — pulls messages within a reporting window from a configured mailbox/backend, applies ownership analysis to surface only threads where the ball is in the user's court, drops automated mail and bare calendar notices while keeping invitations, and records each thread with a Gmail deep link for citation.
Gathers recent feed entries via the feed_entries sweep collector, reading a configurable set of feed categories (with per-category limits and a widened lookback after weekends) and storing the entries as compact, citable findings for the gatherer to select and phrase.
Gathers Slack DMs and @-mentions via the slack_waiting sweep collector — reads up to three pages per thread or DM, drops answered conversations, marks the rest waiting or acknowledged (reaction-only), flags unverified/long threads, and separates new in-window asks from still-open ones, with a configurable workspace URL.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mcp-metsuke-crunchtoolspull the freshest output for the 'weekly vuln scan' report"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-metsuke-crunchtools
Stateful reports catalog MCP server. Metsuke (目付 — the Sengoku intelligence officer who gathered field reports and compiled them for the daimyō) is the durable, cross-agent home for report definitions and their gathered outputs.
An autonomous gatherer reads a definition, sweeps the configured sources, and writes findings — each carrying its own source URL. A compiler later reads the freshest output to draft a fully cited report. This replaces ad-hoc, run-scoped research caches with a real, queryable store.
Features
Definitions + outputs — separate what-to-gather from what-was-gathered
Built-in scheduler — Metsuke fires each definition's gather when its cron schedule comes due, or on demand via
trigger_report— no external timerRun lifecycle — every fire opens a durable run with a second-granularity
run_id: a provisional row is recorded before the callback (a fire is never lost), and a per-report concurrency lock keeps two gathers from racingCited findings — payloads carry per-finding source URLs for one-click checking
Re-homeable gathering —
owner_agentmakes the gatherer identity data, not codeLocal-first — plain SQLite (WAL), no external services, no per-seat fees
Three transports — stdio, SSE, streamable-http
Related MCP server: Argon Memory
Install
uvx (recommended)
uvx mcp-metsuke-crunchtoolspip
pip install mcp-metsuke-crunchtoolsContainer
podman run --rm -v ~/.local/share/mcp-metsuke:/data:Z \
quay.io/crunchtools/mcp-metsuke \
--transport streamable-http --host 0.0.0.0 --port 8009Claude Code Integration
claude mcp add mcp-metsuke-crunchtools -- uvx mcp-metsuke-crunchtoolsTools (10)
Definitions (4)
Tool | Description |
| List all report definitions in the catalog, each with its next scheduled fire time. |
| Return the gather prompt + source config for one definition. |
| Create or update a definition (name, prompt, owner, cron schedule, timezone, sources). |
| Fire a report gather right now, without waiting for its schedule. For a swept definition it returns at once with |
Outputs (6)
Tool | Description |
| Complete a run by persisting its gathered findings (pass the |
| Read the freshest output (or a specific day's) for compiling a report. |
| Browse the run history — metadata + |
| Delete one saved output by id. |
| Bulk-prune a report's outputs — keep the N newest, or drop those before a date. |
| Read a swept run's pre-gathered records: the index, or one page of one section. |
Environment Variables
Variable | Default | Description |
|
| SQLite database path |
| (none) | Path whose contents override |
| (none) | Base URL of the Trentina gateway; the scheduler POSTs gather callbacks to |
| (none) | Alert token identifying the reports profile; enables the scheduler when set with |
| (none) | Path whose contents override |
| (none) | Trentina gateway MCP endpoint for the sweep profile, e.g. |
| (none) | Bearer token for that profile; |
|
| Upper bound on one run's sweep |
|
| How often the scheduler checks for due reports |
| (auto) | Force the scheduler on/off; defaults to on when the callback is configured |
|
| How long an in-flight run holds the per-report lock before it is expired (self-heals a dead gatherer) |
The scheduler runs only under the sse and streamable-http transports (the long-lived production processes), never under stdio.
Data Model
report_definitions —
name(PK),gather_prompt,owner_agent,schedule(cron),timezone(IANA),source_config(JSON),last_fired_at,updated_atreport_outputs (one row per run) —
id,report_name(FK),run_id(second-granularity identity),trigger(scheduled/manual/direct),gathered_at,finished_at,window_start/window_end,payload(JSON findings with source URLs),status(gathering/ready/compiled/failed),gatherer_run_ref,detail. A partial unique index on(report_name) WHERE status='gathering'is the per-report concurrency lock.
MCP Registry
io.github.crunchtools/metsuke
License
AGPL-3.0-or-later
Sweeps
A definition can opt into a deterministic pre-gather by adding sweep to its
source_config:
{"sweep": {"timezone": "America/New_York", "window_hour": 6, "steps": [
{"section": "slack", "collector": "slack_waiting",
"options": {"user_id": "U9VN3S1ST", "handle": "smccarty"}},
{"section": "work_email", "collector": "gmail_waiting",
"options": {"backend": "gw-work", "account": "smccarty@redhat.com"}},
{"section": "calendar", "collector": "calendar_day",
"options": {"backend": "gw-work", "account": "smccarty@redhat.com"}},
{"section": "rss", "collector": "feed_entries",
"options": {"categories": {"1": 20, "5": 20}}}]}}On fire, Metsuke runs each step in order through the Trentina gateway (as its
own read-only profile), stores the results on the run, then dispatches the
callback with sweep_status. The gatherer reads records with get_sweep
instead of calling the sources, so the gathering LLM only selects, phrases and
delivers. A failing step is recorded and the sweep continues. Results Trentina
flags are stored as metadata only, with their text withheld.
Collector | What it produces |
| Direct asks over a lookback: an @-mention, or a DM message that reads as a question or request (DM chatter and Slack system notices are not asks). Threads are read as threads; DMs and top-level channel mentions are read from channel history (one conversation per channel). Each is read for up to three pages (longer or partly unreadable ones are marked |
| Inbox threads in the window (or |
| The report day's meetings, pending invites over a lookahead, and hard overlaps |
| Recent entries per feed category (read or unread), with a longer window after a weekend |
| Issues matching each named JQL query, with |
Collector options
Every step is {"section": "<name>", "collector": "<collector>", "options": {...}}.
Section names are [a-z0-9_-], unique, up to 12 steps. Top-level timezone
(IANA, default America/New_York) and window_hour (0-23, default 6) set the
window: from that hour on the previous weekday until the run.
Collector | Option | Required | Default | Bounds |
|
| yes | 2-32 chars | |
| yes | 1-64 chars | ||
|
| ≤64 chars | ||
|
| ≤64 chars | ||
| none | 1-64 chars; in group DMs and channels, a question or request that addresses the user by this name is an ask | ||
| 7 | 1-30 | ||
| 40 | 1-100 | ||
|
| ≤200 chars | ||
|
| yes | e.g. | |
| yes | the mailbox address | ||
|
| ≤500 chars, appended to | ||
| 60 | 1-200 | ||
| 1200 | 0-4000 (0 = no body) | ||
| Gmail | ≤200 chars, or null | ||
| report window | 1-30, search the last N days instead | ||
| 0 | 0-168, drop threads active more recently | ||
|
| ≤20 From-header substrings (1-100 chars, case-insensitive) | ||
|
| yes | ||
| yes | |||
|
| IANA zone | ||
| 3 | 0-14 (pending invites) | ||
|
| yes | 1-12 entries of | |
|
| |||
| 1 | 1-30 | ||
| 3 | 1-30 (Mondays and weekend runs) | ||
|
| yes | 1-6 of | |
|
| ≤64 chars; needs | ||
|
| ≤200 chars, prefixed to the issue key | ||
| 800 | 0-2000 (0 = no form detail line) |
Collectors are code in sweep/collectors.py, not configuration: a definition
can only pick and parameterize them. A bad spec is rejected at
upsert_definition time.
Available Tools
10 toolsdelete_output_toolA
Delete one saved report output by id.
Returns the deleted output's metadata (with deleted=True). Find ids with list_outputs_tool. Raises if no output carries that id.
| Name | Required | Description | Default |
|---|---|---|---|
| output_id | Yes | The id of the output to delete (from list_outputs_tool) |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does well: it discloses the return shape (deleted metadata with deleted=True) and the failure mode (raises if no output has that id). It does not explicitly call out irreversibility or auth requirements, but the destructive nature and error contract are clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action, then return value, then error behavior. No wasted words and the id-discovery hint arrives exactly where needed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter mutation with an output schema, everything an agent needs is present: the action, how to source the id, the success return, and the failure condition. Nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter already documents its type and provenance (from list_outputs_tool). The description only restates 'by id', adding no syntax or format detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Delete), resource (saved report output), and scope (one, by id), which cleanly distinguishes it from list_outputs_tool and prune_outputs_tool. An agent knows exactly what operation this performs without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Directs the agent to list_outputs_tool for obtaining ids, giving a clear prerequisite workflow. It does not explicitly mention the bulk-delete alternative (prune_outputs_tool) or state when-not to use this tool, so it stops short of full sibling routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_output_toolA
Read a gathered report output for compiling a report.
Returns the most recent output by default, or the latest output gathered on a specific date when gathered_date is given.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The report definition name | |
| gathered_date | No | Optional YYYY-MM-DD to fetch that day's output instead |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It discloses the default behavior (most recent output), the date-specific behavior, and the fallback semantics of 'latest output gathered on a specific date.' This is meaningful behavioral context beyond what the schema expresses, though it does not cover edge cases like missing outputs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the main purpose appears in the first sentence, and the only additional sentence explains the optional parameter behavior. Every sentence earns its place with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity read tool, the description is complete: it states the resource, the default behavior, the optional date behavior, and the parameters are fully covered by the schema. An output schema exists, so return-value details are already provided elsewhere, and no critical invocation detail is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Both parameters are already well described in the schema, so the baseline is 3. The description adds context by explaining the default selection behavior and how gathered_date changes the result, but it does not significantly expand on the parameter meanings beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Read') and resource ('gathered report output'), and it adds the purpose ('for compiling a report'). It also distinguishes this tool from siblings by clarifying it retrieves outputs, whereas list_reports_tool, get_spec_tool, and save_output_tool imply different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage context is implied by the phrase 'for compiling a report' and by the optional date parameter, but there is no explicit guidance on when to choose this tool over alternatives. It does not describe exclusions or mention sibling tools, so the agent must infer the appropriate moment to call it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_spec_toolA
Return the gather spec (prompt + source config) for a report definition.
This is what the autonomous gatherer calls on callback to learn what to collect for the named report.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The report definition name (e.g. "core-platform-status") |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It clarifies what the return payload contains (prompt + source config) and the invocation context, but it does not explicitly state that the operation is read-only, what happens for unknown report names, or any auth requirements. 'Return' implies a safe read, but it is not explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The first sentence front-loads the core action and resource; the second adds valuable contextual information about the caller. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read tool with an existing output schema, the description provides enough to invoke it correctly: what it returns and when to use it. Minor gaps such as not-found behavior or explicit read-only assurance are acceptable for a simple getter, but would make it more complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides 100% coverage for the single 'name' parameter, including an example. The description adds only a slight restatement ('named report') and does not provide additional semantic depth such as naming conventions, validation rules, or behavior for missing names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Return' with a clear resource: 'gather spec (prompt + source config) for a report definition.' The second sentence adds a concrete invocation context (autonomous gatherer callback), which makes its distinct role relative to sibling tools obvious.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states when this tool is used: the autonomous gatherer calls it on callback to learn what to collect for the named report. It does not explicitly mention alternatives or exclusions, but the callback context is strong enough to guide an agent without ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sweep_toolA
Read the pre-gathered sweep for a run.
Swept reports are gathered by Metsuke before the gatherer is called: fixed source calls, reply-state checks and noise filters run in code, and the results are stored on the run as sections of compact records. Call with no section for the index (window, overall status, per-section record counts and errors), then page through each section you need.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | 1-based page of records within the section | |
| run_id | No | The run from the callback (preferred) | |
| section | No | A section name from the index; omit for the index itself | |
| page_size | No | Records per page (max 50) | |
| report_name | No | Alternatively, use this report's newest swept run |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden and does meaningful work: it discloses that the records are pre-computed and stored on the run (not fetched live at call time) and that filtering already happened in code, which tells the agent why the results are compact/limited. It says nothing about permissions, error behavior, or freshness beyond that, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action is front-loaded in the first sentence, and the remaining sentences explain the domain model and the call pattern without obvious filler. Slightly dense for a five-parameter read tool, but every sentence carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description covers what the index contains and how paging works. It leaves minor gaps around pagination boundaries and what happens when a section name is invalid, which keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real semantics to the optional `section` parameter by explaining that omitting it returns the index (window, overall status, per-section counts and errors) and that sections are then paged individually. That index/section interaction is not derivable from the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Read the pre-gathered sweep for a run") and then defines what a sweep actually is — reports gathered by Metsuke before the gatherer is called, with fixed source calls, reply-state checks and noise filters run in code. That domain framing lets an agent distinguish this clearly from siblings like get_output_tool or list_reports_tool without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit workflow: "Call with no section for the index... then page through each section you need." The two-step discovery-then-paginate pattern is clear and actionable. It stops short of naming when to prefer an alternative tool or a when-not condition, so it is not a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_outputs_toolA
List saved report outputs (the run history), newest first.
Returns one lightweight row per saved gather — id, report_name, gathered_at, reporting window, status, and a finding_count — but NOT the payloads, so the full history stays cheap to browse. Use get_output_tool to pull one run's findings, and delete_output_tool / prune_outputs_tool to prune.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max rows to return, newest first (default 50, max 500) | |
| report_name | No | Optional report to filter to; None lists across all reports |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the disclosure burden, and it does well: it reveals that results are lightweight one-row-per-run entries, that payloads are deliberately excluded to keep browsing cheap, and enumerates the returned fields. It is silent on auth or rate-limit behavior, which keeps it out of 5 territory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the purpose and ordering, followed by the payload-exclusion rationale and sibling routing. Every clause earns its place and nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is a simple two-parameter read and an output schema exists, so return values need no expansion. The description covers purpose, cost rationale, and sibling routing, leaving only minor gaps such as pagination behavior for large histories.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters (limit, report_name) are fully documented in the schema, so the baseline is 3. The description adds no parameter-specific detail (no syntax, defaults, or filtering semantics) beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List saved report outputs (the run history)') plus ordering ('newest first'), and explicitly contrasts itself with get_output_tool, delete_output_tool, and prune_outputs_tool. An agent can distinguish this browse-the-history tool from siblings without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It routes the agent explicitly: use get_output_tool to pull one run's findings, and delete_output_tool / prune_outputs_tool to prune. That gives clear selection criteria against the closest alternatives, though it does not state an explicit 'when not to use this' condition or any prerequisite/permission requirement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_reports_toolA
List all report definitions in the catalog.
Returns each definition with its gather prompt, owner agent, schedule, source config, and last-updated time.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the behavioral burden. It clearly states this is a listing operation that 'returns each definition' with specific fields, implying a non-mutating read behavior. It does not mention edge cases or pagination, but for a simple list tool this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no filler. The core purpose is front-loaded in the first sentence, and the second sentence efficiently enumerates the return fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With zero parameters, an output schema present, and a clear statement of purpose and returned fields, the description fully supports correct invocation and selection. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. The description reinforces that it lists 'all' definitions and returns each one, making it clear there are no filters or arguments required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('List all report definitions') and a specific resource ('the catalog'), making the tool's function clear. It does not explicitly mention sibling tools, but the 'all' scope and report-definition focus distinguish it from the get/save/upsert siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied: use it when you need to enumerate all report definitions in the catalog. However, it does not explicitly state when not to use it or how it relates to alternatives like get_spec_tool or get_output_tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prune_outputs_toolA
Bulk-prune a report's saved outputs. Returns the count and ids deleted.
Give exactly one criterion: keep_last retains the N most recent outputs and deletes the rest (keep_last=0 deletes them all); before_date deletes every output gathered strictly before that date.
| Name | Required | Description | Default |
|---|---|---|---|
| keep_last | No | Retain this many newest outputs, delete older ones | |
| before_date | No | Delete outputs gathered before this date (YYYY-MM-DD) | |
| report_name | Yes | The report whose outputs to prune |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses what gets destroyed and the precise deletion conditions (keep_last=0 deletes all, before_date is strict), but says nothing about irreversibility, required permissions, or whether a confirmation is expected for a destructive bulk operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and result, followed by a compact criteria explanation with no filler. The 'Returns the count and ids deleted' line is mild overlap with the existing output schema, but it is a single clause and reads well at a glance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be detailed, and the description covers the prune semantics and criteria exclusivity adequately for a 3-param destructive tool. Remaining gaps are the lack of any permission/irreversibility note, which would matter for a bulk delete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds genuine meaning beyond the schema: the two criteria are mutually exclusive (use exactly one) and keep_last=0 has the special meaning of deleting everything. It also clarifies strictness of before_date, which the schema does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (bulk-prune) and resource (a report's saved outputs), with the 'bulk' scope implicitly distinguishing it from the singular delete_output_tool sibling. An agent can tell it apart from delete_output_tool and save_output_tool without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly requires 'exactly one criterion' and explains both options (keep_last, before_date), which is real selection guidance. It does not, however, name an alternative tool or state when a caller should prefer delete_output_tool or list_outputs_tool instead, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_output_toolA
Persist a gathered report output, completing the run that opened it.
The gatherer writes findings here after sweeping the sources. Each finding in payload should carry its own source URL so the compiler can cite it. Pass the run_id Metsuke handed you in the fire callback so this completes that exact in-flight run (one row per fire). Omit run_id for an ad-hoc direct save; Metsuke stamps a fresh run identity either way.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | No | The run to complete (from the fire callback); None = direct save | |
| status | No | One of "gathering", "ready", "compiled", "failed" (default: "ready") | ready |
| payload | Yes | List of findings. Each needs a non-empty summary and should carry its source_url; see Finding for the other allowed keys | |
| window_end | No | End of the reporting window (ISO date/datetime) | |
| report_name | Yes | The report definition this output belongs to | |
| window_start | No | Start of the reporting window (ISO date/datetime) | |
| gatherer_run_ref | No | Opaque reference to the gatherer run that produced this |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose meaningful behavior: one row per fire, that omitting run_id still produces a fresh run identity ('Metsuke stamps a fresh run identity either way'), and that findings should carry source URLs for downstream citation. It omits idempotency/duplicate handling and permission requirements, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose in the first sentence, then layers usage detail. Four sentences, each earning its place, though domain jargon ('fire callback', 'Metsuke') assumes shared context and slightly reduces standalone clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is unnecessary, and the description correctly focuses on the input-side lifecycle: run_id provenance, payload expectations, and the two save modes. Adequate for a mutation tool of this complexity, with only minor gaps around failure/duplicate behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description mildly enriches run_id semantics ('one row per fire', completes 'that exact in-flight run') and reiterates the source_url expectation, but adds little beyond what the schema descriptions already document for each field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource: 'Persist a gathered report output', and immediately clarifies its transactional role ('completing the run that opened it'). This clearly positions it as the write path, distinguishing it from read-oriented siblings like get_output_tool, list_outputs_tool, and delete_output_tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit timing ('The gatherer writes findings here after sweeping the sources') and two distinct invocation modes: pass run_id from the fire callback to complete an in-flight run, or omit it for an ad-hoc save. It stops short of naming sibling alternatives (e.g., get_output_tool) to route away from, but the when-to-use context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trigger_report_toolA
Fire a report gather right now, without waiting for its schedule.
Dispatches the same callback the scheduler uses: Metsuke POSTs the trigger to the Trentina alert endpoint, which forwards it to the owning gatherer. Requires the callback to be configured (TRENTINA_ALERT_URL + METSUKE_ALERT_TOKEN).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The report definition name to fire (e.g. "core-platform-status") |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the external callback chain (Metsuke POSTs to the Trentina alert endpoint, forwarded to the owning gatherer) and the configuration prerequisites (TRENTINA_ALERT_URL + METSUKE_ALERT_TOKEN). It does not say whether the call blocks until the gather finishes, nor what happens on failure when the callback is unconfigured.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the imperative action, then mechanism, then prerequisite. The middle sentence on the alert endpoint is somewhat internal but earns its place as behavioral context; overall tight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the prerequisites and side-effect path are stated. The main remaining gap is the synchronous/asynchronous nature of the trigger and error behavior, but for a one-parameter tool this is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single parameter is documented in-schema with an example value, so the baseline is 3. The description adds no extra meaning about the name parameter beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: fire a report gather immediately rather than on its schedule. The action is unambiguous and clearly distinct from the list/get/upsert/delete siblings, though it never names or contrasts them explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Right now, without waiting for its schedule" gives a clear condition for selecting this tool over simply letting the scheduler run. No explicit when-not or alternative tool is named, but the triggering context is well defined.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upsert_definition_toolB
Create or update a report definition.
The owner_agent field makes the gatherer identity data, not code, so gathering can be re-homed to another agent without a rebuild. The schedule is a live cron expression: Metsuke's built-in scheduler fires the report when it comes due, in the given timezone.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Unique report name (e.g. "core-platform-status") | |
| schedule | No | Cron expression for when to fire (e.g. "0 6 * * 5"); None = manual only | |
| timezone | No | IANA timezone the schedule runs in (default: "UTC") | UTC |
| owner_agent | No | Which agent gathers this report (default: "kagetora") | kagetora |
| gather_prompt | Yes | The instruction the gatherer runs to collect findings | |
| source_config | No | Which sources to sweep, as a JSON object |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the behavioral burden. It usefully discloses upsert semantics, that owner_agent makes identity data rather than code (re-homing without a rebuild), and that schedule is a live cron fired by the built-in scheduler. It omits update-overwrite behavior, permission requirements, and side effects, so it adds context but not a full behavioral picture.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action is front-loaded in the first sentence, followed by two tight sentences that each explain a distinct field behavior. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation. For a six-parameter upsert mutation with no annotations, the description covers the two most conceptually important fields but leaves source_config, gather_prompt, and the update-vs-create behavioral distinction unaddressed, making it adequate but incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, but the description adds genuine meaning beyond it: owner_agent carries the re-homing rationale and schedule is framed as a live cron fired by Metsuke's scheduler. This enriches at least two parameters above the schema baseline, though source_config and gather_prompt remain unaddressed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair (create or update) and resource (report definition), which is clear. However it does not differentiate from siblings like list_reports_tool, trigger_report_tool, or delete_output_tool, so the agent gets a precise action but no routing contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The text explains field semantics rather than when to invoke this tool versus alternatives. There is no when-to-use, when-not-to-use, or sibling routing guidance (e.g., trigger_report_tool for firing an existing report), leaving the agent to infer selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v1.1.0- Added
delete_output_tool - Added
get_sweep_tool - Added
list_outputs_tool - Added
prune_outputs_tool - Changed
save_output_tool8 fields changed- changed
Input schema / properties / payload / descriptionPrevious value: -"List of finding objects, each ideally carrying a source URL"New value: +"List of findings. Each needs a non-empty summary and should\ncarry its source_url; see Finding for the other allowed keys" - changed
Input schema / properties / payload / items / additionalPropertiesPrevious value: -trueNew value: +false - added
Input schema / properties / payload / items / descriptionAdded value: +"One gathered finding.\n\nThe fields are declared, not left as a free-form dict, because a model\ncalling the tool fills in what the schema names: with a bare ``object``\nitem, strict tool-calling models emitted ``{}`` for every finding (RT #1505).\n``summary`` is the one field every report uses and the one a finding is\nworthless without. The rest are the keys the live reports use; a report\nthat needs one more adds it here.\n\nExample::\n\n {\"summary\": \"Fedora 45 Beta shipped with Podman 6.\",\n \"source_url\": \"https://fedoramagazine.org/...\",\n \"section\": \"rss-news-roundup\", \"theme\": \"RHEL/Linux\"}" - added
Input schema / properties / payload / items / propertiesAdded value: +{ + "actors": { + "anyOf": [ + { + "items": { + "maxLength": 200, + "type": "string" + }, + "maxItems": 2000, + "type": "array" + }, + { + "type": "null" + } + ], + "default": null, + "description": "People or teams involved, by name." + }, + "category": { + "anyOf": [ + { + "maxLength": 200, + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Source category, e.g. the feed category the item came from." + }, + "date": { + "anyOf": [ + { + "maxLength": 200, + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "When it happened, ISO date (YYYY-MM-DD)." + }, + "outcome_ref": { + "anyOf": [ + { + "maxLength": 2000, + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Tracker key this finding advances, e.g. a Jira issue key." + }, + "section": { + "anyOf": [ + { + "maxLength": 200, + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Report section this belongs in, as named by the report's gather prompt." + }, + "source_type": { + "anyOf": [ + { + "maxLength": 200, + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Kind of source, e.g. 'jira', 'slack', 'email', 'rss', 'web'." + }, + "source_url": { + "anyOf": [ + { + "maxLength": 2000, + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Clickable link to the evidence; null when no source exists." + }, + "summary": { + "description": "One or two sentences stating the finding itself.", + "maxLength": 2000, + "minLength": 1, + "type": "string" + }, + "theme": { + "anyOf": [ + { + "maxLength": 200, + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Grouping within a section, e.g. 'Security' or 'AI/Agentic'." + }, + "title": { + "anyOf": [ + { + "maxLength": 2000, + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "Headline of the source item, when it has one." + } +} - added
Input schema / properties / payload / items / requiredAdded value: +[ + "summary" +] - added
Input schema / properties / run_idAdded value: +{ + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "description": "The run to complete (from the fire callback); None = direct save" +} - changed
Input schema / properties / status / descriptionPrevious value: -"One of \"gathering\", \"ready\", \"compiled\" (default: \"ready\")"New value: +"One of \"gathering\", \"ready\", \"compiled\", \"failed\" (default: \"ready\")" - changed
Input schema / properties / status / enumPrevious value: -[ - "gathering", - "ready", - "compiled" -]New value: +[ + "gathering", + "ready", + "compiled", + "failed" +]
- Added
trigger_report_tool - Changed
upsert_definition_tool2 fields changed- changed
Input schema / properties / schedule / descriptionPrevious value: -"Descriptive cron/time metadata (the actual firing is external)"New value: +"Cron expression for when to fire (e.g. \"0 6 * * 5\"); None = manual only" - added
Input schema / properties / timezoneAdded value: +{ + "default": "UTC", + "description": "IANA timezone the schedule runs in (default: \"UTC\")", + "type": "string" +}
5 tool updates
v0.2.0- First observed
get_output_tool - First observed
get_spec_tool - First observed
list_reports_tool - First observed
save_output_tool - First observed
upsert_definition_tool
TDQS
Scored across 10 tools
The tools split cleanly along resource lines (definitions vs. outputs vs. sweep vs. trigger), and single-vs-bulk deletion (delete_output_tool vs. prune_outputs_tool) is clearly delineated. The only mild overlap is get_spec_tool versus list_reports_tool, since list_reports already returns the gather prompt and source config that get_spec_tool exists to expose.
All ten tools follow a predictable verb_noun_tool pattern (get_output_tool, list_outputs_tool, prune_outputs_tool), which is easy to scan and reason about. Minor inconsistency in the noun chosen for the same entity: report definitions are called both 'definition' (upsert_definition_tool) and 'reports' (list_reports_tool).
Ten tools is well-scoped for a report gatherer/compiler: definition management, output lifecycle, triggering, and sweep reading each get exactly the operations they need. No redundant or filler tools.
Output lifecycle is fully covered (save, list, get, delete, prune) and definitions support list/create/update plus read via get_spec_tool. The one clear gap is the absence of a delete/remove operation for report definitions, so stale reports cannot be retired through the tool surface.
Maintenance
Related MCP Connectors
A public commons for agents to search and share reusable findings and open research questions.
One place for every AI agent's pages and docs: versioned links to share, search and update.
Machine-native research commons for agent evidence, discovery, rooms, and bounded research quests.
Search and share cited agent findings. Public reads; authenticated writes.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceProvides AI agents with persistent memory across sessions, enabling recall of decisions, clients, and deadlines with verifiable citations.-
- AlicenseNot gradedqualityAmaintenanceEnables MCP agents to maintain durable, evidence-aware project knowledge, retrieve precise excerpts on demand, and track decisions, conflicts, and revisions across sessions.1Apache 2.0
- AlicenseCqualityAmaintenanceEnables engineering agents to maintain persistent knowledge across sessions by storing decisions, invariants, gotchas, and rejected ideas, with staleness detection, conflict detection, full-text search, and structured context assembly.371MIT
- AlicenseNot gradedqualityAmaintenanceProvides persistent, searchable memory across coding projects and machines, letting agents record and retrieve projects, reusable assets, sessions, decisions, commits, and handoffs via MCP.24 PyPI2Apache 2.0