Skip to main content
Glama
crunchtools

mcp-metsuke-crunchtools

Official
by crunchtools

mcp-metsuke-crunchtools

Stateful reports catalog MCP server. Metsuke (目付 — the Sengoku intelligence officer who gathered field reports and compiled them for the daimyō) is the durable, cross-agent home for report definitions and their gathered outputs.

An autonomous gatherer reads a definition, sweeps the configured sources, and writes findings — each carrying its own source URL. A compiler later reads the freshest output to draft a fully cited report. This replaces ad-hoc, run-scoped research caches with a real, queryable store.

Features

  • Definitions + outputs — separate what-to-gather from what-was-gathered

  • Built-in scheduler — Metsuke fires each definition's gather when its cron schedule comes due, or on demand via trigger_report — no external timer

  • Run lifecycle — every fire opens a durable run with a second-granularity run_id: a provisional row is recorded before the callback (a fire is never lost), and a per-report concurrency lock keeps two gathers from racing

  • Cited findings — payloads carry per-finding source URLs for one-click checking

  • Re-homeable gathering — owner_agent makes the gatherer identity data, not code

  • Local-first — plain SQLite (WAL), no external services, no per-seat fees

  • Three transports — stdio, SSE, streamable-http

Related MCP server: Argon Memory

Install

uvx mcp-metsuke-crunchtools

pip

pip install mcp-metsuke-crunchtools

Container

podman run --rm -v ~/.local/share/mcp-metsuke:/data:Z \
  quay.io/crunchtools/mcp-metsuke \
  --transport streamable-http --host 0.0.0.0 --port 8009

Claude Code Integration

claude mcp add mcp-metsuke-crunchtools -- uvx mcp-metsuke-crunchtools

Tools (10)

Definitions (4)

Tool

Description

list_reports

List all report definitions in the catalog, each with its next scheduled fire time.

get_spec

Return the gather prompt + source config for one definition.

upsert_definition

Create or update a definition (name, prompt, owner, cron schedule, timezone, sources).

trigger_report

Fire a report gather right now, without waiting for its schedule. For a swept definition it returns at once with dispatched: "after_sweep"; the sweep and the callback follow in the background.

Outputs (6)

Tool

Description

save_output

Complete a run by persisting its gathered findings (pass the run_id from the fire callback; each finding ideally carrying a source URL).

get_output

Read the freshest output (or a specific day's) for compiling a report.

list_outputs

Browse the run history — metadata + finding_count per saved gather, no payloads.

delete_output

Delete one saved output by id.

prune_outputs

Bulk-prune a report's outputs — keep the N newest, or drop those before a date.

get_sweep

Read a swept run's pre-gathered records: the index, or one page of one section.

Environment Variables

Variable

Default

Description

METSUKE_DB

~/.local/share/mcp-metsuke/metsuke.db

SQLite database path

METSUKE_DB_FILE

(none)

Path whose contents override METSUKE_DB (container secret-file convention)

TRENTINA_ALERT_URL

(none)

Base URL of the Trentina gateway; the scheduler POSTs gather callbacks to <url>/alert with the token as Authorization: Bearer

METSUKE_ALERT_TOKEN

(none)

Alert token identifying the reports profile; enables the scheduler when set with TRENTINA_ALERT_URL

METSUKE_ALERT_TOKEN_FILE

(none)

Path whose contents override METSUKE_ALERT_TOKEN (container secret-file convention)

TRENTINA_GATEWAY_URL

(none)

Trentina gateway MCP endpoint for the sweep profile, e.g. http://mcp-trentina:8019/gateway/metsuke-sweep/mcp. Plain HTTP is accepted only for internal hosts (single-label service names, localhost, private IPs); anything else must be HTTPS

METSUKE_SWEEP_TOKEN

(none)

Bearer token for that profile; METSUKE_SWEEP_TOKEN_FILE overrides it

METSUKE_SWEEP_TIMEOUT_SECONDS

1200

Upper bound on one run's sweep

METSUKE_SCHEDULER_POLL_SECONDS

60

How often the scheduler checks for due reports

METSUKE_SCHEDULER_ENABLED

(auto)

Force the scheduler on/off; defaults to on when the callback is configured

METSUKE_RUN_LOCK_TTL_SECONDS

1800

How long an in-flight run holds the per-report lock before it is expired (self-heals a dead gatherer)

The scheduler runs only under the sse and streamable-http transports (the long-lived production processes), never under stdio.

Data Model

  • report_definitions — name (PK), gather_prompt, owner_agent, schedule (cron), timezone (IANA), source_config (JSON), last_fired_at, updated_at

  • report_outputs (one row per run) — id, report_name (FK), run_id (second-granularity identity), trigger (scheduled/manual/direct), gathered_at, finished_at, window_start/window_end, payload (JSON findings with source URLs), status (gathering/ready/compiled/failed), gatherer_run_ref, detail. A partial unique index on (report_name) WHERE status='gathering' is the per-report concurrency lock.

MCP Registry

io.github.crunchtools/metsuke

License

AGPL-3.0-or-later

Sweeps

A definition can opt into a deterministic pre-gather by adding sweep to its source_config:

{"sweep": {"timezone": "America/New_York", "window_hour": 6, "steps": [
  {"section": "slack", "collector": "slack_waiting",
   "options": {"user_id": "U9VN3S1ST", "handle": "smccarty"}},
  {"section": "work_email", "collector": "gmail_waiting",
   "options": {"backend": "gw-work", "account": "smccarty@redhat.com"}},
  {"section": "calendar", "collector": "calendar_day",
   "options": {"backend": "gw-work", "account": "smccarty@redhat.com"}},
  {"section": "rss", "collector": "feed_entries",
   "options": {"categories": {"1": 20, "5": 20}}}]}}

On fire, Metsuke runs each step in order through the Trentina gateway (as its own read-only profile), stores the results on the run, then dispatches the callback with sweep_status. The gatherer reads records with get_sweep instead of calling the sources, so the gathering LLM only selects, phrases and delivers. A failing step is recorded and the sweep continues. Results Trentina flags are stored as metadata only, with their text withheld.

Collector

What it produces

slack_waiting

Direct asks over a lookback: an @-mention, or a DM message that reads as a question or request (DM chatter and Slack system notices are not asks). Threads are read as threads; DMs and top-level channel mentions are read from channel history (one conversation per channel). Each is read for up to three pages (longer or partly unreadable ones are marked unverified). An ask the user replied to or reacted to is answered and dropped; the rest are waiting, with in_window separating new asks from still-open ones

gmail_waiting

Inbox threads in the window (or lookback_days) where the backend's ownership analysis says the ball is in the user's court; automated mail, bare calendar notices and threads newer than min_age_hours dropped, invitations kept, priority_senders marked priority

calendar_day

The report day's meetings, pending invites over a lookahead, and hard overlaps

feed_entries

Recent entries per feed category (read or unread), with a longer window after a weekend

jira_issues

Issues matching each named JQL query, with {since} substituted by the window start. Each record carries the key, browse link, status, components and age; where the description is a web intake form, its Field: value pairs are parsed into contact and a capped detail line. A failing query is recorded and the rest still run

Collector options

Every step is {"section": "<name>", "collector": "<collector>", "options": {...}}. Section names are [a-z0-9_-], unique, up to 12 steps. Top-level timezone (IANA, default America/New_York) and window_hour (0-23, default 6) set the window: from that hour on the previous weekday until the run.

Collector

Option

Required

Default

Bounds

slack_waiting

user_id

yes

2-32 chars

handle

yes

1-64 chars

backend

slack

≤64 chars

self_label

you

≤64 chars

first_name

none

1-64 chars; in group DMs and channels, a question or request that addresses the user by this name is an ask

lookback_days

7

1-30

max_conversations

40

1-100

workspace_url

https://redhat-internal.slack.com

≤200 chars

gmail_waiting

backend

yes

e.g. gw-work, gw-personal

account

yes

the mailbox address

query_extra

""

≤500 chars, appended to in:inbox after:<window>

max_threads

60

1-200

body_chars

1200

0-4000 (0 = no body)

link_template

Gmail #all/{thread_id}

≤200 chars, or null

lookback_days

report window

1-30, search the last N days instead

min_age_hours

0

0-168, drop threads active more recently

priority_senders

[]

≤20 From-header substrings (1-100 chars, case-insensitive)

calendar_day

backend

yes

account

yes

timezone

America/New_York

IANA zone

lookahead_days

3

0-14 (pending invites)

feed_entries

categories

yes

1-12 entries of "<id>": <limit>, id ≤9 digits, limit 1-100

backend

feeds

since_days

1

1-30

since_days_after_weekend

3

1-30 (Mondays and weekend runs)

jira_issues

queries

yes

1-6 of {"label", "jql", "limit"}: label [a-z0-9_-] ≤32, jql ≤600 chars ({since} → window start as YYYY-MM-DD HH:mm), limit 1-50

backend

jira

≤64 chars; needs jira_search allowed

browse_url

https://redhat.atlassian.net/browse/

≤200 chars, prefixed to the issue key

detail_chars

800

0-2000 (0 = no form detail line)

Collectors are code in sweep/collectors.py, not configuration: a definition can only pick and parameterize them. A bad spec is rejected at upsert_definition time.

Available Tools

10 tools
delete_output_toolA

Delete one saved report output by id.

Returns the deleted output's metadata (with deleted=True). Find ids with list_outputs_tool. Raises if no output carries that id.

ParametersJSON Schema
NameRequiredDescriptionDefault
output_idYesThe id of the output to delete (from list_outputs_tool)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and does well: it discloses the return shape (deleted metadata with deleted=True) and the failure mode (raises if no output has that id). It does not explicitly call out irreversibility or auth requirements, but the destructive nature and error contract are clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action, then return value, then error behavior. No wasted words and the id-discovery hint arrives exactly where needed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter mutation with an output schema, everything an agent needs is present: the action, how to source the id, the success return, and the failure condition. Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter already documents its type and provenance (from list_outputs_tool). The description only restates 'by id', adding no syntax or format detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Delete), resource (saved report output), and scope (one, by id), which cleanly distinguishes it from list_outputs_tool and prune_outputs_tool. An agent knows exactly what operation this performs without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Directs the agent to list_outputs_tool for obtaining ids, giving a clear prerequisite workflow. It does not explicitly mention the bulk-delete alternative (prune_outputs_tool) or state when-not to use this tool, so it stops short of full sibling routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_output_toolA

Read a gathered report output for compiling a report.

Returns the most recent output by default, or the latest output gathered on a specific date when gathered_date is given.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe report definition name
gathered_dateNoOptional YYYY-MM-DD to fetch that day's output instead

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It discloses the default behavior (most recent output), the date-specific behavior, and the fallback semantics of 'latest output gathered on a specific date.' This is meaningful behavioral context beyond what the schema expresses, though it does not cover edge cases like missing outputs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the main purpose appears in the first sentence, and the only additional sentence explains the optional parameter behavior. Every sentence earns its place with no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity read tool, the description is complete: it states the resource, the default behavior, the optional date behavior, and the parameters are fully covered by the schema. An output schema exists, so return-value details are already provided elsewhere, and no critical invocation detail is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Both parameters are already well described in the schema, so the baseline is 3. The description adds context by explaining the default selection behavior and how gathered_date changes the result, but it does not significantly expand on the parameter meanings beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('Read') and resource ('gathered report output'), and it adds the purpose ('for compiling a report'). It also distinguishes this tool from siblings by clarifying it retrieves outputs, whereas list_reports_tool, get_spec_tool, and save_output_tool imply different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The usage context is implied by the phrase 'for compiling a report' and by the optional date parameter, but there is no explicit guidance on when to choose this tool over alternatives. It does not describe exclusions or mention sibling tools, so the agent must infer the appropriate moment to call it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_spec_toolA

Return the gather spec (prompt + source config) for a report definition.

This is what the autonomous gatherer calls on callback to learn what to collect for the named report.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe report definition name (e.g. "core-platform-status")

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It clarifies what the return payload contains (prompt + source config) and the invocation context, but it does not explicitly state that the operation is read-only, what happens for unknown report names, or any auth requirements. 'Return' implies a safe read, but it is not explicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The first sentence front-loads the core action and resource; the second adds valuable contextual information about the caller. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read tool with an existing output schema, the description provides enough to invoke it correctly: what it returns and when to use it. Minor gaps such as not-found behavior or explicit read-only assurance are acceptable for a simple getter, but would make it more complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides 100% coverage for the single 'name' parameter, including an example. The description adds only a slight restatement ('named report') and does not provide additional semantic depth such as naming conventions, validation rules, or behavior for missing names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb 'Return' with a clear resource: 'gather spec (prompt + source config) for a report definition.' The second sentence adds a concrete invocation context (autonomous gatherer callback), which makes its distinct role relative to sibling tools obvious.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly states when this tool is used: the autonomous gatherer calls it on callback to learn what to collect for the named report. It does not explicitly mention alternatives or exclusions, but the callback context is strong enough to guide an agent without ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_sweep_toolA

Read the pre-gathered sweep for a run.

Swept reports are gathered by Metsuke before the gatherer is called: fixed source calls, reply-state checks and noise filters run in code, and the results are stored on the run as sections of compact records. Call with no section for the index (window, overall status, per-section record counts and errors), then page through each section you need.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNo1-based page of records within the section
run_idNoThe run from the callback (preferred)
sectionNoA section name from the index; omit for the index itself
page_sizeNoRecords per page (max 50)
report_nameNoAlternatively, use this report's newest swept run

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden and does meaningful work: it discloses that the records are pre-computed and stored on the run (not fetched live at call time) and that filtering already happened in code, which tells the agent why the results are compact/limited. It says nothing about permissions, error behavior, or freshness beyond that, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core action is front-loaded in the first sentence, and the remaining sentences explain the domain model and the call pattern without obvious filler. Slightly dense for a five-parameter read tool, but every sentence carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers what the index contains and how paging works. It leaves minor gaps around pagination boundaries and what happens when a section name is invalid, which keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real semantics to the optional `section` parameter by explaining that omitting it returns the index (window, overall status, per-section counts and errors) and that sections are then paged individually. That index/section interaction is not derivable from the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Read the pre-gathered sweep for a run") and then defines what a sweep actually is — reports gathered by Metsuke before the gatherer is called, with fixed source calls, reply-state checks and noise filters run in code. That domain framing lets an agent distinguish this clearly from siblings like get_output_tool or list_reports_tool without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit workflow: "Call with no section for the index... then page through each section you need." The two-step discovery-then-paginate pattern is clear and actionable. It stops short of naming when to prefer an alternative tool or a when-not condition, so it is not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_outputs_toolA

List saved report outputs (the run history), newest first.

Returns one lightweight row per saved gather — id, report_name, gathered_at, reporting window, status, and a finding_count — but NOT the payloads, so the full history stays cheap to browse. Use get_output_tool to pull one run's findings, and delete_output_tool / prune_outputs_tool to prune.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax rows to return, newest first (default 50, max 500)
report_nameNoOptional report to filter to; None lists across all reports

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the disclosure burden, and it does well: it reveals that results are lightweight one-row-per-run entries, that payloads are deliberately excluded to keep browsing cheap, and enumerates the returned fields. It is silent on auth or rate-limit behavior, which keeps it out of 5 territory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the purpose and ordering, followed by the payload-exclusion rationale and sibling routing. Every clause earns its place and nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is a simple two-parameter read and an output schema exists, so return values need no expansion. The description covers purpose, cost rationale, and sibling routing, leaving only minor gaps such as pagination behavior for large histories.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters (limit, report_name) are fully documented in the schema, so the baseline is 3. The description adds no parameter-specific detail (no syntax, defaults, or filtering semantics) beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List saved report outputs (the run history)') plus ordering ('newest first'), and explicitly contrasts itself with get_output_tool, delete_output_tool, and prune_outputs_tool. An agent can distinguish this browse-the-history tool from siblings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It routes the agent explicitly: use get_output_tool to pull one run's findings, and delete_output_tool / prune_outputs_tool to prune. That gives clear selection criteria against the closest alternatives, though it does not state an explicit 'when not to use this' condition or any prerequisite/permission requirement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_reports_toolA

List all report definitions in the catalog.

Returns each definition with its gather prompt, owner agent, schedule, source config, and last-updated time.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the behavioral burden. It clearly states this is a listing operation that 'returns each definition' with specific fields, implying a non-mutating read behavior. It does not mention edge cases or pagination, but for a simple list tool this is adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no filler. The core purpose is front-loaded in the first sentence, and the second sentence efficiently enumerates the return fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With zero parameters, an output schema present, and a clear statement of purpose and returned fields, the description fully supports correct invocation and selection. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. The description reinforces that it lists 'all' definitions and returns each one, making it clear there are no filters or arguments required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('List all report definitions') and a specific resource ('the catalog'), making the tool's function clear. It does not explicitly mention sibling tools, but the 'all' scope and report-definition focus distinguish it from the get/save/upsert siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied: use it when you need to enumerate all report definitions in the catalog. However, it does not explicitly state when not to use it or how it relates to alternatives like get_spec_tool or get_output_tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

prune_outputs_toolA

Bulk-prune a report's saved outputs. Returns the count and ids deleted.

Give exactly one criterion: keep_last retains the N most recent outputs and deletes the rest (keep_last=0 deletes them all); before_date deletes every output gathered strictly before that date.

ParametersJSON Schema
NameRequiredDescriptionDefault
keep_lastNoRetain this many newest outputs, delete older ones
before_dateNoDelete outputs gathered before this date (YYYY-MM-DD)
report_nameYesThe report whose outputs to prune

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses what gets destroyed and the precise deletion conditions (keep_last=0 deletes all, before_date is strict), but says nothing about irreversibility, required permissions, or whether a confirmation is expected for a destructive bulk operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and result, followed by a compact criteria explanation with no filler. The 'Returns the count and ids deleted' line is mild overlap with the existing output schema, but it is a single clause and reads well at a glance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be detailed, and the description covers the prune semantics and criteria exclusivity adequately for a 3-param destructive tool. Remaining gaps are the lack of any permission/irreversibility note, which would matter for a bulk delete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds genuine meaning beyond the schema: the two criteria are mutually exclusive (use exactly one) and keep_last=0 has the special meaning of deleting everything. It also clarifies strictness of before_date, which the schema does not convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (bulk-prune) and resource (a report's saved outputs), with the 'bulk' scope implicitly distinguishing it from the singular delete_output_tool sibling. An agent can tell it apart from delete_output_tool and save_output_tool without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly requires 'exactly one criterion' and explains both options (keep_last, before_date), which is real selection guidance. It does not, however, name an alternative tool or state when a caller should prefer delete_output_tool or list_outputs_tool instead, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_output_toolA

Persist a gathered report output, completing the run that opened it.

The gatherer writes findings here after sweeping the sources. Each finding in payload should carry its own source URL so the compiler can cite it. Pass the run_id Metsuke handed you in the fire callback so this completes that exact in-flight run (one row per fire). Omit run_id for an ad-hoc direct save; Metsuke stamps a fresh run identity either way.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idNoThe run to complete (from the fire callback); None = direct save
statusNoOne of "gathering", "ready", "compiled", "failed" (default: "ready")ready
payloadYesList of findings. Each needs a non-empty summary and should carry its source_url; see Finding for the other allowed keys
window_endNoEnd of the reporting window (ISO date/datetime)
report_nameYesThe report definition this output belongs to
window_startNoStart of the reporting window (ISO date/datetime)
gatherer_run_refNoOpaque reference to the gatherer run that produced this

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose meaningful behavior: one row per fire, that omitting run_id still produces a fresh run identity ('Metsuke stamps a fresh run identity either way'), and that findings should carry source URLs for downstream citation. It omits idempotency/duplicate handling and permission requirements, which keeps it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads purpose in the first sentence, then layers usage detail. Four sentences, each earning its place, though domain jargon ('fire callback', 'Metsuke') assumes shared context and slightly reduces standalone clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is unnecessary, and the description correctly focuses on the input-side lifecycle: run_id provenance, payload expectations, and the two save modes. Adequate for a mutation tool of this complexity, with only minor gaps around failure/duplicate behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description mildly enriches run_id semantics ('one row per fire', completes 'that exact in-flight run') and reiterates the source_url expectation, but adds little beyond what the schema descriptions already document for each field.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource: 'Persist a gathered report output', and immediately clarifies its transactional role ('completing the run that opened it'). This clearly positions it as the write path, distinguishing it from read-oriented siblings like get_output_tool, list_outputs_tool, and delete_output_tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit timing ('The gatherer writes findings here after sweeping the sources') and two distinct invocation modes: pass run_id from the fire callback to complete an in-flight run, or omit it for an ad-hoc save. It stops short of naming sibling alternatives (e.g., get_output_tool) to route away from, but the when-to-use context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

trigger_report_toolA

Fire a report gather right now, without waiting for its schedule.

Dispatches the same callback the scheduler uses: Metsuke POSTs the trigger to the Trentina alert endpoint, which forwards it to the owning gatherer. Requires the callback to be configured (TRENTINA_ALERT_URL + METSUKE_ALERT_TOKEN).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe report definition name to fire (e.g. "core-platform-status")

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the external callback chain (Metsuke POSTs to the Trentina alert endpoint, forwarded to the owning gatherer) and the configuration prerequisites (TRENTINA_ALERT_URL + METSUKE_ALERT_TOKEN). It does not say whether the call blocks until the gather finishes, nor what happens on failure when the callback is unconfigured.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the imperative action, then mechanism, then prerequisite. The middle sentence on the alert endpoint is somewhat internal but earns its place as behavioral context; overall tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the prerequisites and side-effect path are stated. The main remaining gap is the synchronous/asynchronous nature of the trigger and error behavior, but for a one-parameter tool this is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single parameter is documented in-schema with an example value, so the baseline is 3. The description adds no extra meaning about the name parameter beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: fire a report gather immediately rather than on its schedule. The action is unambiguous and clearly distinct from the list/get/upsert/delete siblings, though it never names or contrasts them explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Right now, without waiting for its schedule" gives a clear condition for selecting this tool over simply letting the scheduler run. No explicit when-not or alternative tool is named, but the triggering context is well defined.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upsert_definition_toolB

Create or update a report definition.

The owner_agent field makes the gatherer identity data, not code, so gathering can be re-homed to another agent without a rebuild. The schedule is a live cron expression: Metsuke's built-in scheduler fires the report when it comes due, in the given timezone.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesUnique report name (e.g. "core-platform-status")
scheduleNoCron expression for when to fire (e.g. "0 6 * * 5"); None = manual only
timezoneNoIANA timezone the schedule runs in (default: "UTC")UTC
owner_agentNoWhich agent gathers this report (default: "kagetora")kagetora
gather_promptYesThe instruction the gatherer runs to collect findings
source_configNoWhich sources to sweep, as a JSON object

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral burden. It usefully discloses upsert semantics, that owner_agent makes identity data rather than code (re-homing without a rebuild), and that schedule is a live cron fired by the built-in scheduler. It omits update-overwrite behavior, permission requirements, and side effects, so it adds context but not a full behavioral picture.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core action is front-loaded in the first sentence, followed by two tight sentences that each explain a distinct field behavior. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation. For a six-parameter upsert mutation with no annotations, the description covers the two most conceptually important fields but leaves source_config, gather_prompt, and the update-vs-create behavioral distinction unaddressed, making it adequate but incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, but the description adds genuine meaning beyond it: owner_agent carries the re-homing rationale and schedule is framed as a live cron fired by Metsuke's scheduler. This enriches at least two parameters above the schema baseline, though source_config and gather_prompt remain unaddressed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair (create or update) and resource (report definition), which is clear. However it does not differentiate from siblings like list_reports_tool, trigger_report_tool, or delete_output_tool, so the agent gets a precise action but no routing contrast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The text explains field semantics rather than when to invoke this tool versus alternatives. There is no when-to-use, when-not-to-use, or sibling routing guidance (e.g., trigger_report_tool for firing an existing report), leaving the agent to infer selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv1.1.0
    • Addeddelete_output_tool
    • Addedget_sweep_tool
    • Addedlist_outputs_tool
    • Addedprune_outputs_tool
    • Changedsave_output_tool8 fields changed
      • changedInput schema / properties / payload / description
        Previous value: -"List of finding objects, each ideally carrying a source URL"New value: +"List of findings. Each needs a non-empty summary and should\ncarry its source_url; see Finding for the other allowed keys"
      • changedInput schema / properties / payload / items / additionalProperties
        Previous value: -trueNew value: +false
      • addedInput schema / properties / payload / items / description
        Added value: +"One gathered finding.\n\nThe fields are declared, not left as a free-form dict, because a model\ncalling the tool fills in what the schema names: with a bare ``object``\nitem, strict tool-calling models emitted ``{}`` for every finding (RT #1505).\n``summary`` is the one field every report uses and the one a finding is\nworthless without. The rest are the keys the live reports use; a report\nthat needs one more adds it here.\n\nExample::\n\n    {\"summary\": \"Fedora 45 Beta shipped with Podman 6.\",\n     \"source_url\": \"https://fedoramagazine.org/...\",\n     \"section\": \"rss-news-roundup\", \"theme\": \"RHEL/Linux\"}"
      • addedInput schema / properties / payload / items / properties
        Added value: +{
        +  "actors": {
        +    "anyOf": [
        +      {
        +        "items": {
        +          "maxLength": 200,
        +          "type": "string"
        +        },
        +        "maxItems": 2000,
        +        "type": "array"
        +      },
        +      {
        +        "type": "null"
        +      }
        +    ],
        +    "default": null,
        +    "description": "People or teams involved, by name."
        +  },
        +  "category": {
        +    "anyOf": [
        +      {
        +        "maxLength": 200,
        +        "type": "string"
        +      },
        +      {
        +        "type": "null"
        +      }
        +    ],
        +    "default": null,
        +    "description": "Source category, e.g. the feed category the item came from."
        +  },
        +  "date": {
        +    "anyOf": [
        +      {
        +        "maxLength": 200,
        +        "type": "string"
        +      },
        +      {
        +        "type": "null"
        +      }
        +    ],
        +    "default": null,
        +    "description": "When it happened, ISO date (YYYY-MM-DD)."
        +  },
        +  "outcome_ref": {
        +    "anyOf": [
        +      {
        +        "maxLength": 2000,
        +        "type": "string"
        +      },
        +      {
        +        "type": "null"
        +      }
        +    ],
        +    "default": null,
        +    "description": "Tracker key this finding advances, e.g. a Jira issue key."
        +  },
        +  "section": {
        +    "anyOf": [
        +      {
        +        "maxLength": 200,
        +        "type": "string"
        +      },
        +      {
        +        "type": "null"
        +      }
        +    ],
        +    "default": null,
        +    "description": "Report section this belongs in, as named by the report's gather prompt."
        +  },
        +  "source_type": {
        +    "anyOf": [
        +      {
        +        "maxLength": 200,
        +        "type": "string"
        +      },
        +      {
        +        "type": "null"
        +      }
        +    ],
        +    "default": null,
        +    "description": "Kind of source, e.g. 'jira', 'slack', 'email', 'rss', 'web'."
        +  },
        +  "source_url": {
        +    "anyOf": [
        +      {
        +        "maxLength": 2000,
        +        "type": "string"
        +      },
        +      {
        +        "type": "null"
        +      }
        +    ],
        +    "default": null,
        +    "description": "Clickable link to the evidence; null when no source exists."
        +  },
        +  "summary": {
        +    "description": "One or two sentences stating the finding itself.",
        +    "maxLength": 2000,
        +    "minLength": 1,
        +    "type": "string"
        +  },
        +  "theme": {
        +    "anyOf": [
        +      {
        +        "maxLength": 200,
        +        "type": "string"
        +      },
        +      {
        +        "type": "null"
        +      }
        +    ],
        +    "default": null,
        +    "description": "Grouping within a section, e.g. 'Security' or 'AI/Agentic'."
        +  },
        +  "title": {
        +    "anyOf": [
        +      {
        +        "maxLength": 2000,
        +        "type": "string"
        +      },
        +      {
        +        "type": "null"
        +      }
        +    ],
        +    "default": null,
        +    "description": "Headline of the source item, when it has one."
        +  }
        +}
      • addedInput schema / properties / payload / items / required
        Added value: +[
        +  "summary"
        +]
      • addedInput schema / properties / run_id
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "The run to complete (from the fire callback); None = direct save"
        +}
      • changedInput schema / properties / status / description
        Previous value: -"One of \"gathering\", \"ready\", \"compiled\" (default: \"ready\")"New value: +"One of \"gathering\", \"ready\", \"compiled\", \"failed\" (default: \"ready\")"
      • changedInput schema / properties / status / enum
        Previous value: -[
        -  "gathering",
        -  "ready",
        -  "compiled"
        -]New value: +[
        +  "gathering",
        +  "ready",
        +  "compiled",
        +  "failed"
        +]
    • Addedtrigger_report_tool
    • Changedupsert_definition_tool2 fields changed
      • changedInput schema / properties / schedule / description
        Previous value: -"Descriptive cron/time metadata (the actual firing is external)"New value: +"Cron expression for when to fire (e.g. \"0 6 * * 5\"); None = manual only"
      • addedInput schema / properties / timezone
        Added value: +{
        +  "default": "UTC",
        +  "description": "IANA timezone the schedule runs in (default: \"UTC\")",
        +  "type": "string"
        +}
  2. 5 tool updatesv0.2.0
    • First observedget_output_tool
    • First observedget_spec_tool
    • First observedlist_reports_tool
    • First observedsave_output_tool
    • First observedupsert_definition_tool

TDQS

A3.9/5.0

Scored across 10 tools

Disambiguation4/5

The tools split cleanly along resource lines (definitions vs. outputs vs. sweep vs. trigger), and single-vs-bulk deletion (delete_output_tool vs. prune_outputs_tool) is clearly delineated. The only mild overlap is get_spec_tool versus list_reports_tool, since list_reports already returns the gather prompt and source config that get_spec_tool exists to expose.

Naming Consistency4/5

All ten tools follow a predictable verb_noun_tool pattern (get_output_tool, list_outputs_tool, prune_outputs_tool), which is easy to scan and reason about. Minor inconsistency in the noun chosen for the same entity: report definitions are called both 'definition' (upsert_definition_tool) and 'reports' (list_reports_tool).

Tool Count5/5

Ten tools is well-scoped for a report gatherer/compiler: definition management, output lifecycle, triggering, and sweep reading each get exactly the operations they need. No redundant or filler tools.

Completeness4/5

Output lifecycle is fully covered (save, list, get, delete, prune) and definitions support list/create/update plus read via get_spec_tool. The one clear gap is the absence of a delete/remove operation for report definitions, so stale reports cannot be retired through the tool surface.

Maintenance

ActivityActive
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables MCP agents to maintain durable, evidence-aware project knowledge, retrieve precise excerpts on demand, and track decisions, conflicts, and revisions across sessions.
    1
    Apache 2.0
  • A
    license
    C
    quality
    A
    maintenance
    Enables engineering agents to maintain persistent knowledge across sessions by storing decisions, invariants, gotchas, and rejected ideas, with staleness detection, conflict detection, full-text search, and structured context assembly.
    37
    1
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Provides persistent, searchable memory across coding projects and machines, letting agents record and retrieve projects, reusable assets, sessions, decisions, commits, and handoffs via MCP.
    24 PyPI
    2
    Apache 2.0