Skip to main content
Glama

Forge MCP — freelance pipeline operations

Version: 1.0.0
Quality gate: .\make.ps1 (black, ruff, mypy, pytest cov≥90%, Martin limits)

Forge hunts public, ToS-safe job APIs, ranks listings by capability match and money heat, pushes HOT cards (Telegram + local digest), and tracks the pipeline from NEW → paid. Client pitches stay in the listing language; operator briefs are in Russian.

What is in scope / out of scope

In scope

Out of scope

RemoteOK, Remotive, Himalayas, Arbeitnow(+UK), Jobicy, HN Who's Hiring

Upwork / Freelancer.com scraping

Manual capture_lead for off-board finds

Auto-apply without a human

Telegram buttons + rituals

Secret logging / token echo

Related MCP server: freelancehunt-mcp

Install

pip install -e ".[dev]"
copy .env.example .env   # fill BOT_TOKEN + CHAT_ID
.\make.ps1               # must exit 0

CLIs: forge-hunt, forge-telegram-bot, forge-morning, forge-evening, forge-mcp.

Configuration

Secrets and knobs live in .env (gitignored). See .env.example.

Variable

Role

BOT_TOKEN / CHAT_ID

Telegram (aliases → FORGE_TELEGRAM_*)

FORGE_CONTRACT_ONLY=1

Strict contract/freelance HOT gate

FORGE_ANCHOR_* / FORGE_WALK_AWAY

One killer offer + floor

FORGE_PROOF_URL

Optional case link appended to SHORT

FORGE_PAUSED_BOARDS

Manual board deprioritization

FORGE_DB / FORGE_DIGEST_PATH

Storage / digest overrides

Operations (money loop)

  1. set_watchlist([...]) — narrow magnets

  2. forge-hunt every ~30m — HOT + Telegram buttons

  3. Copy PITCH → send → Proposed (target <15 min NEW→proposed)

  4. forge-morning / forge-evening — queue + follow-ups + board pause

  5. Won + log_win_reason → learning / A/B / packages

Windows schedules: scripts\schedule_hunt.ps1, schedule_rituals.ps1, schedule_bot.ps1
(scripts resolve project root from their own location).

MCP

{
  "mcpServers": {
    "forge": {
      "command": "python",
      "args": ["-m", "forge.presentation.mcp_server"]
    }
  }
}

Core tools: hunt_jobs, act_now, morning_digest, evening_followup, draft_proposal, brief_lead, rate_advice, check_red_flags, follow_ups, sla_status, chase_leads, package_prices, pause_weak_boards, platform_roi, board_advice, win_journal, log_win_reason, pitch_ab_stats, set_watchlist, pipeline_status, earnings_summary, capture_lead, scan_for_leads, hot_leads.

Architecture

src/forge/
  entities/         # Lead FSM, match, heat, gates, pricing, i18n
  use_cases/        # hunt, rituals, pipeline, learning, act-now
  interfaces/       # ports (repos, sources, notifier)
  infrastructure/   # SQLite, HTTP boards, Telegram, .env loader
  presentation/     # MCP + CLIs + composition root

Clean architecture: use cases depend on ports; composition wires adapters. Persistence: SQLite (~/.forge/forge.db by default) with additive migrations.

Quality / acceptance

See docs/ACCEPTANCE.md. Summary:

  • .\make.ps1 exits 0

  • ≥90% line coverage on forge (CLIs/MCP omitted by design)

  • No secrets in repo; Telegram errors redacted

  • Martin limits: file ≤300 lines, fn ≤25 lines, ≤4 args, CC≤10, nesting≤3

License

MIT — see LICENSE.

Available Tools

31 tools
act_nowC

Priority queue of NEW leads (payability + age + rate).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, yet it never states that this is a read-only listing. The imperative name 'act_now' could be read as a mutating action, and the description does not resolve that ambiguity, nor mention auth, pagination, or side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler; the ranking criteria are compactly conveyed. It is terse to the point of being under-specified, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so result format need not be explained, but for a tool embedded among ~30 overlapping lead/outreach siblings the definition supplies no routing guidance and no limit semantics. It is too thin to let an agent confidently choose it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single 'limit' parameter has 0% schema description coverage and a bare default of 8, and the description says nothing about it (e.g., max queue size or truncation behavior). With one undocumented param, the description does not compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource ('priority queue of NEW leads') and even the scoring criteria (payability, age, rate), but it uses no verb and gives no hint of the operation performed (return? act on?). It also fails to differentiate from near-identical siblings such as hot_leads, chase_leads, and queue_outreach.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance: nothing says when to call act_now versus hot_leads, chase_leads, or follow_ups, all of which plausibly serve the same lead-prioritization need. The '(payability + age + rate)' parenthetical hints at ranking logic but not at selection conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

add_capabilityC

Add a reusable capability asset for future matching.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
keywordsYes
descriptionYes
proposal_snippetYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It implies a persistent write but says nothing about permissions, whether duplicates are rejected, whether the asset is immediately matchable, or what the response contains.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler and the action front-loaded. It is efficient, though efficiency here reflects under-specification as much as discipline.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. However, for a mutation tool with no annotations and four undocumented required parameters, the definition omits persistence, idempotency, and usage context that an agent needs to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across four required parameters, and the schema supplies only bare titles. The description's phrase "capability asset" loosely frames the fields but adds no meaning for keywords or proposal_snippet, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb+resource ("Add a reusable capability asset") is identifiable, but "for future matching" is vague about what the asset is or how it will be used. It gives no differentiation from the sibling suggest_capabilities, leaving the agent unsure which of the two to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to add a capability versus when to call suggest_capabilities or any other sibling. The agent must infer the context entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

board_adviceC

Deprioritize boards with paid=0 and enough leads.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not disclose whether the tool merely recommends deprioritization (as the name 'advice' suggests) or actually mutates board state, nor does it mention permissions, thresholds, or side effects. The 'paid=0' condition is the only behavioral hint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the selection criterion front-loaded and no filler. It is efficient, though the brevity edges toward under-specification rather than true conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, but the description omits essential context: what a 'board' is, what 'enough leads' means, and how this relates to 'pause_weak_boards' or 'rate_advice'. For a no-argument advisory tool sitting among many similar siblings, more routing and threshold context is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate. Baseline 4 applies since no parameter-level meaning is required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description pairs a verb ('Deprioritize') with a resource ('boards') and a condition ('paid=0 and enough leads'), so the general intent is discernible. However, it never defines what a 'board' is or whether this is advisory-only or an actual state change, which matters given the name 'board_advice' conflicts with the imperative 'Deprioritize'. It is not clearly distinguishable from the sibling 'pause_weak_boards'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to call this tool, no prerequisites, and no reference to alternatives. The sibling 'pause_weak_boards' appears closely related, yet the description offers no guidance on which to choose or in what sequence.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

brief_leadC

Explain a vacancy in clear Russian.

ParametersJSON Schema
NameRequiredDescriptionDefault
lead_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, yet it discloses almost nothing: no indication of read-only vs. mutating behavior, permission needs, or cost. The one genuine trait it does reveal is output localization to Russian, which is a small but real piece of context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler, front-loaded on the action. However, the brevity reflects under-specification rather than tight writing, so it does not earn above a middling score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and for a one-parameter read tool the core action is covered. What is missing is any situational context explaining how it fits alongside the many sibling lead-advice tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter lead_id is undocumented in both schema and description. The word 'vacancy' only loosely hints that lead_id identifies a job lead; no format, ID type, or source is given.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('explain') and resource ('vacancy'), so the basic action is inferable, but it never reconciles 'vacancy' with the tool name brief_lead or the lead_id parameter. It also fails to distinguish this from siblings like check_red_flags, rate_advice, or board_advice that plausibly also 'explain' something.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, no prerequisites, and no named alternative among the ~30 sibling tools. The only implicit guidance is that a lead_id is required, which the schema already enforces.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capture_leadC

Track a job lead. Dedupes by URL and title fingerprint.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYes
platformYes
source_urlNo
descriptionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose a genuinely useful non-obvious trait — deduplication by URL and title fingerprint — which tells the agent repeated calls won't create duplicates. But it says nothing about permissions, side effects, or failure behavior, so the disclosure is partial despite being relevant.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences with no filler. The purpose leads and the dedupe behavior follows. Nothing redundant, though the dedupe clause is doing double duty as the only behavioral disclosure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter mutation tool with 3 required inputs, no annotations, and 0% schema coverage, the description is too thin. An output schema exists so return values need not be explained, but required-field semantics and usage context are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description is the only source of parameter meaning. Mentioning URL and title fingerprint implicitly maps to source_url and title, but platform and description (both required) are entirely unexplained, leaving half the interface opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb and resource: 'Track a job lead.' An agent can tell this records a lead rather than listing or scanning for one (contrast scan_for_leads, list_outreach). It stops short of distinguishing itself from the many other lead-mutating siblings like update_lead_status or brief_lead.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No indication of when to call this versus the crowded set of lead-related siblings (update_lead_status, brief_lead, hot_leads, chase_leads). No prerequisites, no exclusions, no context about the intended workflow.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

chase_leadsC

NEW leads past SLA — propose these first.

ParametersJSON Schema
NameRequiredDescriptionDefault
stale_after_hoursNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the filtering criterion (new leads past SLA) but nothing about the action taken, whether anything is mutated, auth requirements, or how SLA is determined relative to stale_after_hours.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short, front-loaded line with no filler; the qualifier 'NEW ... past SLA' comes before the directive. It is arguably under-specified rather than padded, but every token does work.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, but for a tool with zero annotations and an undocumented parameter the description should say what 'chase/propose' entails and how staleness is computed. Those gaps are not covered anywhere else.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter stale_after_hours (default 2) has 0% schema description coverage and is never mentioned in the description. The description's 'past SLA' hints at a threshold but does not clarify whether it maps to stale_after_hours or is a separate SLA definition.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource and scope: NEW leads that are past SLA, which tells an agent what set it returns. However, the metaphorical verb 'chase' and the phrase 'propose these first' leave the actual operation vague, and it never distinguishes itself from siblings like hot_leads, sla_status, or follow_ups.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Propose these first' gives an implied priority directive, suggesting the agent should surface these ahead of other lead sets. It does not state when-not to use it, nor name a sibling alternative for a different priority or lead type.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_red_flagsC

Red-flag checklist before you propose.

ParametersJSON Schema
NameRequiredDescriptionDefault
lead_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden and delivers almost nothing: no indication of read-only vs mutating, no permission requirements, no note on what the check inspects or how severe a 'flag' is. The output schema covers return shape, but side-effect and safety behavior remain undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Seven words is not concision but under-specification: it is a noun phrase with no verb, no object of the check, and no stated output. Brevity here removes information the agent needs rather than trimming waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, but everything else is missing: no annotation coverage, an undocumented required parameter, and no statement of what the check evaluates or what the agent should do with the result. Inadequate for even a simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single lead_id parameter has 0% schema description coverage and the description adds no meaning at all - it does not say which lead record to pass, whether it accepts an ID versus a name, or what happens for an unknown lead. For a 1-param tool, some guidance is expected.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The fragment 'Red-flag checklist' plus the name check_red_flags identifies the resource, and 'before you propose' implies it evaluates a lead for risks ahead of a proposal. However, it never states what the tool actually does (compute flags? return a list? block a proposal?) and does not differentiate it from advice-oriented siblings like rate_advice or board_advice.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Before you propose' gives a clear timing cue that ties it to the proposal workflow (draft_proposal), which is more than most siblings offer. It stops short of naming an alternative tool or stating when not to run it, leaving the agent to infer the trigger condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

draft_proposalC

SHORT + FULL pitches with i18n opener, A/B variant, anchor price.

ParametersJSON Schema
NameRequiredDescriptionDefault
lead_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral-disclosure burden. It mentions output features like i18n openers and A/B variants, but says nothing about side effects, persistence, permissions, rate limits, or whether a proposal is created or merely drafted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short but reads as a cryptic fragment rather than a structured, front-loaded statement. Its brevity comes at the cost of clarity, leaving the core action and context under-specified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. However, with no annotations, a required but undocumented parameter, and no usage guidance, the description is not complete enough to ensure correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool requires one parameter, lead_id, and schema description coverage is 0%. The description does not mention lead_id or add any meaning beyond the bare schema field name, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names output artifacts—short and full pitches with i18n opener, A/B variant, and anchor price—but never states the action explicitly. It is more informative than a tautology, yet fragmented and not clearly distinguished from siblings like pitch_ab_stats or package_prices.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no named alternatives among the many sibling outreach and lead tools. The description leaves all routing decisions to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

earnings_summaryB

Show total won, paid, and outstanding amounts.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full behavioral burden and largely fails to meet it. It does not say whether this is read-only, what time window or scope it covers, whether it needs auth, or whether amounts are gross/net or filtered by anything.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. Every word contributes and nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. However, the description omits scope details (time period, filtering, per-workspace vs global) that materially affect interpretation of an earnings aggregate, leaving an agent guessing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema has nothing to document and there is no parameter semantics burden. Baseline 4 applies; the description mentions the metric set, which is the only meaningful content here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb (Show) and a specific set of metrics (total won, paid, outstanding amounts). It is distinguishable from sibling write tools like record_won_amount and record_payment, though it does not explicitly name them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No statement of when to use this versus alternatives such as pipeline_status or platform_roi, and no scope/timeframe conditions. The agent must infer that this is the aggregate earnings read from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evening_followupC

Evening: SLA + proposed pings + board pause advice.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It hints that the tool synthesizes SLA data, proposes pings, and advises on board pausing, but says nothing about whether it mutates state, what 'proposed pings' means operationally, or any permissions/rate considerations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The entire description is a single telegraphic fragment with no complete sentence, which is under-specification rather than true conciseness. It saves characters at the cost of clarity about what the tool actually returns or does.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the return shape itself need not be explained in prose. Still, for a compound digest tool with several implied sub-behaviors and no annotations, the description is too thin to tell the agent exactly what invoking it produces.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing to disambiguate; the baseline of 4 applies. The description neither adds nor detracts from parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names three concrete components (SLA, proposed pings, board pause advice) and the 'Evening:' prefix implies a periodic digest, which loosely distinguishes it from sibling morning_digest. However, it is a noun fragment with no verb, so an agent must infer that it aggregates/returns these items rather than performing an action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites, and no reference to any of the many siblings such as morning_digest, sla_status, board_advice, or pause_weak_boards. The agent is left to guess the timing and trigger conditions from the word 'Evening' alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

follow_upsD

Proposed leads silent too long — copy the EN ping.

ParametersJSON Schema
NameRequiredDescriptionDefault
after_hoursNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and delivers nothing. It never states what the tool actually does (send? draft? queue?), whether it mutates lead state, what permissions are needed, or what 'copy the EN ping' implies operationally.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The text is short but telegraphic rather than concise: the fragments are not self-explanatory and 'copy the EN ping' spends characters on obscure internal shorthand instead of stating the action. Brevity here reflects under-specification, not efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, but everything else is missing: no annotations, no parameter meaning at 0% coverage, and no indication of the action or its side effects. For a lead-mutating follow-up tool this is wholly inadequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the sole parameter 'after_hours' (default 48) is never mentioned in the description. Its meaning — a threshold, a quiet-hours window, or something else — is left entirely to guesswork.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Proposed leads silent too long' hints at a follow-up action on stale leads, but the verb is never stated and 'copy the EN ping' is unexplained jargon. It cannot be distinguished from siblings like chase_leads or evening_followup, which appear to cover near-identical ground.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of alternatives such as chase_leads or queue_outreach. The agent has no basis for choosing this tool over its many similar siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hot_leadsA

List fresh (status=new) leads waiting for your proposal.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral burden. 'List' strongly implies a read-only query, and the status=new filter is useful context, but the description omits any explicit side-effect, permission, pagination, or ordering behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single front-loaded sentence with no filler. The key qualifier (status=new) and purpose (waiting for proposal) are both immediately visible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter list tool with an output schema, the description is nearly complete: it identifies the resource and the relevant status filter. A minor gap remains in not clarifying how this list relates to other lead-oriented siblings such as list_outreach or chase_leads.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has 0 parameters, so there is no parameter semantics to document. The baseline score for zero-parameter tools is 4, and the description appropriately focuses on the tool's filtering semantics instead.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (List) and resource (leads), plus a clear scope filter (status=new, waiting for your proposal). It distinguishes this from generic lead lists through the proposal-ready stage, but it does not explicitly name or contrast a sibling tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'waiting for your proposal' implies the tool should be used when the agent needs new leads ready for proposal outreach. However, it gives no explicit when-not guidance and does not route to alternatives among the many sibling lead-handling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hunt_jobsD

Hunt boards; push freelance-only HOT with RU brief + EN short pitch.

ParametersJSON Schema
NameRequiredDescriptionDefault
min_heatNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. The verb "push" implies a write/side-effecting action (possibly outreach or queueing), yet the description does not state what is pushed, where, whether it is reversible, or what permissions are required. "Freelance-only" and the dual-language output are the only real behavioral disclosures.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single, front-loaded sentence with no filler, so it is structurally efficient. However, the density of unexplained abbreviations (HOT, RU, EN) trades clarity for brevity, which is a conciseness failure of a different kind. Adequate length, poor legibility.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, but for a pipeline-style tool with no annotations and an undocumented parameter the description is far from complete. It omits selection criteria, side effects, and how the "push" interacts with the sibling outreach tools. An agent could not invoke this correctly without guessing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the single parameter min_heat, and it does not mention it at all. Defaults and range for min_heat are left entirely to the raw schema. Only one parameter, but the description adds zero semantic value over the schema type.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase "Hunt boards" largely restates the tool name, and the second clause is a compressed jargon string ("HOT", "RU brief", "EN short pitch") whose meaning is not self-evident. Siblings like scan_for_leads, hot_leads, and chase_leads overlap heavily, and nothing here distinguishes hunt_jobs from them. An agent cannot confidently tell what this tool does that the siblings do not.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool, when not to, or which sibling to prefer for adjacent tasks (scanning, chasing, drafting). No prerequisites, cadence, or trigger conditions are given. The only hint is the word "freelance-only", which filters scope but not usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_outreachB

List pending outreach drafts.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full behavioral burden. 'List' weakly implies a non-mutating read, but nothing states that explicitly, nor covers pagination, ordering, or permission requirements. It adds essentially no behavioral context beyond the verb.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the resource and scope front-loaded and zero filler. Nothing in the text is redundant with the title or name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is parameterless and has an output schema, so return values need not be explained, keeping the bar low. Even so, with no annotations and no routing against its many siblings, the definition is only minimally adequate for correct selection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter meaning to convey and the baseline of 4 applies. The description correctly avoids inventing parameters that do not exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with a scope qualifier ('List pending outreach drafts'), so the agent knows exactly what it returns. However, it makes no attempt to distinguish itself from near-neighbours like queue_outreach, follow_ups, or chase_leads, leaving sibling differentiation to guesswork.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus the many other outreach-oriented siblings (queue_outreach, follow_ups, chase_leads). The single sentence names no condition, prerequisite, or alternative, so the agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

log_win_reasonC

Win journal: one phrase why the client chose you.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonYes
lead_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden and adds essentially nothing: it does not say the entry is persisted, whether it is append-only or overwrites a prior reason, whether it requires an existing lead, or what it returns. For a mutation tool with zero annotation coverage this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler, and the essential qualifier ('one phrase') is front-loaded. It is appropriately sized for the operation, though its brevity is partly under-specification rather than deliberate economy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, but for a two-param mutation with no annotations the description omits usage context, lead_id semantics, and any behavioral disclosure. It is too thin to let an agent invoke it with confidence.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. 'One phrase why the client chose you' usefully constrains the reason parameter to a short free-text phrase, but lead_id is left completely unexplained and its relationship to capture_lead is unstated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Win journal: one phrase why the client chose you' conveys the intent of recording the reason a deal was won, so an agent can roughly infer the resource. However, the action verb is only implied by the tool name, and it does not distinguish itself from nearby siblings like record_won_amount or win_journal.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus record_won_amount (which captures the deal value) or win_journal, nor any mention of prerequisites such as the lead already being marked won. The agent must guess the trigger condition entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

morning_digestC

Push top NEW with RU brief + SHORT EN (Telegram buttons).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose that output is delivered to Telegram with interactive buttons (a real external side effect), but says nothing about whether this sends messages repeatedly, requires a configured Telegram channel, is idempotent, or costs anything.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short sentence with no padding, so nothing is wasted. However, the extreme telegraphic compression (RU/SHORT EN) trades away clarity rather than achieving efficient structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and with zero parameters the schema is trivial. But with no annotations and an undecoded description, an agent cannot infer what 'top NEW' means, what the brief contains, or what side effects occur.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to explain and the baseline of 4 applies. No parameter-related gaps exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb 'Push' is identifiable, but the resource is compressed into jargon: 'top NEW', 'RU brief', 'SHORT EN' are not decoded anywhere. It is close to restating the name (morning_digest) with unexplained abbreviations, and it does not distinguish itself from siblings such as evening_followup.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is given. The 'morning' in the name and in the description hints at a daily cadence, but there is no statement of prerequisites, no contrast with evening_followup or brief_lead, and no conditions under which the agent should or should not call it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

package_pricesC

Discovery / milestone / retainer from average won + anchor.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the full behavioral burden, and it discloses almost nothing — not whether this is a pure computation or a mutation, whether it depends on previously recorded wins, or what happens when no 'average won' exists. It hints at inputs but never states the side-effect or precondition profile.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but brevity here is under-specification rather than economy: a bare noun list with a prepositional modifier, no front-loaded statement of purpose. Every word 'counts' but the sentence never lands as a claim about the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, but for a pricing-generation tool the description never explains what it computes or when to reach for it. The combination of zero annotations and a cryptic fragment leaves the agent materially under-informed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema baseline of 4 applies. The phrase 'from average won + anchor' usefully signals that the inputs are derived from recorded data rather than passed in, which is more than the empty schema conveys.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The fragment names package types (discovery / milestone / retainer) and two inputs (average won + anchor), so the subject matter is inferable, but there is no verb and no statement of what the tool actually produces. An agent cannot confidently tell this apart from rate_advice, board_advice, or earnings_summary without guessing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Discovery / milestone / retainer from average won + anchor' gives no when-to-use, no prerequisites, and no alternative. Several siblings (rate_advice, board_advice, pipeline_status) could plausibly cover adjacent pricing questions, and nothing routes the agent between them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pause_weak_boardsB

Auto-pause boards with paid=0 and enough leads.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full burden. 'Auto-pause' is a state-changing mutation, yet the description never says whether pausing is reversible, what permissions it needs, what 'enough leads' thresholds, or what happens to already-paused boards.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the condition stated up front and no filler. It is terse to the point of underspecification, but not padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no explanation, but for an unannotated bulk-mutation tool the description omits too much: what qualifies as 'weak', how many boards are affected, and whether the action can be undone. An agent cannot predict the blast radius before calling it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing parameter-level for the description to explain; the baseline is 4. The criteria wording ('paid=0', 'enough leads') hints at internal filters but does not correspond to any caller-supplied input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb+resource ('Auto-pause boards') and adds the selection criteria (paid=0 and enough leads), so an agent can tell it apart from advisory siblings like board_advice or package_prices. It is clear but the criteria terms are undefined.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies the trigger condition (weak boards) but gives no explicit when-to-use guidance, no mention of alternatives, and no prerequisites for invoking it. An agent must infer all of this from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pipeline_statusB

Show pipeline counts and win rate.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing about scope, data freshness, time window, or permissions. For a metrics tool with zero annotation coverage, this is a real gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the key output (counts and win rate) front-loaded. No waste, though it is arguably terse to the point of under-specification.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and with no parameters the surface is simple. Still, the description omits what 'pipeline' scopes over and any time-window or filtering assumption, leaving the agent to guess the metric's boundaries.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema defines zero parameters, so there are no parameter semantics to clarify; the baseline for a parameterless tool applies. Nothing in the description is needed to compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete verb ('Show') and a specific resource ('pipeline counts and win rate'), so the agent knows it returns aggregate pipeline metrics. It does not, however, differentiate itself from nearby siblings such as earnings_summary or platform_roi, which likely also surface aggregate numbers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to reach for this tool versus the many overlapping analytics siblings (earnings_summary, platform_roi, sla_status, win_journal). The agent is left to infer that this is the general pipeline-health view.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pitch_ab_statsB

A/B win rates for short vs full pitches.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden, and it discloses nothing beyond the metric itself — no data window, no sample-size or minimum-volume caveats, no indication of what counts as a 'short' vs 'full' pitch, and no statement of read-only/non-mutating behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded clause with no filler. Nothing can be trimmed without losing the metric definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and zero parameters keeps the surface small. However, with no annotations and no usage or behavioral context, the definition is only minimally viable — it never says what a 'short' or 'full' pitch means or over what period the win rates are computed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4; the description reasonably implies the comparison dimension (pitch length) is baked in rather than parameterized.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific metric and scope: win rates comparing short vs full pitches. The agent can infer this returns an A/B comparison statistic, though no verb (returns/computes/compares) is given and no sibling is named for differentiation — none of the listed siblings are obviously confusable, so the gap is modest.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no trigger conditions, and no mention of alternatives. The agent must guess whether this is a reporting tool to run proactively, a diagnostic to run after a batch of pitches, or something gated on prior activity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

platform_roiB

Show which job boards actually convert to paid work.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it says nothing about whether this is a read-only analysis, what time window it covers, or whether it requires existing paid-work data to be meaningful. Only the vague qualifier 'actually' hints at real conversion data rather than projections.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the intent lands immediately. Nothing in it is redundant given the sparse structured metadata.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and there are no parameters to document. However, with no annotations and no mention of scope, timeframe, or data prerequisites, the description is thin for an analytics tool whose interpretation depends on context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline of 4 applies. No parameter-related guidance is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (show) and a concrete subject: which job boards convert to paid work. This is a distinct analytical purpose, distinguishable from siblings like board_advice or pause_weak_boards, though it never explicitly names those alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to reach for this tool versus board_advice, pause_weak_boards, or earnings_summary, and no stated prerequisites or cadence. The agent must infer usage from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

queue_outreachC

Queue Merchant/MCP outreach draft → Telegram Send/Skip buttons.

ParametersJSON Schema
NameRequiredDescriptionDefault
angleNomerchant
emailYes
companyYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It hints at a side effect (a Telegram approval message with Send/Skip buttons) but omits whether the email is sent immediately, what Send/Skip actually triggers, permission/auth requirements, or any rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single terse clause with no filler, front-loading the action and the resulting UI. It is efficient though cryptic in its jargon.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, but for a mutation tool with zero annotation coverage and three undocumented parameters, the description leaves too much unstated: no usage context, no parameter meaning, and no behavioral detail beyond the Telegram buttons.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 3 parameters, so the description must compensate and largely does not. Only 'angle' is faintly implied by 'Merchant/MCP'; 'email' and 'company' (both required) receive no explanation of expected format or role.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb 'Queue' and resource 'outreach draft' are identifiable, and the Telegram Send/Skip delivery is mentioned. However, 'Merchant/MCP' is unexplained jargon and the description does not distinguish this from siblings like draft_proposal, list_outreach, or capture_lead, which also touch outreach.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus alternatives, no prerequisites, and no exclusions. The agent is left to infer that it should call this after identifying a lead, but nothing in the text says so.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rate_adviceD

Your avg won vs listing budget mid.

ParametersJSON Schema
NameRequiredDescriptionDefault
lead_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

D1.6/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, yet it discloses nothing about whether the tool reads or mutates data, permission needs, or what the advice consists of. Only a garbled phrase is provided.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short, but the brevity reflects under-specification rather than efficient front-loading. The single fragment is not a coherent sentence and fails to state the action.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, but the description still leaves the core action, the lead_id input, and the use case entirely unclear. For a tool with a required lead_id, this is inadequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter (lead_id) is never mentioned in the description. The phrase 'Your avg won' only weakly implies per-user/per-lead data, so the description fails to compensate for the undocumented parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The fragment 'Your avg won vs listing budget mid' hints at a comparison between average won amounts and a listing budget midpoint, but gives no clear verb or resource. It is neither a clean statement of purpose nor distinguishable from siblings like board_advice, pitch_ab_stats, or earnings_summary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to use this tool, when not to, or which sibling (board_advice, pitch_ab_stats, earnings_summary) it replaces. The agent is left to guess based on a cryptic fragment.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_paymentC

Record that you've actually been paid for a lead.

ParametersJSON Schema
NameRequiredDescriptionDefault
amountYes
lead_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It implies a mutation ('record') but says nothing about permissions, idempotency or duplicate-payment handling, or how this affects downstream aggregates like earnings_summary.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler and the action front-loaded, but its terseness is under-specification rather than efficiency given the total absence of behavioral and parameter detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be described, but for a mutation tool with no annotations and zero parameter documentation the single sentence leaves too much unstated about inputs and side effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two required parameters, and the description never mentions lead_id or amount. The agent must infer that amount is a currency value and lead_id identifies the paying lead purely from the property names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource ('Record ... payment') scoped to a lead, so the agent knows the core action. It does not, however, distinguish itself from nearby siblings like record_won_amount or log_win_reason, which an agent could easily confuse with this.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Actually been paid' hints at a real-payment trigger, but there is no explicit when-to-use guidance, no prerequisite conditions, and no routing to or away from alternatives such as record_won_amount or earnings_summary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_won_amountC

Record the agreed price once you win a lead.

ParametersJSON Schema
NameRequiredDescriptionDefault
amountYes
lead_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It never says whether the amount overwrites a previously recorded value, whether the write is idempotent, what permissions are needed, or how it interacts with the lead's status; for a mutation tool this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler and no repetition of the tool name. It is efficient, though its brevity borders on under-specification rather than true conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. However, with zero annotations, zero parameter documentation, and a mutation of a CRM record, the description leaves out the write semantics, the relationship to record_payment/log_win_reason, and any preconditions an agent would need.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the definition must compensate. "The agreed price" loosely maps to the amount parameter, but lead_id is left entirely undefined (is it the lead's ID, a slug, an external reference?) and no format, currency, or unit guidance is given for the numeric amount.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (record the agreed price) tied to a named trigger (winning a lead), so the action is unambiguous. It does not, however, differentiate itself from near-neighbors like record_payment or log_win_reason, which an agent must disambiguate on its own.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Once you win a lead" gives a clear triggering condition, which is more than most siblings offer. But it names no alternatives and gives no exclusion — e.g. when record_payment or log_win_reason would be the correct call instead — leaving the routing decision implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scan_for_leadsB

Scan public job boards and auto-track capability matches.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It implies both a read (scan boards) and a write (auto-track matches), yet says nothing about whether leads are created or deduplicated, what 'capability matches' are compared against, or any rate limits — the write side-effect is the most important thing to disclose and it is left implicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. It is tight, though 'auto-track capability matches' is jargon that could be replaced with plainer language at no length cost.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and there are no parameters to document. However, for a zero-config tool with no annotations and a crowded sibling set, the definition needs at least a pointer on how it differs from hunt_jobs and what 'auto-track' mutates — that gap keeps it at minimum viable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate and the baseline of 4 applies. No compensating detail is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Scan public job boards') plus a secondary effect ('auto-track capability matches'), so the agent knows this reads external boards and writes matches. It does not distinguish itself from close siblings like hunt_jobs or chase_leads, which leaves an ambiguity an agent must resolve elsewhere.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to invoke this versus hunt_jobs, hot_leads, or chase_leads, nor any stated preconditions or frequency limits. 'Auto-track' hints at a downstream workflow but never says when or why to trigger the scan.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_watchlistC

Set personal hunt keywords (money magnets).

ParametersJSON Schema
NameRequiredDescriptionDefault
keywordsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and discloses almost nothing. It does not say whether 'set' replaces or appends to existing keywords, whether it persists, or whether any permissions are required. 'Money magnets' adds flavor but no operational meaning.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence that front-loads the verb, with no filler. The unusual 'money magnets' phrasing slightly muddies an otherwise efficient statement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, but critical behavior for a preference-setting tool — replace vs merge semantics and persistence — is absent, and no annotations fill the gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single required parameter 'keywords' is only restated as 'hunt keywords (money magnets)'. No format, count limits, or expected values are given, so the description does not compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb 'Set' plus resource 'watchlist keywords' gives a rough sense of the operation, but the phrase 'money magnets' is unexplained jargon and there is no differentiation from siblings like hunt_jobs or scan_for_leads. An agent can guess this configures search terms, but the actual effect is vague.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to call this tool or how it relates to sibling tools such as hunt_jobs or scan_for_leads. The agent must infer the trigger condition entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sla_statusC

SLA: burning NEW leads and proposed needing a 48h ping.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It gives a vague data scope but says nothing about whether this is a read-only status check, what permissions are needed, what triggers the SLA, or how results are refreshed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is brief with no wasted words, but it is under-specified rather than effectively concise. It is front-loaded as a cryptic fragment and does not provide enough structure for an agent to act on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although an output schema exists and need not be explained, the description still fails to define the tool's purpose, usage context, or SLA mechanics. For a zero-parameter status tool surrounded by many status-oriented siblings, this leaves too much ambiguous.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero input parameters, so the baseline is 4 per the scoring rules. There are no parameter semantics for the description to add or omit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names an SLA concept and two lead conditions, but it does not state what the tool actually does. It reads as a label for a reporting scope rather than a clear verb+resource definition, so an agent cannot confidently distinguish it from siblings like pipeline_status or follow_ups.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, no when-not-to-use guidance, and no named alternative among the many sibling tools. The phrase about burning NEW leads and proposed 48h pings hints at context, but it does not tell the agent when this tool should be selected.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

suggest_capabilitiesC

Rank your capability assets by relevance to a lead.

ParametersJSON Schema
NameRequiredDescriptionDefault
lead_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. "Rank" implies a non-mutating recommendation operation, but the description never confirms it is read-only, says nothing about what state or permissions are required, and gives no sense of what the ranking output contains despite an output schema existing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the action and the criterion appear immediately. It is arguably too terse for the ambiguity it sits in, but nothing in it is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is low-complexity (one required parameter) and an output schema exists, so return values need not be explained. Still, against a large sibling set that includes add_capability, the description omits any disambiguation or usage context, leaving real gaps for correct tool selection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single required lead_id parameter, so the description must compensate. It partially does by describing the output framing ("relevance to a lead"), which implies lead_id anchors the ranking, but it adds no format, source, or constraint details for the parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ("Rank your capability assets") and states the ranking criterion ("by relevance to a lead"), which is clear and unambiguous. It does not, however, distinguish itself from the sibling add_capability, so an agent must infer that this reads/ranks rather than mutates.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus alternatives such as add_capability or draft_proposal, nor any stated prerequisites (e.g., that capability assets must already exist). The agent is left to infer the usage context entirely from the one-line purpose.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_lead_statusC

Move a lead through the pipeline; wins feed learned magnets.

ParametersJSON Schema
NameRequiredDescriptionDefault
statusYes
lead_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It hints that setting a 'win' status has a side effect ("feed learned magnets"), but this is unexplained and the description omits permission needs, validation behavior, reversibility, and what happens to the lead record.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is a single short, front-loaded sentence with no padding, which is good. However, the second clause is cryptic rather than informative, so brevity comes at the cost of clarity instead of earning its keep.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained. But for a mutation tool with zero annotations and two completely undocumented parameters (including an enum-less status field), the description is far too thin to let an agent invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and neither parameter is documented, so the description must compensate — it does not. The crucial 'status' parameter has no enum and no listed valid values, so an agent cannot know what strings to pass, and the phrase "move through the pipeline" only vaguely implies status is a stage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Move a lead through the pipeline" conveys a mutating action on a lead resource, so the core purpose is inferable. But the second clause "wins feed learned magnets" is opaque jargon that adds confusion rather than clarity, and nothing distinguishes this from siblings like capture_lead or chase_leads.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no indication of when to call this versus the many adjacent lead tools (capture_lead, chase_leads, brief_lead, hot_leads). No prerequisites, no context about which status transitions are valid or when a status change should be initiated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

win_journalC

Recent wins with amounts, pitch variant, and reasons.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure and fails to meet it. It never states that this is a read-only retrieval (the word 'journal' could be misread as a logging/write action), nor does it mention time window, result caps, or ordering beyond the vague word 'Recent'. Only the content of the returned fields is conveyed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single nine-word fragment with no filler and the key information ('recent wins') front-loaded. It is appropriately sized for a zero-argument tool, though the missing verb makes it feel like a label rather than an instruction.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because an output schema exists, the description need not explain return values, and with no parameters the input side is trivial. What remains missing is the read-vs-write disambiguation against log_win_reason/record_won_amount and any notion of the 'recent' window, which for a listing tool is the main open question.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has zero parameters, so per the rubric the baseline is 4 and there is nothing extra to document. The description appropriately spends its words on output content instead of nonexistent inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase names a resource (recent wins) and the fields it carries (amounts, pitch variant, reasons), so an agent can guess it is a read/listing tool. But there is no verb at all, and the name/description does not distinguish it from siblings like log_win_reason or record_won_amount, which are the write-side counterparts. Purpose is implied rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use, when-not-to-use, or alternative is given. With siblings named log_win_reason and record_won_amount sitting right next to it, some routing guidance (this reads past wins vs. those record new ones) would be essential. The agent must infer the distinction entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 31 tool updatesv1.0.0
    • First observedact_now
    • First observedadd_capability
    • First observedboard_advice
    • First observedbrief_lead
    • First observedcapture_lead
    • First observedchase_leads
    • First observedcheck_red_flags
    • First observeddraft_proposal
    • First observedearnings_summary
    • First observedevening_followup
    • First observedfollow_ups
    • First observedhot_leads
    • First observedhunt_jobs
    • First observedlist_outreach
    • First observedlog_win_reason
    • First observedmorning_digest
    • First observedpackage_prices
    • First observedpause_weak_boards
    • First observedpipeline_status
    • First observedpitch_ab_stats
    • First observedplatform_roi
    • First observedqueue_outreach
    • First observedrate_advice
    • First observedrecord_payment
    • First observedrecord_won_amount
    • First observedscan_for_leads
    • First observedset_watchlist
    • First observedsla_status
    • First observedsuggest_capabilities
    • First observedupdate_lead_status
    • First observedwin_journal

TDQS

C2.3/5.0

Scored across 31 tools

Disambiguation2/5

Several tool clusters overlap heavily: act_now/hot_leads/chase_leads all surface fresh NEW leads to act on, scan_for_leads/hunt_jobs both scan boards for leads, and pause_weak_boards/board_advice both deprioritize paid=0 boards. Descriptions hint at nuances (age, SLA, push vs track) but an agent could easily misselect within these pairs.

Naming Consistency3/5

Casing is consistently snake_case, but the verb/noun pattern is mixed: some tools are verb_noun (capture_lead, record_payment, draft_proposal) while many others are bare noun phrases (pipeline_status, hot_leads, follow_ups, win_journal, rate_advice). Readable but not a predictable convention.

Tool Count2/5

31 tools is heavy for a lead-hunting/pipeline server, and the count is inflated by overlapping capabilities (multiple lead-surfacing, board-advice, and outreach tools) rather than distinct operations. Bloat suggests several tools could be merged or parameterized.

Completeness4/5

The lifecycle is well covered: capture/scan leads, qualify (red flags, pricing, rate advice), outreach/proposals, pipeline status, payments, board ROI, and win analytics. Only minor gaps (e.g., no explicit lead delete/archive or contact management) that agents can work around.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    An AI job-hunt copilot that enables searching live job boards, shortlisting openings, tracking application pipelines, and generating tailored resumes and cover letters from any MCP client.
    14
    Apache 2.0
  • A
    license
    B
    quality
    C
    maintenance
    Enables MCP clients to search and manage Freelancehunt projects, bids, profiles, reviews, messages, and workspaces, with local filtering by keywords, budget, bid count, and employer quality.
    31
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables AI agents to discover Gibwork bounties, verify their Solana escrow accounts on-chain, rank opportunities by value, competition and deadline, diff captures over time, and export Markdown/CSV/JSON reports through four read-only MCP tools requiring no wallet or API key.
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables tracking a job-search pipeline through MCP tools, a resource, and a prompt, providing stale-application alerts, weekly reviews, and gated follow-up drafting with human approval.
    MIT