Skip to main content
Glama
timmKal01

ke-tenders-audit-mcp

by timmKal01

Ke-Tenders Audit

An agentic AI that prepares a sourced review file on Kenyan public procurement awards for a human evaluation committee, then stops. It flags patterns. It never awards a tender, never recommends a winner and never declares wrongdoing. Every write needs a named human's approval.

Built for the African Agentic AI Design Challenge, Governance track (The Bid Box Challenge), theme Value for money.

Live demo: https://ke-tenders-audit-bxxfc4gmxjc9mpuuvkzypz.streamlit.app/ (opens on recorded real runs; switch to "Live review" in the sidebar for up to 3 live reviews a day on a free model tier)

pip install -r requirements.txt && streamlit run app.py

(Python 3.12. Put a model key in .env first: see Setup.)


Problem statement

Kenyan public bodies (county governments, TVET colleges, schools, NG-CDF offices, water companies) publish every tender and award on the PPRA portal, tenders.go.ke, in the Open Contracting Data Standard (OCDS). Almost nobody has time to use it. Evaluation committees and internal auditors spend weeks checking bidders, comparing prices and writing reports; the challenge brief puts it at 26 days of committee work per tender. Oversight bodies face the opposite problem: thousands of records and no way to see which ones deserve a second look.

The warning signs are already in the data. In the FY2026/27 release (4,830 tenders, 136 with awards), 23 awards had a single bidder, some suppliers win repeatedly from the same buyer, some contracts are signed before the tender closed, and some awards sit just below round amounts.

Related MCP server: Tender MCP

Which half of the data this is built on

The brief notes that portals publish what was tendered and awarded, not the bids. This project is built on the tendered-and-awarded half: award values, bidder lists on awarded tenders, dates, methods and buyers. It does not evaluate bid documents, because none are published.

Solution overview

A reviewer asks a question such as "Review all published awards by PC Kinyanjui Technical Training Institute". The agent plans the review, calls tools on our MCP server to search awards, run integrity checks, compare prices and profile suppliers, then proposes flags. Each proposed flag pauses the run until a named reviewer approves, rejects or requests changes. Approved flags go into a case file, and the agent drafts a committee report where every finding cites its OCDS record.

What the code enforces, not just the prompt:

  • A flag must cite an ocid that a tool actually returned in this run, or it is refused before a human sees it.

  • A flag naming a company that is not a bidder or winner on that record is refused, with the correct name.

  • Only the human gate can fill approved_by; the model cannot approve itself.

  • A write the reviewer rejected is refused if the model proposes it again.

  • Wording that decides an award or declares guilt is refused.

  • Every report ends with a table of all patterns found, computed from the data, so a wrong model summary cannot hide anything.

  • No em dashes or en dashes in anything written to the case file.

Target users

  • Procurement evaluation committees and heads of procurement in counties, colleges and state agencies.

  • Internal audit units preparing for the Office of the Auditor-General.

  • Also useful to oversight bodies, journalists and civil society, and to honest suppliers who gain from fairer competition.

Architecture

Architecture

One page: ARCHITECTURE.md.

Setup

Requirements: Python 3.12, Node.js 18+ (for the borrowed Filesystem MCP server, run with npx).

  1. Copy .env.example to .env and choose a model:

    • Hosted Qwen on Groq (free key at console.groq.com): set KTA_LLM_BASE_URL=https://api.groq.com/openai/v1, KTA_LLM_MODEL=qwen/qwen3.8-27b and your key in KTA_LLM_API_KEY.

    • Fully local with Ollama (data never leaves the machine): install Ollama, run ollama pull qwen3:4b, and keep the default http://localhost:11434/v1 settings.

  2. Install and start:

pip install -r requirements.txt && streamlit run app.py

The cleaned data ships in data/clean/releases.jsonl, so no download is needed. To rebuild it from a fresh PPRA download, save the JSON from https://tenders.go.ke/api/ocds/tenders?fy=2026-2027 into data/raw/ and run ke-tenders-ingest.

Usage

Review screen: enter your full name, ask a question, watch each tool call with its input and output, then approve, reject or request changes on each proposed flag. Each proposal shows the record's own facts (title, winner, amount, bidders, method) beside the model's claim, with a red warning on any mismatch. Download the report when done.

Terminal:

python -m ke_tenders_audit.cli "Review Kilifi County Government's awards." --reviewer "Your Full Name"

Tests and evals:

pytest
python evals/run_evals.py --runs 2

Technology stack

Layer

Choice

Orchestration

LangGraph (state graph, interrupt() for the human gate, SQLite checkpointer)

Model

Qwen3.8-27B (open weights) on Groq; Qwen3 4B locally through Ollama; any OpenAI-compatible endpoint

MCP

Official MCP Python SDK (our server), langchain-mcp-adapters (client), official Filesystem MCP server (borrowed)

Data

DuckDB in memory, built from cleaned OCDS JSONL; RapidFuzz for name matching

Interface

Streamlit review screen, plus a terminal runner

Logs

Append-only JSONL audit log per run

Tests

pytest (31 tests), plus a graded eval harness

Agent architecture

plan -> agent -> read tools -> tools -> verify -> agent ...
             -> write tools -> grounding check -> human gate -> tools -> verify -> agent ...
             -> no tool calls -> done
  • plan: the model writes a short numbered plan for the request.

  • agent: the model chooses tools, at most 3 per step.

  • verify: after each tool round, deterministic checks look for errors, refusals, hints and "insufficient comparables", and tell the model plainly so it re-plans instead of guessing.

  • grounding check and human gate: described above.

  • memory: LangGraph checkpoints in SQLite, so a run pauses for approval and resumes exactly where it stopped, even after an error ("Retry last step").

  • budget: older tool results are shortened to keep each request under about 7,000 tokens, which fits free model tiers. Red-flag results put their totals first so shortening never loses them.

MCP implementation

Two MCP servers, both connected over stdio through langchain-mcp-adapters. A tool interceptor writes every call (server, tool, inputs, output, duration) to the audit log.

MCP tools and servers

Built: ke-tenders-audit-mcp (server.py)

Tool

Kind

What it does

search_awards

read

Find awards by buyer, item, supplier, category, method or date. Explains empty results and suggests close buyer names.

check_red_flags

read

Single bidder, direct method, signed before close, just below a round amount, repeat winner at the same buyer. Totals first.

price_benchmark

read

Compare an award with similar awards (same item group and category). Says so when there are fewer than 5.

supplier_profile

read

Bids, wins, win rate, buyers and single-bidder wins for a company.

file_flag

write, gated

Add a sourced flag to the case file, with the record's own facts attached.

draft_report

write, gated

Build the committee report from approved flags, plus the data-computed totals table.

Resources: ocds://release/{ocid} (the cleaned record behind any finding) and ocds://data-quality-notes (known problems in the feed).

Borrowed: the official Filesystem MCP server (@modelcontextprotocol/server-filesystem), scoped to the cases/ folder, read tools only. It lets the agent read earlier case files. It is maintained, tested and sandboxed to one folder, so writing our own file access would add risk and no value.

Human-in-the-loop workflow

  1. The agent proposes one or more writes. The run pauses (LangGraph interrupt()).

  2. The reviewer sees each proposal with the record's facts and the full source record.

  3. For each one: Approve, Request changes (with a note the agent uses to revise) or Reject (never proposed again in this run).

  4. The gate writes the reviewer's name into approved_by. The audit log records the decision.

  5. The final report says who approved it. The committee makes the award decision, outside this system.

Data conduct

Open data only, from PPRA. Cleaning removes contactPoint, street addresses and postal codes, and strips emails and phone numbers that publishers typed into company names. Company records are used; no individual's personal details are. No live tender in progress is evaluated: awards are by definition closed.

Cost per run

Measured on Groq, Qwen3.8-27B at $0.80 and $4.00 per million input and output tokens: a full buyer review (31 awards, 7 model calls, 13 tool calls) used about 24,000 input tokens, about $0.03. A single-tender review costs under $0.02. Locally with Ollama the cost is zero. Every run's tokens and cost are in its audit log. Evaluation results: EVALS.md.

Limitations

  • No bid documents. The feed has no bids, so technical and financial evaluation of bids is out of scope.

  • Lump sums, not unit prices. Price comparisons point to scope questions, not proven overpricing.

  • Small award sample. Only 136 tenders have published awards in FY2026/27, so many items have too few comparables.

  • Item grouping is rule-based. Free-text titles are grouped with keyword rules; about 7% of awards stay ungrouped.

  • Thresholds are round numbers, not confirmed law. "Just below a round amount" needs checking against the PPADA Regulations 2020.

  • Free model tier. On Groq's free tier a review takes 1 to 3 minutes, mostly rate-limit waiting, and about 6 full reviews fit in a day.

  • The model's summary can still be wrong. The data-computed table at the end of each report is the safeguard.

Future improvements

  • Load past fiscal years (FY2025/26 and earlier) for stronger price comparisons and supplier histories.

  • Run against other OCDS portals (Rwanda, Tanzania, Nigeria, South Africa); the tools are schema-generic.

  • Confirm PPADA thresholds and replace round amounts with the real limits per procurement method.

  • Match suppliers on registration numbers if PPRA publishes them.

  • Unit-price extraction from contract documents, where those become public.

Licence

MIT. See LICENSE. Data: Public Procurement Regulatory Authority (Kenya), published as open data.

Available Tools

6 tools
check_red_flagsB

Integrity checks for one tender (ocid) or all awards of a buyer.

Single bidder, direct method, signed before tender closed, just below a round amount, repeat winner at the same buyer. Data problems are listed separately and are not flags.

ParametersJSON Schema
NameRequiredDescriptionDefault
ocidNo
buyerNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It helpfully discloses that data problems are reported separately and not treated as flags — a genuine output-semantics detail — but says nothing about permissions, cost, rate limits, or whether checks are exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loading the core purpose before the flag list. The middle line is telegraphic to the point of being a fragment, but nothing is padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the flag list plus the data-problems caveat cover the essentials. Missing are the parameter-interaction rules (both optional/none supplied) and any routing versus the sibling analysis tools, leaving gaps for a zero-coverage schema with no annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two parameters, so the description must compensate. It does clarify the scope of each parameter (ocid = one tender, buyer = all of that buyer's awards), which adds meaning, but it never states whether both can be combined or what happens when neither or both are supplied (both are optional with null defaults).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete capability — integrity/red-flag checks — and enumerates the specific patterns detected (single bidder, direct method, signed before close, just-below-round amounts, repeat winners), which is well beyond a tautology. It distinguishes itself implicitly from siblings like price_benchmark or supplier_profile, but never explicitly routes against them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for one tender (ocid) or all awards of a buyer' implies the two operating modes, which is some usage context. However, there is no guidance on when to prefer this over price_benchmark or supplier_profile, no exclusions, and no statement about what triggers the buyer-wide mode.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

draft_reportA

WRITE. Build the committee report (report.md) from the approved flags in a case.

Needs a named human approver. Each finding in the report cites its OCDS record.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYes
case_idYes
summaryYes
approved_byNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden and does reasonably well: the leading 'WRITE.' signals a side-effecting mutation, it states the concrete output artifact (report.md), it discloses the human-approval gate (an auth-like requirement), and it notes each finding cites its OCDS record. It omits overwrite/error behavior and permissions details, keeping it from a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the operation type ('WRITE.'), then purpose, then prerequisites in three short lines with no filler. The fragmentary line breaks cost it a little readability, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is unnecessary, and the write/approval prerequisites are covered. However, for a mutation tool with zero annotations and 0% parameter documentation, the description should say more about what happens on write (overwrite, failures) and what the remaining parameters mean.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it only partially does: 'approved flags in a case' loosely implies case_id and 'needs a named human approver' maps to approved_by. The title and summary parameters are undocumented in both schema and description, and approved_by is optional in the schema despite being framed as a requirement.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: 'Build the committee report (report.md) from the approved flags in a case' names both the artifact and its input source. This is clearly distinct from siblings like file_flag or check_red_flags, though it never explicitly contrasts itself with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States two concrete preconditions for use: flags must already be approved and a named human approver is required. It gives no explicit when-not guidance or alternative (e.g. what to do if flags are not yet approved), so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_flagA

WRITE, needs human approval. Add one flag to the case file.

finding: the pattern in one plain sentence. evidence: the numbers behind it. Refused if the ocid is unknown or the wording decides the award or declares wrongdoing.

ParametersJSON Schema
NameRequiredDescriptionDefault
ocidYes
case_idYes
findingYes
evidenceYes
flag_typeYes
approved_byNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full load and does well: it flags the write nature, the human-approval requirement, and the exact content conditions that cause rejection. It does not cover idempotency, duplicate-flag behavior, or what happens to case state after writing, which are the remaining behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the two most important facts (WRITE, needs approval) in the first clause, then adds the field guidance and refusal rules without padding. The fragmentary phrasing around finding/evidence is terse but readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be described, and the description covers the mutation/approval profile and rejection rules. The unresolved gap is the four undocumented parameters, including flag_type, whose semantics and allowed values are unavailable anywhere in the definition.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across six parameters, so the description must supply meaning for all of them. It explains only two ('finding' = the pattern in one sentence, 'evidence' = the numbers behind it), leaving ocid, case_id, flag_type, and approved_by entirely undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Add one flag to the case file,' which is concrete and tells the agent this creates an artifact rather than querying one. It does not explicitly distinguish itself from the sibling check_red_flags, but the write framing is clear enough to separate the two in most cases.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a hard precondition ('needs human approval') and explicit refusal conditions (unknown ocid, wording that decides the award or declares wrongdoing), which functions as meaningful when-not guidance. It never names an alternative tool such as check_red_flags, so routing among siblings is left partly to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

price_benchmarkB

Compare one tender's award value with similar awards (same item group and category).

Returns median, quartiles, ratio to median, outlier yes/no. Says so when under 5 comparables.

ParametersJSON Schema
NameRequiredDescriptionDefault
ocidYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden. It usefully discloses a data-sufficiency behavior ('says so when under 5 comparables'), but omits anything about read-only safety, permissions, or comparability constraints. Its enumeration of return fields is largely redundant given an output schema exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler, and the core action is front-loaded ahead of the return-value and caveat details. Every clause adds information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, annotated-free tool with an output schema, the description covers purpose, output shape, and a meaningful edge-case behavior. The remaining gap is the unexplained 'ocid' parameter, which neither the schema nor the description clarifies.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single 'ocid' parameter is never explained in the description. The phrase 'one tender's award value' only loosely implies the identifier's role; format, requiredness, and lookup semantics are left entirely to inference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('compare one tender's award value with similar awards') and pins the comparison scope to 'same item group and category'. This is enough to separate it from siblings like search_awards or supplier_profile, though it never explicitly names which sibling to prefer over check_red_flags, whose outlier-detection role overlaps.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the verb ('compare one tender's award value') but there is no explicit when-to-use, when-not-to-use, or named alternative among the sibling tools. The only conditional guidance is a data caveat, not a routing rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_awardsB

Find published awards. All filters optional.

buyer: part of the buyer name. item: word from the title or item group. supplier: company name. category: goods|services|works. method: open|selective|direct. date_from/date_to: YYYY-MM-DD signing date. If nothing matches, explains why and suggests close buyer names.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemNo
buyerNo
limitNo
methodNo
date_toNo
categoryNo
supplierNo
date_fromNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose one genuinely useful trait: on no matches it 'explains why and suggests close buyer names,' which tells the agent how to interpret empty results. It says nothing about pagination behavior, the meaning of limit, or result-set size, which is a real gap for a search tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core statement and the optionality note, then gives a compact per-parameter glossary with no filler. Slightly clipped phrasing in the last sentence ('explains why') costs it a point but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value documentation is not required, and the description covers nearly every input including enum vocabularies. Missing annotation coverage and the undocumented 'limit' parameter are the only substantive holes for an 8-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and largely does: it defines buyer ('part of the buyer name'), item ('word from the title or item group'), supplier, date format (YYYY-MM-DD signing date), and supplies enum values for category (goods|services|works) and method (open|selective|direct) that the schema does not contain. Only 'limit' is left undefined.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource with scope: 'Find published awards.' An agent immediately knows this is a search over award records. However, it never distinguishes itself from siblings like price_benchmark or check_red_flags, so sibling routing must be inferred.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'All filters optional' is the only usage hint, and it describes cardinality rather than when to reach for this tool over its siblings. There is no statement of when-not to use it or which alternative handles related questions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

supplier_profileA

Profile a supplier company: bids listed, awards won, win rate, buyers, single-bidder wins.

Uses fuzzy matching on the company name. Company records only, never personal details.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose two non-obvious traits: input is resolved by fuzzy matching (so the name need not be exact), and output is restricted to company records with no personal details. It stops short of confirming read-only status, coverage limits, or rate/cost behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero filler. The output contents are front-loaded and the two caveats (fuzzy matching, no personal data) follow immediately. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be explained, and input semantics are covered. The remaining gap is temporal scope — win rate, buyers, and single-bidder wins are period-dependent, and the description never states the time window covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% for the single 'name' parameter, so the description must compensate — and it does, specifying that matching against the supplied name is fuzzy rather than exact. That is the key semantic an agent needs, though no example format or minimum-length guidance is given.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Profile) and resource (supplier company), then enumerates exactly what the profile contains: bids listed, awards won, win rate, buyers, single-bidder wins. This distinguishes it implicitly from siblings like search_awards (raw award lookup) and check_red_flags (anomaly detection), though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied — an agent can infer you call this to get an aggregate picture of one company, versus search_awards for individual records. There is no explicit when-to-use, when-not-to-use, or named alternative among the five siblings, and no stated prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv0.1.0
    • First observedcheck_red_flags
    • First observeddraft_report
    • First observedfile_flag
    • First observedprice_benchmark
    • First observedsearch_awards
    • First observedsupplier_profile

TDQS

A3.6/5.0

Scored across 6 tools

Disambiguation4/5

Each tool targets a distinct action: search, benchmark, profile, flag detection, flag filing, and reporting. There is mild overlap between price_benchmark and check_red_flags (both analyze award patterns) and between supplier_profile and check_red_flags, but the descriptions make the boundaries reasonably clear.

Naming Consistency3/5

Four tools follow a verb-first pattern (search_awards, check_red_flags, file_flag, draft_report), while price_benchmark and supplier_profile use a noun-first pattern. Readable overall, but the convention is mixed rather than predictable.

Tool Count4/5

Six tools is well-scoped for an audit workflow spanning discovery, analysis, flagging, and reporting. It is on the leaner side but each tool clearly earns its place; no redundancy.

Completeness4/5

The surface covers the full audit lifecycle: find awards, benchmark prices, profile suppliers, detect red flags, file flags, and produce a report. A detail-retrieval tool (e.g. get_award/get_tender by ocid) is a minor gap, but agents can work around it via search.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Enables AI agents to find, score, and monitor government contract opportunities across UK, EU, and US with AI-powered relevance scoring.
    2
    49 npm
    1
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    MCP server that gives AI agents structured access to Public Procurement Registry data, allowing natural language queries for planned procurements, open tenders, contract awards, buyers, and winning suppliers.
    -
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables searching and analyzing government tenders, contract awards, and pre-tender pipelines from 21 official sources, with tools for tender search, award intelligence, and detailed notice retrieval.
    MIT