ke-tenders-audit-mcp
This MCP server lets an agent search Kenyan public procurement awards, run integrity checks, and support a human-gated review workflow without making award decisions.
Search published awards by buyer, item, supplier, category, method, or signing date.
Run red-flag checks for single bidder, direct method, signed before tender close, just below a round amount, and repeat winners at the same buyer.
Benchmark one award's value against similar awards, with median, quartiles, ratio, and outlier status.
Profile a supplier's bids, awards won, win rate, buyers, and single-bidder wins.
File sourced flags into a case file (write; requires named human approval).
Draft a committee report from approved flags with OCDS citations (write; requires named human approval).
Read cleaned OCDS release records and data-quality notes via MCP resources.
Enforce grounding, refuse unsupported or guilt-deciding wording, and prevent the model from approving itself.
Read-only access to the local cases folder via the official Filesystem MCP server, letting the agent open and read earlier case files while preparing a procurement review.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ke-tenders-audit-mcpReview all published awards by PC Kinyanjui Technical Training Institute"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Ke-Tenders Audit
An agentic AI that prepares a sourced review file on Kenyan public procurement awards for a human evaluation committee, then stops. It flags patterns. It never awards a tender, never recommends a winner and never declares wrongdoing. Every write needs a named human's approval.
Built for the African Agentic AI Design Challenge, Governance track (The Bid Box Challenge), theme Value for money.
Live demo: https://ke-tenders-audit-bxxfc4gmxjc9mpuuvkzypz.streamlit.app/ (opens on recorded real runs; switch to "Live review" in the sidebar for up to 3 live reviews a day on a free model tier)
pip install -r requirements.txt && streamlit run app.py(Python 3.12. Put a model key in .env first: see Setup.)
Problem statement
Kenyan public bodies (county governments, TVET colleges, schools, NG-CDF offices, water companies) publish every tender and award on the PPRA portal, tenders.go.ke, in the Open Contracting Data Standard (OCDS). Almost nobody has time to use it. Evaluation committees and internal auditors spend weeks checking bidders, comparing prices and writing reports; the challenge brief puts it at 26 days of committee work per tender. Oversight bodies face the opposite problem: thousands of records and no way to see which ones deserve a second look.
The warning signs are already in the data. In the FY2026/27 release (4,830 tenders, 136 with awards), 23 awards had a single bidder, some suppliers win repeatedly from the same buyer, some contracts are signed before the tender closed, and some awards sit just below round amounts.
Related MCP server: Tender MCP
Which half of the data this is built on
The brief notes that portals publish what was tendered and awarded, not the bids. This project is built on the tendered-and-awarded half: award values, bidder lists on awarded tenders, dates, methods and buyers. It does not evaluate bid documents, because none are published.
Solution overview
A reviewer asks a question such as "Review all published awards by PC Kinyanjui Technical Training Institute". The agent plans the review, calls tools on our MCP server to search awards, run integrity checks, compare prices and profile suppliers, then proposes flags. Each proposed flag pauses the run until a named reviewer approves, rejects or requests changes. Approved flags go into a case file, and the agent drafts a committee report where every finding cites its OCDS record.
What the code enforces, not just the prompt:
A flag must cite an ocid that a tool actually returned in this run, or it is refused before a human sees it.
A flag naming a company that is not a bidder or winner on that record is refused, with the correct name.
Only the human gate can fill
approved_by; the model cannot approve itself.A write the reviewer rejected is refused if the model proposes it again.
Wording that decides an award or declares guilt is refused.
Every report ends with a table of all patterns found, computed from the data, so a wrong model summary cannot hide anything.
No em dashes or en dashes in anything written to the case file.
Target users
Procurement evaluation committees and heads of procurement in counties, colleges and state agencies.
Internal audit units preparing for the Office of the Auditor-General.
Also useful to oversight bodies, journalists and civil society, and to honest suppliers who gain from fairer competition.
Architecture

One page: ARCHITECTURE.md.
Setup
Requirements: Python 3.12, Node.js 18+ (for the borrowed Filesystem MCP server, run with npx).
Copy
.env.exampleto.envand choose a model:Hosted Qwen on Groq (free key at console.groq.com): set
KTA_LLM_BASE_URL=https://api.groq.com/openai/v1,KTA_LLM_MODEL=qwen/qwen3.8-27band your key inKTA_LLM_API_KEY.Fully local with Ollama (data never leaves the machine): install Ollama, run
ollama pull qwen3:4b, and keep the defaulthttp://localhost:11434/v1settings.
Install and start:
pip install -r requirements.txt && streamlit run app.pyThe cleaned data ships in data/clean/releases.jsonl, so no download is needed. To rebuild it from a
fresh PPRA download, save the JSON from https://tenders.go.ke/api/ocds/tenders?fy=2026-2027 into
data/raw/ and run ke-tenders-ingest.
Usage
Review screen: enter your full name, ask a question, watch each tool call with its input and output, then approve, reject or request changes on each proposed flag. Each proposal shows the record's own facts (title, winner, amount, bidders, method) beside the model's claim, with a red warning on any mismatch. Download the report when done.
Terminal:
python -m ke_tenders_audit.cli "Review Kilifi County Government's awards." --reviewer "Your Full Name"Tests and evals:
pytestpython evals/run_evals.py --runs 2Technology stack
Layer | Choice |
Orchestration | LangGraph (state graph, |
Model | Qwen3.8-27B (open weights) on Groq; Qwen3 4B locally through Ollama; any OpenAI-compatible endpoint |
MCP | Official MCP Python SDK (our server), |
Data | DuckDB in memory, built from cleaned OCDS JSONL; RapidFuzz for name matching |
Interface | Streamlit review screen, plus a terminal runner |
Logs | Append-only JSONL audit log per run |
Tests | pytest (31 tests), plus a graded eval harness |
Agent architecture
plan -> agent -> read tools -> tools -> verify -> agent ...
-> write tools -> grounding check -> human gate -> tools -> verify -> agent ...
-> no tool calls -> doneplan: the model writes a short numbered plan for the request.
agent: the model chooses tools, at most 3 per step.
verify: after each tool round, deterministic checks look for errors, refusals, hints and "insufficient comparables", and tell the model plainly so it re-plans instead of guessing.
grounding check and human gate: described above.
memory: LangGraph checkpoints in SQLite, so a run pauses for approval and resumes exactly where it stopped, even after an error ("Retry last step").
budget: older tool results are shortened to keep each request under about 7,000 tokens, which fits free model tiers. Red-flag results put their totals first so shortening never loses them.
MCP implementation
Two MCP servers, both connected over stdio through langchain-mcp-adapters. A tool interceptor
writes every call (server, tool, inputs, output, duration) to the audit log.
MCP tools and servers
Built: ke-tenders-audit-mcp (server.py)
Tool | Kind | What it does |
| read | Find awards by buyer, item, supplier, category, method or date. Explains empty results and suggests close buyer names. |
| read | Single bidder, direct method, signed before close, just below a round amount, repeat winner at the same buyer. Totals first. |
| read | Compare an award with similar awards (same item group and category). Says so when there are fewer than 5. |
| read | Bids, wins, win rate, buyers and single-bidder wins for a company. |
| write, gated | Add a sourced flag to the case file, with the record's own facts attached. |
| write, gated | Build the committee report from approved flags, plus the data-computed totals table. |
Resources: ocds://release/{ocid} (the cleaned record behind any finding) and
ocds://data-quality-notes (known problems in the feed).
Borrowed: the official Filesystem MCP server (@modelcontextprotocol/server-filesystem), scoped to
the cases/ folder, read tools only. It lets the agent read earlier case files. It is maintained,
tested and sandboxed to one folder, so writing our own file access would add risk and no value.
Human-in-the-loop workflow
The agent proposes one or more writes. The run pauses (LangGraph
interrupt()).The reviewer sees each proposal with the record's facts and the full source record.
For each one: Approve, Request changes (with a note the agent uses to revise) or Reject (never proposed again in this run).
The gate writes the reviewer's name into
approved_by. The audit log records the decision.The final report says who approved it. The committee makes the award decision, outside this system.
Data conduct
Open data only, from PPRA. Cleaning removes contactPoint, street addresses and postal codes, and
strips emails and phone numbers that publishers typed into company names. Company records are used;
no individual's personal details are. No live tender in progress is evaluated: awards are by
definition closed.
Cost per run
Measured on Groq, Qwen3.8-27B at $0.80 and $4.00 per million input and output tokens: a full buyer review (31 awards, 7 model calls, 13 tool calls) used about 24,000 input tokens, about $0.03. A single-tender review costs under $0.02. Locally with Ollama the cost is zero. Every run's tokens and cost are in its audit log. Evaluation results: EVALS.md.
Limitations
No bid documents. The feed has no bids, so technical and financial evaluation of bids is out of scope.
Lump sums, not unit prices. Price comparisons point to scope questions, not proven overpricing.
Small award sample. Only 136 tenders have published awards in FY2026/27, so many items have too few comparables.
Item grouping is rule-based. Free-text titles are grouped with keyword rules; about 7% of awards stay ungrouped.
Thresholds are round numbers, not confirmed law. "Just below a round amount" needs checking against the PPADA Regulations 2020.
Free model tier. On Groq's free tier a review takes 1 to 3 minutes, mostly rate-limit waiting, and about 6 full reviews fit in a day.
The model's summary can still be wrong. The data-computed table at the end of each report is the safeguard.
Future improvements
Load past fiscal years (FY2025/26 and earlier) for stronger price comparisons and supplier histories.
Run against other OCDS portals (Rwanda, Tanzania, Nigeria, South Africa); the tools are schema-generic.
Confirm PPADA thresholds and replace round amounts with the real limits per procurement method.
Match suppliers on registration numbers if PPRA publishes them.
Unit-price extraction from contract documents, where those become public.
Licence
MIT. See LICENSE. Data: Public Procurement Regulatory Authority (Kenya), published as open data.
Available Tools
6 toolscheck_red_flagsB
Integrity checks for one tender (ocid) or all awards of a buyer.
Single bidder, direct method, signed before tender closed, just below a round amount, repeat winner at the same buyer. Data problems are listed separately and are not flags.
| Name | Required | Description | Default |
|---|---|---|---|
| ocid | No | ||
| buyer | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It helpfully discloses that data problems are reported separately and not treated as flags — a genuine output-semantics detail — but says nothing about permissions, cost, rate limits, or whether checks are exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loading the core purpose before the flag list. The middle line is telegraphic to the point of being a fragment, but nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the flag list plus the data-problems caveat cover the essentials. Missing are the parameter-interaction rules (both optional/none supplied) and any routing versus the sibling analysis tools, leaving gaps for a zero-coverage schema with no annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for two parameters, so the description must compensate. It does clarify the scope of each parameter (ocid = one tender, buyer = all of that buyer's awards), which adds meaning, but it never states whether both can be combined or what happens when neither or both are supplied (both are optional with null defaults).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete capability — integrity/red-flag checks — and enumerates the specific patterns detected (single bidder, direct method, signed before close, just-below-round amounts, repeat winners), which is well beyond a tautology. It distinguishes itself implicitly from siblings like price_benchmark or supplier_profile, but never explicitly routes against them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for one tender (ocid) or all awards of a buyer' implies the two operating modes, which is some usage context. However, there is no guidance on when to prefer this over price_benchmark or supplier_profile, no exclusions, and no statement about what triggers the buyer-wide mode.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
draft_reportA
WRITE. Build the committee report (report.md) from the approved flags in a case.
Needs a named human approver. Each finding in the report cites its OCDS record.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | ||
| case_id | Yes | ||
| summary | Yes | ||
| approved_by | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden and does reasonably well: the leading 'WRITE.' signals a side-effecting mutation, it states the concrete output artifact (report.md), it discloses the human-approval gate (an auth-like requirement), and it notes each finding cites its OCDS record. It omits overwrite/error behavior and permissions details, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the operation type ('WRITE.'), then purpose, then prerequisites in three short lines with no filler. The fragmentary line breaks cost it a little readability, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is unnecessary, and the write/approval prerequisites are covered. However, for a mutation tool with zero annotations and 0% parameter documentation, the description should say more about what happens on write (overwrite, failures) and what the remaining parameters mean.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it only partially does: 'approved flags in a case' loosely implies case_id and 'needs a named human approver' maps to approved_by. The title and summary parameters are undocumented in both schema and description, and approved_by is optional in the schema despite being framed as a requirement.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: 'Build the committee report (report.md) from the approved flags in a case' names both the artifact and its input source. This is clearly distinct from siblings like file_flag or check_red_flags, though it never explicitly contrasts itself with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States two concrete preconditions for use: flags must already be approved and a named human approver is required. It gives no explicit when-not guidance or alternative (e.g. what to do if flags are not yet approved), so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
file_flagA
WRITE, needs human approval. Add one flag to the case file.
finding: the pattern in one plain sentence. evidence: the numbers behind it. Refused if the ocid is unknown or the wording decides the award or declares wrongdoing.
| Name | Required | Description | Default |
|---|---|---|---|
| ocid | Yes | ||
| case_id | Yes | ||
| finding | Yes | ||
| evidence | Yes | ||
| flag_type | Yes | ||
| approved_by | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full load and does well: it flags the write nature, the human-approval requirement, and the exact content conditions that cause rejection. It does not cover idempotency, duplicate-flag behavior, or what happens to case state after writing, which are the remaining behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the two most important facts (WRITE, needs approval) in the first clause, then adds the field guidance and refusal rules without padding. The fragmentary phrasing around finding/evidence is terse but readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be described, and the description covers the mutation/approval profile and rejection rules. The unresolved gap is the four undocumented parameters, including flag_type, whose semantics and allowed values are unavailable anywhere in the definition.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across six parameters, so the description must supply meaning for all of them. It explains only two ('finding' = the pattern in one sentence, 'evidence' = the numbers behind it), leaving ocid, case_id, flag_type, and approved_by entirely undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Add one flag to the case file,' which is concrete and tells the agent this creates an artifact rather than querying one. It does not explicitly distinguish itself from the sibling check_red_flags, but the write framing is clear enough to separate the two in most cases.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a hard precondition ('needs human approval') and explicit refusal conditions (unknown ocid, wording that decides the award or declares wrongdoing), which functions as meaningful when-not guidance. It never names an alternative tool such as check_red_flags, so routing among siblings is left partly to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
price_benchmarkB
Compare one tender's award value with similar awards (same item group and category).
Returns median, quartiles, ratio to median, outlier yes/no. Says so when under 5 comparables.
| Name | Required | Description | Default |
|---|---|---|---|
| ocid | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It usefully discloses a data-sufficiency behavior ('says so when under 5 comparables'), but omits anything about read-only safety, permissions, or comparability constraints. Its enumeration of return fields is largely redundant given an output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler, and the core action is front-loaded ahead of the return-value and caveat details. Every clause adds information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, annotated-free tool with an output schema, the description covers purpose, output shape, and a meaningful edge-case behavior. The remaining gap is the unexplained 'ocid' parameter, which neither the schema nor the description clarifies.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single 'ocid' parameter is never explained in the description. The phrase 'one tender's award value' only loosely implies the identifier's role; format, requiredness, and lookup semantics are left entirely to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('compare one tender's award value with similar awards') and pins the comparison scope to 'same item group and category'. This is enough to separate it from siblings like search_awards or supplier_profile, though it never explicitly names which sibling to prefer over check_red_flags, whose outlier-detection role overlaps.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the verb ('compare one tender's award value') but there is no explicit when-to-use, when-not-to-use, or named alternative among the sibling tools. The only conditional guidance is a data caveat, not a routing rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_awardsB
Find published awards. All filters optional.
buyer: part of the buyer name. item: word from the title or item group. supplier: company name. category: goods|services|works. method: open|selective|direct. date_from/date_to: YYYY-MM-DD signing date. If nothing matches, explains why and suggests close buyer names.
| Name | Required | Description | Default |
|---|---|---|---|
| item | No | ||
| buyer | No | ||
| limit | No | ||
| method | No | ||
| date_to | No | ||
| category | No | ||
| supplier | No | ||
| date_from | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose one genuinely useful trait: on no matches it 'explains why and suggests close buyer names,' which tells the agent how to interpret empty results. It says nothing about pagination behavior, the meaning of limit, or result-set size, which is a real gap for a search tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core statement and the optionality note, then gives a compact per-parameter glossary with no filler. Slightly clipped phrasing in the last sentence ('explains why') costs it a point but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value documentation is not required, and the description covers nearly every input including enum vocabularies. Missing annotation coverage and the undocumented 'limit' parameter are the only substantive holes for an 8-parameter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate and largely does: it defines buyer ('part of the buyer name'), item ('word from the title or item group'), supplier, date format (YYYY-MM-DD signing date), and supplies enum values for category (goods|services|works) and method (open|selective|direct) that the schema does not contain. Only 'limit' is left undefined.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with scope: 'Find published awards.' An agent immediately knows this is a search over award records. However, it never distinguishes itself from siblings like price_benchmark or check_red_flags, so sibling routing must be inferred.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'All filters optional' is the only usage hint, and it describes cardinality rather than when to reach for this tool over its siblings. There is no statement of when-not to use it or which alternative handles related questions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
supplier_profileA
Profile a supplier company: bids listed, awards won, win rate, buyers, single-bidder wins.
Uses fuzzy matching on the company name. Company records only, never personal details.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose two non-obvious traits: input is resolved by fuzzy matching (so the name need not be exact), and output is restricted to company records with no personal details. It stops short of confirming read-only status, coverage limits, or rate/cost behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler. The output contents are front-loaded and the two caveats (fuzzy matching, no personal data) follow immediately. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be explained, and input semantics are covered. The remaining gap is temporal scope — win rate, buyers, and single-bidder wins are period-dependent, and the description never states the time window covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single 'name' parameter, so the description must compensate — and it does, specifying that matching against the supplied name is fuzzy rather than exact. That is the key semantic an agent needs, though no example format or minimum-length guidance is given.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Profile) and resource (supplier company), then enumerates exactly what the profile contains: bids listed, awards won, win rate, buyers, single-bidder wins. This distinguishes it implicitly from siblings like search_awards (raw award lookup) and check_red_flags (anomaly detection), though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied — an agent can infer you call this to get an aggregate picture of one company, versus search_awards for individual records. There is no explicit when-to-use, when-not-to-use, or named alternative among the five siblings, and no stated prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
check_red_flags - First observed
draft_report - First observed
file_flag - First observed
price_benchmark - First observed
search_awards - First observed
supplier_profile
TDQS
Scored across 6 tools
Each tool targets a distinct action: search, benchmark, profile, flag detection, flag filing, and reporting. There is mild overlap between price_benchmark and check_red_flags (both analyze award patterns) and between supplier_profile and check_red_flags, but the descriptions make the boundaries reasonably clear.
Four tools follow a verb-first pattern (search_awards, check_red_flags, file_flag, draft_report), while price_benchmark and supplier_profile use a noun-first pattern. Readable overall, but the convention is mixed rather than predictable.
Six tools is well-scoped for an audit workflow spanning discovery, analysis, flagging, and reporting. It is on the leaner side but each tool clearly earns its place; no redundancy.
The surface covers the full audit lifecycle: find awards, benchmark prices, profile suppliers, detect red flags, file flags, and produce a report. A detail-retrieval tool (e.g. get_award/get_tender by ocid) is a minor gap, but agents can work around it via search.
Maintenance
Related MCP Connectors
License-clean African & EM data as citable agent tools: company fundamentals, tenders (OCDS), macro.
Search, verify & screen 1M+ African companies + their government contracts across 18 registries.
Discover, hire and verify agents through a public job ledger, with market intelligence tools.
Agent-native security, trust, reliability, data and procurement tools for AI workflows.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables querying international public procurement data, including government tenders, through the Open Contracting standard.399 npmMIT
- AlicenseAqualityBmaintenanceEnables AI agents to find, score, and monitor government contract opportunities across UK, EU, and US with AI-powered relevance scoring.249 npm1MIT
- FlicenseNot gradedqualityBmaintenanceMCP server that gives AI agents structured access to Public Procurement Registry data, allowing natural language queries for planned procurements, open tenders, contract awards, buyers, and winning suppliers.-
- AlicenseNot gradedqualityBmaintenanceEnables searching and analyzing government tenders, contract awards, and pre-tender pipelines from 21 official sources, with tools for tender search, award intelligence, and detailed notice retrieval.MIT