Skip to main content
Glama
zeybekci-lab

Payables MCP

by zeybekci-lab

Payables MCP

A small MCP server for accounts payable. I built it to work out what an AP agent should look like if you give a model tools instead of a UI: what it should be allowed to do on its own, what needs a person, and where the line between the two sits.

Everything runs against a mock ERP in data.json, so you can clone it and click through it in a couple of minutes. There are no API keys.

Running it

uv venv && uv sync
uv run mcp dev server.py      # opens the MCP inspector
uv run pytest                 # 5 tests

AP_PROFILE=analyst or AP_PROFILE=processor in front of the run command gives you a smaller server. More on that below.

Related MCP server: tessera-mcp

What you get

Tools

tools

Eleven tools. The first five read and write bills and vendors, the same shape as any ERP connector. The interesting ones are in the middle: check_duplicates, match_invoice and recommend_approval are where the AP judgment lives. They call plain functions in ap_logic.py and hand back a dict with the reasons, so a person can see why a bill got held.

request_payment_release is the one to notice. It queues approved bills for a payment run and that is all it does. There is no tool that releases the run or sends money. That part stays in the ERP, done by a person, on purpose.

Prompts

prompts

One prompt for now, the weekly payment run. It tells the assistant which resource to read and which tools to run in what order. It carries no data itself.

Resources

resources

Things the assistant reads often and should not need a tool call for: all bills, one bill by id, bills due in the next N days, and the tolerance policy.

Profiles

Not every deployment should get every tool. AP_PROFILE decides what gets registered when the server starts:

profile

can

analyst

read bills and vendors, run the checks, see the audit log

processor

all of the above, plus create bills and record bank changes

full

all of the above, plus queue payment runs

An analyst server does not refuse create_bill. It never lists it, so the model cannot try.

The files

file

server.py

the tools, resources, prompt, and the profile setup at the bottom

ap_logic.py

matching, duplicate check, screening. No MCP in here, so it tests on its own

erp.py

the mock ERP. Swap this for a real backend and the server stays the same

data.json

3 vendors, 3 purchase orders, 5 bills

test_server.py

the logic on its own, one bill through the whole flow, and a check on tool names

The five bills are set up so each one shows something: B-1001 is clean, B-1002 is a duplicate of it that came in through a different channel, B-1003 bills more than was received, B-1004 is from a vendor that changed bank details by email last week and has a memo asking for a payment change, and B-1005 has no purchase order.

Available Tools

11 tools
check_duplicatesB

Check a bill against the vendor's earlier bills. A match sets the bill to exception.

ParametersJSON Schema
NameRequiredDescriptionDefault
bill_idYesBill id

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses a state-changing side effect ('a match sets the bill to exception'), which is genuine behavioral value, but it omits permissions needed, reversibility, and what happens/returns when there is no match.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler, and the core action is front-loaded before the side effect. It could be slightly richer, but nothing here is wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with no output schema and no annotations, the description covers the action and one outcome (match → exception) but is silent on the no-match case and the return value. Adequate but with a visible gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter (bill_id) with 100% schema description coverage, so the schema already documents the input fully. The description adds no format, source, or semantics beyond what the schema provides; baseline 3 is appropriate for a single fully-documented parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (check) and resource (a bill against the vendor's earlier bills), which is well beyond a tautology. It is distinguishable from read-only siblings like get_bill or match_invoice, though it never explicitly names which sibling handles the non-duplicate path.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the phrasing (run this to detect duplicate bills from the same vendor), but there is no explicit when-to-use, when-not-to-use, or alternative tool named. An agent must infer the trigger condition from the description alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_billA

Create a bill. Bills over 25,000 ask the user to confirm first.

ParametersJSON Schema
NameRequiredDescriptionDefault
qtyYesQuantity
po_idNoPurchase order id, if any
due_dateYesYYYY-MM-DD
vendor_idYesVendor id
unit_priceYesUnit price
invoice_dateYesYYYY-MM-DD
invoice_numberYesInvoice number from the vendor

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It helpfully discloses the >25,000 confirmation threshold, a genuinely useful interaction rule, but says nothing about side effects (record creation), permissions, or whether repeated invoice numbers are rejected.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, both earning their place: the first establishes the action, the second front-loads the one non-obvious behavioral constraint. No filler whatsoever.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter mutation with no annotations and no output schema, the description is thin. It omits what the tool returns (bill id?), whether the bill is created in draft/pending state, and how errors like duplicate invoice numbers are handled.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 7 parameters are already documented in the schema with dates, ids, and quantity/price labels. The description adds no additional parameter meaning, which is the baseline expectation when the schema does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a bill'), which is unambiguous, but it offers no differentiation from siblings like match_invoice or request_payment_release and does not clarify scope (single bill, required fields).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The tool's purpose implies when it is used, and the description adds a conditional workflow rule (confirmation above 25,000). However, it names no alternatives and gives no guidance on prerequisites such as whether a vendor or PO must already exist.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_billB

Fetch one bill by id, with its total and status.

ParametersJSON Schema
NameRequiredDescriptionDefault
bill_idYesBill id, e.g. B-1001

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. 'Fetch' implies a read-only lookup and it does disclose the returned fields (total, status), which is useful, but it says nothing about lookup failure behavior for an unknown id, permissions, or whether the result is cached/live.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the operation front-loaded and zero filler; every clause (verb, lookup key, return fields) earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-required-parameter lookup with no output schema, the description covers the operation, the key, and the returned fields, which is enough to invoke it correctly. It loses a point only for not clarifying the single-id versus search distinction among siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema already supplies an example ('B-1001'), so the parameter is fully documented in structured data. The phrase 'by id' adds nothing beyond that, making the baseline 3 correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Fetch') and resource ('one bill by id') and adds the returned shape ('total and status'), so the operation is unambiguous. However, it never distinguishes itself from the sibling search_bills, which is the obvious adjacent retrieval tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use context is given. With siblings like search_bills and get_pending_releases, the description should say this is for a single known bill id versus list/search retrieval, but it offers no routing guidance or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_pending_releasesB

List payment runs waiting to be released.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not disclose whether the operation is read-only, what permissions are required, pagination behavior, or any rate limits; 'List' implies a safe read but this is not made explicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler or repetition. Every word contributes to stating the tool's purpose efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool takes no parameters and has an output schema, the description adequately states what is returned. However, with no annotations, it could have added a note about read-only safety or scope (e.g., all runs vs. filtered), leaving a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there are no parameter semantics to explain. Per scoring guidance, a zero-parameter tool receives a baseline of 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('payment runs waiting to be released'), making the tool's purpose immediately clear. It does not explicitly distinguish itself from the sibling request_payment_release, but the wording is specific enough for correct identification.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives like request_payment_release or search_bills. Usage is only implied by the phrase 'waiting to be released,' leaving the agent to infer the context without any stated conditions or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_vendorC

Fetch one vendor by id.

ParametersJSON Schema
NameRequiredDescriptionDefault
vendor_idYesVendor id, e.g. V-01

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden and does almost nothing with it. It does not say what happens when the vendor id does not exist, whether soft-deleted vendors are returned, what permissions are required, or what shape the response takes — all of which matter for a lookup tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The single sentence is front-loaded and free of waste, but its terseness is under-specification rather than disciplined brevity. It is appropriately short, not notably well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a trivial one-parameter read with no output schema and a fully documented parameter, the description is minimally sufficient to invoke the tool. Given the absence of annotations, though, an agent gets no signal about failure modes or auth, leaving it only barely adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the sole parameter is documented with an example format ('V-01'), so the schema does the heavy lifting. The description adds no syntax, format, or validation detail beyond what the schema already provides, which is the baseline case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Fetch) and resource (vendor) and constrains it to a single record by id. No sibling tool fetches a vendor, so sibling collision is not a real risk here, but the description never references any sibling to anchor its place in the workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of prerequisites, and no pointer to any alternative such as a search or list tool. The agent must infer that this is the single-record lookup path rather than a query path.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

match_invoiceA

Match a bill against its PO and goods receipt. Returns matched, exception or no_po.

ParametersJSON Schema
NameRequiredDescriptionDefault
bill_idYesBill id

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the three possible outcomes (matched, exception, no_po), which is genuine behavioral context. However, it says nothing about side effects (does matching persist an exception record or trigger downstream workflow?), permission requirements, or whether the operation is read-only.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two terse sentences that front-load the action and end with the outcome set; no filler or restated field names. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with no output schema and no annotations, the description is adequate but thin: it never explains what drives each outcome, what an 'exception' constitutes, or whether a match mutates state. An agent could call it correctly but not predict its side effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for the single bill_id parameter, so the schema already carries the parameter contract and the baseline is 3. The description adds no format, source, or identifier-type detail beyond what the schema states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Match') and names both the input resource (a bill) and what it is matched against (PO and goods receipt). No sibling tool (get_bill, create_bill, check_duplicates, etc.) performs matching, so an agent can route to it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the action but gives no when-to-use guidance, prerequisites, or sequencing relative to siblings like check_duplicates or recommend_approval. The closest thing to an alternative, check_duplicates, is never mentioned, and no conditions or exclusions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

read_audit_logB

Show the most recent audit entries.

ParametersJSON Schema
NameRequiredDescriptionDefault
nNoHow many

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It implies a read operation via 'Show' but does not state that the tool is read-only, non-destructive, or whether it needs any permissions or has rate limits. Only the most basic behavioral hint is present.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It is appropriately sized for a simple read tool and every part of the sentence contributes to meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema covers return values and the schema fully documents the one parameter, so those areas need not be described. However, with no annotations, the description omits the safety profile and does not mention the default count or that it returns the most recent n entries. It is minimally adequate but leaves notable gaps for a tool with zero annotation support.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the single parameter 'n' and its default of 50. The description adds no parameter semantics beyond what the schema provides, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Show') and resource ('audit entries'), plus a scope ('most recent'). It clearly distinguishes this tool from the sibling tools by resource domain, though it does not explicitly name or contrast with any alternative. A 4 is appropriate because the purpose is immediately legible without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, no mention of prerequisites or context. It simply states a function, leaving the agent to infer usage entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

recommend_approvalA

Recommend approve or hold for a bill, with the reasons. Does not approve anything.

ParametersJSON Schema
NameRequiredDescriptionDefault
bill_idYesBill id

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral load, and it does disclose the critical non-mutating fact ('Does not approve anything') plus that reasons are returned. It stops short of saying whether anything is persisted, whether a recommendation is required before approval, or whether it needs specific permissions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, and the safety-relevant clause ('Does not approve anything') is placed immediately after the capability statement where it will be read.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter advisory tool with no output schema, the description covers what the tool does, the possible outcomes, and the returned reasons. Only the relationship to the downstream approval flow is left implicit, which is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single bill_id parameter is fully documented in the schema, so the baseline is 3. The description adds no format, sourcing, or semantics beyond 'for a bill'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (recommend), the decision space (approve or hold), the resource (a bill), and the return content (with the reasons). That is enough to distinguish it from siblings like get_bill or match_invoice, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: 'Does not approve anything' hints that a separate tool (e.g. request_payment_release) performs the actual approval, but the description never states when to call this versus those alternatives or what precondition a bill must meet.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

request_payment_releaseA

Queue approved bills for a payment run. Someone releases the run in the ERP.

ParametersJSON Schema
NameRequiredDescriptionDefault
bill_idsYesBill ids, all approved
run_dateYesYYYY-MM-DD

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It usefully discloses that this queues rather than pays, but says nothing about permissions, idempotency, whether queuing is reversible, or what happens to already-queued bills.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences, front-loaded with the action and followed by the scoping caveat; every sentence does work. It is a touch terse given the unanswered behavioral questions, but there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter mutation with no output schema, the description covers intent and scope adequately. It leaves the return value, failure behavior for unapproved bills, and auth requirements unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (bill_ids, run_date) are already documented in the schema. The description's only added signal, that bills must be approved, duplicates the schema's own 'all approved' note and adds no format or syntax detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Queue approved bills for a payment run'), which an agent can distinguish from read-side siblings such as get_pending_releases or get_bill. It does not explicitly contrast itself with those siblings, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The prerequisite ('approved bills') and the scope boundary ('Someone releases the run in the ERP') give clear context for when this tool applies and where its responsibility ends. No alternatives or when-not conditions are named, so no exclusions are documented.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_billsB

List bills, optionally filtered by status and vendor.

ParametersJSON Schema
NameRequiredDescriptionDefault
statusNoreceived, matched, exception, approved or release_requested
vendor_idNoVendor id, e.g. V-01

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full behavioral burden. "List" weakly implies a read-only operation, but nothing is said about ordering, pagination, result limits, or scope — significant gaps for a search tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the resource and filters front-loaded and no wasted words. It is minimal but functional rather than padded or rambling.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and complexity is low with two optional parameters. Still, with no annotations and no guidance on result ordering or how it relates to get_bill, the definition is only adequate rather than complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented with their valid values, establishing a baseline of 3. The description only restates the filter fields (status, vendor) without adding format or matching semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("List bills") plus the two filter dimensions, which distinguishes it from the singular sibling get_bill. It does not, however, explicitly name get_bill as the alternative for single-record lookups.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Optionally filtered by status and vendor" implies the filtering use case but gives no explicit when-to-use vs get_bill or when-not-to-use guidance. Usage is inferable from the name and description rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_vendor_bankB

Record a change to a vendor's bank details. The vendor is frozen until someone confirms it in the ERP.

ParametersJSON Schema
NameRequiredDescriptionDefault
vendor_idYesVendor id
new_accountYesNew account number

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose a meaningful consequence: the vendor is frozen until ERP confirmation. It omits other key traits such as required permissions, whether this is a direct write or a queued request, and reversibility.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, the action front-loaded and the consequence immediately following. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations and no output schema, the description covers the action and one consequence but leaves gaps around authorization, the approval/freeze workflow hinted at by siblings like request_payment_release, and what the caller receives back.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both parameters, so the schema already documents vendor_id and new_account. The description adds no format, validation, or account-structure detail beyond the schema, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Record a change) and resource (vendor's bank details), which is far better than a tautology. It does not, however, explicitly distinguish itself from sibling tools such as get_vendor or request_payment_release.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No statement of when to use this versus alternatives, nor any prerequisite such as needing duplicate checks (check_duplicates) or approval routing (recommend_approval). Usage is only implied by the verb.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 11 tool updatesv0.1.0
    • First observedcheck_duplicates
    • First observedcreate_bill
    • First observedget_bill
    • First observedget_pending_releases
    • First observedget_vendor
    • First observedmatch_invoice
    • First observedread_audit_log
    • First observedrecommend_approval
    • First observedrequest_payment_release
    • First observedsearch_bills
    • First observedupdate_vendor_bank

TDQS

A3.5/5.0

Scored across 11 tools

Disambiguation4/5

Most tools have clearly distinct roles (create/get/search bills, match against PO, check duplicates, recommend approval, release payments). The only potential confusion is between check_duplicates and match_invoice, both bill-validation steps, but their descriptions ('earlier bills' vs 'PO and goods receipt') clearly separate them.

Naming Consistency5/5

All tools follow a consistent snake_case verb_noun pattern (get_bill, search_bills, create_bill, match_invoice, request_payment_release). No mixing of conventions or vague verbs; the pattern is predictable throughout.

Tool Count5/5

11 tools is well-scoped for a payables/AP workflow. Each tool maps to a distinct step (bill CRUD, validation, approval recommendation, payment release, audit) and earns its place without redundancy.

Completeness4/5

Core bill lifecycle (create, get, search, validate, recommend, release) and audit are covered, but there is no update_bill/void, no vendor search or create (only get_vendor by id), and no explicit approve action. These are workable gaps for a controlled approval workflow but leave some dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    An accounting-ops agent that reconciles payments against open orders, auto-books provably safe payments through a deterministic policy gate, and escalates exceptions to a human queue with audit trails.
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables AI agents to manage a fictional B2B workspace SaaS (Tessera) with tools for ticketing, invoicing, customer management, and trial extensions, featuring a human-in-the-loop confirm pattern for safety.
    7
    13 npm
    MIT
  • F
    license
    A
    quality
    B
    maintenance
    Exposes a ledger system (invoice queue, duplicate control, VAT register, contractor history, decision journal) as MCP tools for AI agents, enabling accurate invoice processing with deterministic validation.
    7
    -