NitroWatch Billing Demo
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@NitroWatch Billing Demosend invoice inv_2215 to the customer"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
NitroWatch Billing Demo
An intentionally ungoverned MCP server. This is the thing that needs governing.
This is the companion demo target for NitroWatch — a governance layer for MCP.
This server has no safety of its own. Every tool below does exactly what it is told, immediately, for anyone who asks. delete_account will permanently destroy an account and all of its invoices without a single confirmation.
That is the point. It is indistinguishable from thousands of real MCP servers — and that is exactly why NitroWatch exists.
The tools
Five tools, chosen so one lands on each of NitroWatch's risk tiers:
Tool | Effect | NitroWatch classifies it as |
| Fetch one invoice |
|
| List invoices, optionally by account |
|
| Emails an invoice, marks it sent |
|
| Archives an account | unmatched → defaults to |
| Permanently deletes an account and purges its invoices |
|
Why archive_customer is the interesting one
Nothing in its name or description reveals that archived accounts here are purged after 30 days and cannot be restored. It reads as recoverable. No classifier — heuristic or LLM — could know otherwise.
That is precisely the case NitroWatch's set_tier exists for: a human who knows the domain corrects the classification, and the override is recorded in the audit trail.
Related MCP server: mcp-trust-demo
Running it
npm install
npm run build
npx nitrostack-cli start --port 3100⚠️ Use the
--portflag. ThePORTenvironment variable is silently ignored — the server binds 3000 regardless, which collides with NitroWatch. This is tested behaviour, not a guess.
When it's up you'll see:
Legacy SDK SSE http://localhost:3100/sse
Streamable HTTP http://localhost:3100/mcpPutting it under governance
From NitroWatch:
register_server({ name: "Billing API", endpoint: "http://localhost:3100/sse" })
discover_capabilities({ serverId: "billing-api" })
classify_tools({ serverId: "billing-api" })Then try to delete something, and watch it get stopped:
request_action({
serverId: "billing-api",
toolName: "delete_account",
args: { accountId: "acc_991", purgeInvoices: true }
})
// → { decision: "blocked", tier: "irreversible", ... }Note on deployment: if you run NitroWatch from NitroCloud rather than locally, it cannot reach
localhoston your machine — the cloud instance resolves its own localhost. Either run both locally, or deploy this server too and register its public URL.
Seed data
Two accounts and three invoices, held in memory. A restart resets everything — convenient when you have just deleted an account to prove a point.
Account | Invoices |
|
|
|
|
Built with
NitroStack TypeScript SDK — NitroStack × SRM Hackathon 2026.
Available Tools
5 toolsarchive_customerA
Archive a customer account so it no longer appears in active listings
| Name | Required | Description | Default |
|---|---|---|---|
| accountId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses a key behavioral outcome (no longer appears in active listings), but with no annotations provided, it carries the full burden. It does not mention reversibility, permissions, or effects on associated data like invoices. This leaves important behavioral traits undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, short sentence that front-loads the verb 'Archive' and provides the essential outcome. Every word is meaningful, with no redundancy or unnecessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of annotations and an output schema, the description should cover more context to be complete. It fails to mention return values, reversibility, or how archiving differs from permanent deletion (relevant due to the sibling delete_account). The tool is simple, but the description is too sparse to be fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema defines a single required 'accountId' parameter with no description. The tool description does not mention the parameter at all, and schema description coverage is 0%. Since the description fails to explain what the ID refers to or how to find it, it offers no additional value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Archive') and the target ('a customer account'), and specifies the consequence ('it no longer appears in active listings'). This effectively differentiates it from siblings like delete_account, which would imply permanent removal.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a use case: hide an account from active listings. However, it does not explicitly mention when to use this tool versus delete_account or any conditions that would make archiving preferable. The guidance is only implied through the word 'archive' rather than directly stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_accountA
Permanently delete a customer account and all associated invoices. Cannot be undone.
| Name | Required | Description | Default |
|---|---|---|---|
| accountId | Yes | ||
| purgeInvoices | No | Also destroy every invoice for this account |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and effectively discloses the destructive, irreversible behavior and the deletion of associated invoices. It does not mention permissions or edge cases, but the core behavioral traits are clearly stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the key information, and no waste. Very concise and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core behavior and irreversibility, but it lacks an explicit contrast with archive_customer (e.g., when to use delete vs archive) and does not clarify the conditional behavior of purgeInvoices relative to the claimed unconditional deletion of invoices. Given no output schema, return value details are absent but not critical for a simple delete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%; accountId has no description but is self-explanatory, and purgeInvoices is described in the schema. The tool description adds context about deleting invoices, but it does not reconcile the unconditional statement with the purgeInvoices parameter, leaving slight ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('delete') with a clear resource ('customer account') and scope ('all associated invoices'). It also states the irreversible nature, distinguishing it from archive_customer and invoice-specific tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Permanently delete' and 'Cannot be undone' gives clear context for finality, implying this is for hard deletion rather than archiving. However, it does not explicitly name alternatives like archive_customer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_invoiceA
Fetch a single invoice by id
| Name | Required | Description | Default |
|---|---|---|---|
| invoiceId | Yes | Invoice id, e.g. inv_2214 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must communicate behavioral traits. It only states the fetch action, which is read-only, but does not disclose any error handling, return format, or side-effect concerns, adding little beyond the name.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence with no unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-fetch tool with one parameter, the description is functional but leaves out details about return values or error conditions, especially given no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides a full description of invoiceId ('Invoice id, e.g. inv_2214') with 100% coverage, so the description adds no additional parameter semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the verb 'Fetch' and specifies the resource 'single invoice by id', clearly distinguishing it from sibling list_invoices that fetches multiple invoices.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly state when to use this tool versus list_invoices, though the 'single invoice by id' phrasing implies retrieval of one specific invoice. No alternatives or exclusions are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_invoicesA
List all invoices, optionally filtered by account
| Name | Required | Description | Default |
|---|---|---|---|
| account | No | Account id filter |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must carry the behavioral disclosure burden. It states that the tool lists all invoices, which implies a read-only operation, but it does not disclose potential side effects, pagination, rate limits, or return format. While 'List' conveys a safe read action, the lack of explicit behavior beyond that leaves some gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence: 'List all invoices, optionally filtered by account.' Every word contributes to the purpose, with no filler or redundancy. It is optimally concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with one optional parameter and no output schema. The description covers its core purpose and filter capability, which is sufficient for a basic list operation. It could benefit from mentioning response characteristics or scope verification, but given the low complexity, the description is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% because the 'account' parameter has a description ('Account id filter'). The tool description merely echoes this ('optionally filtered by account') without adding further contextual meaning such as value format, requiredness aside from optionality, or behavior when omitted. This meets the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List') and resource ('invoices'), clearly indicating the tool's function. It distinguishes from siblings by showing it retrieves multiple invoices ('all invoices') while get_invoice implies a single invoice and send_invoice implies a different action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool (listing invoices, with optional account filtering), but it does not explicitly mention when not to use it or name alternatives. The sibling tools are distinguishable by their names and descriptions, but the description itself lacks direct exclusions or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
send_invoiceB
Email an invoice to the customer. Can be voided and re-sent.
| Name | Required | Description | Default |
|---|---|---|---|
| No | Override the account email | ||
| invoiceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It does mention a reversibility trait ('Can be voided and re-sent'), which adds some transparency. However, it fails to disclose important side effects such as whether the invoice status changes, whether an email is actually sent immediately, permission requirements, or error behavior—critical gaps for a mutation tool with no annotation support.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is exactly two sentences, with the first sentence stating the core purpose and the second adding a key behavioral trait. There is no redundant text or filler, and the information is front-loaded. This is appropriately sized and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (one required parameter), and the description covers the basic action and a notable reversibility feature. However, with no output schema and no annotations, the description should have explained what happens after sending (e.g., return value, success/failure indications) and any constraints on invoice state. The note about voidability helps but is not enough to make the tool fully self-explanatory.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds no information about the parameters. The schema documents 'email' as an override but leaves 'invoiceId' entirely unspecified; with schema description coverage at 50%, the description should have compensated for the missing parameter explanation. The tool name implies invoiceId is the invoice to send, but the description does not clarify this or the relationship between email and invoiceId.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Email an invoice to the customer') with a specific verb and resource. It distinguishes itself from sibling tools (get_invoice, list_invoices, archive_customer, delete_account) which handle viewing, archiving, or deleting rather than sending. The additional note about being voidable and re-sent adds useful scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage is implied through the purpose (use this when you need to email an invoice), but there is no explicit guidance on when to use this tool versus alternatives, nor any exclusions or prerequisites. Sibling tools are listed but not referenced in the description, so the agent must infer the appropriate context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v1.0.0- First observed
archive_customer - First observed
delete_account - First observed
get_invoice - First observed
list_invoices - First observed
send_invoice
TDQS
Scored across 5 tools
Each tool targets a distinct resource and action: get_invoice retrieves a single invoice, list_invoices retrieves multiple, send_invoice emails an invoice, and archive_customer vs delete_account clearly separate soft from permanent customer removal. No two tools overlap in purpose.
All tool names follow a consistent verb_noun pattern (get/list/send/archive/delete + resource). The use of 'get' for singular and 'list' for plural is conventional, and there is no mixing of styles.
With 5 tools, the server is well-scoped for a billing demo. It covers the essential operations without unnecessary bloat, and the count feels appropriate for its purpose.
The set covers invoice retrieval, emailing, and customer lifecycle (archive/delete), but lacks invoice creation, updating, or voiding. This is a minor gap for a billing system, but core workflows are still represented.
Maintenance
Related MCP Connectors
Hosted MCP server for Mini Accountant: invoices, expenses, customers, analytics, tax estimates.
MCP server for Codat — companies, connections, invoices, bills and financial statements.
MCP server for Quaderno — tax-rate calculation, invoices, contacts, products, receipts & expenses.
MCP server for Autumn — read customers, plans, balances & invoices; track usage and attach plans.
Related MCP Servers
- FlicenseNot gradedqualityBmaintenanceDemo MCP server that exposes order and customer data as read-only tools for AI assistants, simulating a business API or internal data source.-
- FlicenseNot gradedqualityCmaintenanceAn intentionally vulnerable MCP server designed as a live demo target for the MCP Trust security scanner. It contains deliberate insecure patterns to demonstrate scanning capabilities.-
- AlicenseNot gradedqualityAmaintenanceMCP server for Askell's payment and subscription API, allowing users to discover API operations, make raw API calls with approval for mutations, and analyze customers, contracts, billing runs, and webhooks.430MIT
- FlicenseNot gradedqualityCmaintenanceEnables AI agents to generate audit reports and query financial ledger data via MCP. This demo version contains intentional security vulnerabilities for security testing and should not be used in production.-