US compliance and books health
Server Details
US company filing obligations, plus a 25-check bookkeeping diagnostic. Sourced, dated, no signup.
- Status
- Healthy
- Uptime
- 100.0% over 43 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- median-labs/median-compliance-skill
- GitHub Stars
- 3
- Server Listing
- median-compliance-skill
TDQS
Scored across 12 tools
Most tools have clearly distinct purposes and are separated by domain and action. The only mild ambiguity is between get_confounders and get_evidence_recipe, since both are per-obligation evidence/refinement tools, but their descriptions do clarify the different stage each serves.
Names consistently follow a verb_noun pattern: list_* for collections, explain_* for single-item detail, get_* for lookups, and score_* for the one scoring operation. Singular and plural nouns are used predictably to distinguish individual items from lists.
Twelve tools is within a reasonable range and each tool has a defined role. However, four of them are company content tools (blog, pricing, services, overview) that expand the server beyond its compliance and books-health core, making the set feel slightly broader than the server name suggests.
The core workflows are well covered: list applicable obligations and checks, explain individual items, refine with evidence guidance, and score books health. The main gap is that there is no compliance-health scoring analog to score_books_health, so compliance coverage is thorough but lacks an equivalent summary step.
Available Tools
12 toolsexplain_books_checkExplain one books checkARead-onlyIdempotentInspect
Full detail on a single books health check: what it is, exactly where to look in the ledger, every innocent explanation to rule out before concluding anything, what it costs if it is real, what fixing it involves, and whether it is inside what Median does.
| Name | Required | Description | Default |
|---|---|---|---|
| check_id | Yes | Check id from list_books_checks, e.g. 'rec-processor-balance-unreconciled'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark read-only and idempotent, and the description adds rich behavioral detail: what it explains (what it is, where to look, innocent explanations, cost, fixing, Median's scope). This goes well beyond the annotations, providing concrete expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, information-dense sentence that front-loads the core purpose and then enumerates the detailed aspects covered. No fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given one parameter, no output schema, and strong annotations, the description fully compensates by detailing what the agent will learn. It covers enough context for correct invocation and expectation setting.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers the single parameter check_id with a clear description and example. The tool description does not add extra parameter semantics, but with 100% schema coverage, baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool provides 'full detail on a single books health check' with a specific verb and resource. It distinguishes from siblings like list_books_checks (listing) and score_books_health (scoring) by focusing on deep explanation of one check.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context—for understanding a specific check in depth—but does not explicitly name alternatives or exclusion criteria like 'use list_books_checks to enumerate checks'. Clear context is present, but no explicit when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explain_obligationExplain one obligationARead-onlyIdempotentInspect
Plain-English detail on a single obligation: what it is, why it exists, the deadline, the cost, the penalty, the exact remediation steps, and the government source.
| Name | Required | Description | Default |
|---|---|---|---|
| obligation_id | Yes | Obligation id, e.g. 'de-registered-agent'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only and idempotent behavior, which the description does not contradict. The description adds details about the output content (deadline, cost, etc.) but does not provide additional behavioral context such as side effects, rate limits, or error handling. Thus it meets the baseline for annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that lists all key aspects in a concise manner without any fluff or redundancy. It is easy to parse and understand.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one parameter and no output schema, the description is highly complete. It specifies exactly what information will be returned (deadline, cost, penalty, etc.), making its purpose and output expectations clear without needing additional details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides full coverage for the only parameter (obligation_id) with a clear example. The description does not add any extra meaning to the parameter beyond what the schema already states, so the score is at the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool explains a single obligation and lists the specific details provided (what it is, why it exists, deadline, cost, penalty, remediation steps, government source). This distinguishes it from sibling tools that target different resources.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies the tool is for explaining an obligation, but it does not explicitly mention when to use it over alternative tools like 'explain_books_check' or mention exclusions. While the name and context make the usage clear, explicit alternative guidance would strengthen it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_blog_postGet blog postARead-onlyIdempotentInspect
Fetch the full text of one published Median blog post by its slug.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | Yes | Post slug, e.g. 'how-to-switch-bookkeeping-services'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so safety is covered. The description adds that the post must be published and that full text is returned, but it does not disclose behavior for nonexistent slugs or draft posts. This is acceptable but not richly transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, tightly worded sentence that front-loads the action, resource, and parameter. Every word earns its place, with no redundant phrasing or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read-only tool with no output schema, this description is complete. The agent knows exactly what to supply (slug) and what to expect (full text of a published post), and the annotations cover the behavioral safety profile.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the only parameter, slug, is already described with an example in the schema. The description references the slug as the lookup key but adds no new semantic detail beyond what the schema provides, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Fetch'), a precise resource ('full text of one published Median blog post'), and the identifying key ('by its slug'). This clearly distinguishes it from the sibling list_blog_posts, which would be used to enumerate posts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the intended context clear: use this tool when you need the full text of a single post and already know its slug. It does not explicitly name alternatives or exclusion conditions, but the wording and sibling list_blog_posts make the usage boundary obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_company_overviewGet company overviewARead-onlyIdempotentInspect
Who Median is, what it does, who it is for, how onboarding works, and how to get in touch or book an intro call.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered. The description adds content scope but no additional behavioral details such as response format, required authentication, or any side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that front-loads the most important content areas and wastes no words. Every phrase contributes meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only overview tool, the description sufficiently covers what callers can expect to learn. It does not describe return shape or format, but no output schema exists and the content enumeration is adequate for this simple use case.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline of 4 applies. There is no parameter detail to add beyond what the empty schema already communicates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly enumerates the content of the overview: company identity, product, audience, onboarding, and contact. It is specific enough to distinguish this from siblings like get_pricing or list_services, though it lacks an explicit action verb in the description itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance or mention of alternatives. The intended use is strongly implied by the tool name and the content list, but the description does not state when a caller should choose this over sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_confoundersGet confounding questionsARead-onlyIdempotentInspect
Refinement step, NOT the place to start: call list_compliance_obligations first, then call this to narrow what is still genuinely open. Returns the facts that flip an answer, such as whether an employer of record is the legal employer. Pass every fact you already know and it returns only what those facts leave unsettled. These are things for YOU to look up in the user's own documents. They are not a questionnaire to send the user.
| Name | Required | Description | Default |
|---|---|---|---|
| answers | No | Answers to the yes/no threshold questions this tool asks, keyed by the threshold id shown in the question, e.g. { "foreign-accounts-over-10k": false, "ny-sales-over-500k": false }. Without this the same questions repeat forever and the obligations behind them can never resolve. Omit a key you genuinely do not know. | |
| answered | No | Only for confounders with no matching fact field above (who-paid, prepaid-multi-year). For everything else pass the fact itself, e.g. payroll_model: 'eor', not answered: { 'payroll-model': 'eor' }. | |
| tax_year | No | Tax year under review. Defaults to last calendar year. | |
| sells_saas | No | Whether the product is SaaS. New York taxes SaaS as prewritten software. | |
| entity_type | No | Legal entity type. | |
| company_name | No | Optional, and the only identifying thing this tool records. Median stores it to see which companies use this tool and may follow up. It changes nothing about the answer, so omit it if the user has not agreed to share it. Send the legal entity name only, never an address, EIN or anything from a document. | |
| revenue_band | No | Gross annual revenue as a band, never an exact figure: 'pre-revenue', 'under-250k', '250k-1m', '1m-5m' or 'over-5m'. Drives revenue-scaled thresholds such as state franchise tax minimums and economic nexus. Leave unset if unknown; the tool asks for it rather than guessing. | |
| sales_states | No | States with customers. | |
| foreign_owned | No | Whether any non-US person owns 25% or more. Leave unset if genuinely unknown. | |
| payroll_model | No | Critical. An employer of record holds state registrations under its own entity, so the answer flips entirely on this. | |
| presence_type | No | Per-state presence, e.g. { NY: 'virtual-mailbox' }. | |
| registered_in | No | Whether the company is already registered to do business in a state, e.g. { NY: false }. Drives obligations that only begin at registration. | |
| formation_date | No | ISO date the entity was formed. | |
| employee_states | No | States where W-2 employees work. | |
| fiscal_year_end | No | MM-DD, e.g. '12-31'. Defaults to 12-31. | |
| formation_state | No | Two-letter state of formation, e.g. 'DE'. | |
| extensions_filed | No | Whether a tax extension was filed, keyed 'US' for federal and by state code, e.g. { US: true, NY: false }. OMIT a key you are unsure about: no deadline is asserted for an unknown key, because assuming an extension tells a late filer they have months in hand. | |
| operating_states | No | States where the company operates. | |
| payroll_provider | No | The provider's name only, e.g. 'Deel', 'Gusto', 'Rippling'. Not account numbers, not employee details, not anything else. | |
| contractor_states | No | States where 1099 contractors work. | |
| formation_platform | No | Decides whether a bundled first-year registered agent explains a missing fee. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true and idempotentHint=true, and the description adds behavioral context: it reveals that the tool returns only unsettled facts and that output should be investigated by the agent. It does not contradict the annotations and enriches them with workflow semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the critical workflow constraint ('NOT the place to start') and every sentence earns its place, covering sequencing, output purpose, input policy, and user interaction guidance without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite 21 parameters and no output schema, the description gives sufficient context for an agent to call the tool correctly: it explains when to call it, what to pass, what it returns, and how to use the results. It could be slightly more explicit about the exact output shape, but the conceptual return type is adequately conveyed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents every parameter in detail. The description adds a global instruction to 'pass every fact you already know,' but it does not add parameter-specific meaning beyond what the schema provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it 'narrows what is still genuinely open' and 'returns the facts that flip an answer.' It clearly distinguishes itself from list_compliance_obligations by naming it as the prerequisite, so an agent can tell this tool apart from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says 'NOT the place to start' and instructs calling list_compliance_obligations first, giving a clear ordering constraint. It also tells the agent to pass every known fact and clarifies that the results are for the agent to look up in user documents, not a questionnaire to send to the user.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_evidence_recipeGet evidence recipeARead-onlyIdempotentInspect
For one obligation, return exactly how to prove it: which sources settle the question, which are only corroborating, what each state's public registry does and does not expose, and the innocent explanations for missing evidence. Use this before concluding anything is missing.
| Name | Required | Description | Default |
|---|---|---|---|
| obligation_id | Yes | Obligation id from list_compliance_obligations. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered. The description adds substantial behavioral context by detailing the exact nature of the output: which sources settle, which are corroborating, what registries expose, and innocent explanations. This goes beyond what annotations provide, giving the agent a clear picture of what to expect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the primary action and object. Each clause adds value—scope, output content, and usage timing. There is no filler or repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has one simple parameter, good annotations, and no output schema, the description carries the full burden of explaining the return content. It does so thoroughly, listing specific categories of information the recipe includes. It also provides a contextual usage pointer. This is complete for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the sole parameter (obligation_id) is already clearly described as 'Obligation id from list_compliance_obligations.' The description only adds 'For one obligation,' which is redundant with the parameter name and schema. No additional syntax or semantics are provided, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action: 'For one obligation, return exactly how to prove it.' It identifies the resource (evidence recipe for an obligation) and enumerates the content (sources that settle, corroborating sources, registry exposure, innocent explanations). This distinguishes it from sibling tools like explain_obligation or get_confounders by focusing on evidence and proof.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on when to use the tool: 'Use this before concluding anything is missing.' This implies a specific workflow step. However, it does not explicitly name alternatives or state when not to use this tool, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_pricingGet pricingBRead-onlyIdempotentInspect
Median's published pricing: the flat monthly fee by revenue band under $1M in annual revenue, the per-account and per-ledger-entry meter that scoping is based on above $1M, what is included, and add-on pricing for tax, R&D credits, and sales tax.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this tool as read-only and idempotent, so the safety profile is clear. The description adds that the returned data is published pricing and describes its category breakdown, but it does not disclose any further behavioral detail such as whether the data is periodically stale, sourced externally, or the canonical source of truth. No contradiction exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single coherent sentence that leads with the resource name and then lists the pricing components in an organized manner. It has no filler, though it is slightly dense with multiple clauses packed in one sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only information tool with no output schema, the description covers the main pricing dimensions an agent needs. It does not include a detailed return shape, but none is expected here, and the content list is enough to judge relevance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no input schema meaning to supplement. The 0-param baseline of 4 applies, and the description appropriately focuses entirely on what the output represents rather than parameter handling.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly names the resource — Median's published pricing — and enumerates the exact content categories (flat monthly fee by revenue band, metered pricing above $1M, included features, add-ons). It is more specific than the title alone, though it does not explicitly distinguish itself from siblings or state a use verb beyond the tool name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance about when to call this tool, what question it answers, or when to prefer sibling tools such as get_company_overview or list_services. It names no exclusions and does not help an agent route between alternatives. All usage context has to be inferred from the pricing content.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_blog_postsList blog postsARead-onlyIdempotentInspect
List published articles from the Median blog, newest first. Optionally filter by category or a search term matched against the title.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many posts to return (default 20, max 50). | |
| search | No | Match against the post title. | |
| category | No | Filter to one category slug or name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and idempotentHint, and the description does not contradict them. It adds minimal behavioral context beyond the annotations, only mentioning ordering ('newest first') and optional filters, which mostly restate parameter definitions. No side effects, auth needs, or rate limits are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, using two sentences to convey the essential purpose and optional parameters without redundancy or fluff. It is well-structured and easy to understand.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list operation with no output schema required, the description is complete. It covers the core functionality (listing, ordering, filtering) and does not need to explain return types or pagination, as these are not specified in the schema. The simplicity of the tool makes this sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides descriptions for all three parameters, achieving 100% coverage. The tool description adds meaningful context by specifying that results are ordered newest first and explicitly stating that filters are optional, which supplements the schema. This slightly exceeds the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: listing published articles from the Median blog, newest first, with optional filters. It distinguishes itself from sibling tools by specifying the resource (blog posts) and the brand (Median), making it distinct from other list_* tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for retrieving multiple blog posts with filtering and ordering, but does not explicitly state when not to use it or mention alternatives like get_blog_post. However, the context of sibling tools (e.g., get_blog_post) makes the distinction implicit, and the description is self-explanatory.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_books_checksList books health checksARead-onlyIdempotentInspect
START HERE for a books health check. Given whatever you know about a business, return the bookkeeping checks that apply to it, grouped by dimension, each with where to look, the innocent explanations to rule out first, and what it costs if it is real. Call it immediately with partial facts rather than interviewing the user first: unknown facts keep checks in the list rather than dropping them, so nothing is silently skipped. Cash-basis businesses correctly get a shorter list. This never says whether books are correct, only what to look at.
| Name | Required | Description | Default |
|---|---|---|---|
| basis | No | Cash or accrual, printed at the top of most profit and loss reports. This matters more than anything else here: several revenue checks do not apply on a cash basis. | |
| ledger | No | Which ledger the books live in. Visible in the header of any export. | |
| has_raised | No | Whether the business has raised on a SAFE, note or priced round, or taken a loan. | |
| processors | No | Payment processors that hold a balance before paying out, e.g. ['stripe','paypal']. Names only. Drives the single most-missed check. | |
| bills_ahead | No | Whether anything is billed ahead of delivery: subscriptions, retainers, annual plans. | |
| company_name | No | Optional, and the only identifying thing this tool records. Median stores it to see which businesses use this tool and may follow up. It changes nothing about the answer, so omit it if the user has not agreed to share it. Send the business name only, never a balance, an account number or anything copied out of a document. | |
| payroll_model | No | How people are paid. | |
| has_contractors | No | Whether 1099 contractors are paid. | |
| charges_sales_tax | No | Whether sales tax is charged anywhere. | |
| needs_segment_reporting | No | Whether any part of the business needs separate reporting: locations, programs, product lines, more than one entity. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate read-only, idempotent, and closed-world behavior. The description adds valuable behavioral context: unknown facts keep checks in the list rather than dropping them, cash-basis shortens the list, and the tool makes no judgment on correctness. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with 'START HERE' and every sentence carries essential information: what it returns, when to call, how partial facts are handled, cash-basis behavior, and its non-judgmental nature. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description thoroughly explains the output structure (checks grouped by dimension, where to look, innocent explanations, cost) and covers behavioral constraints. With 10 optional parameters, it gives enough context for an agent to invoke the tool successfully.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with detailed parameter descriptions, so baseline is 3. The description adds meaningful parameter-agnostic guidance: partial facts are acceptable and unknown facts don't drop checks, which tells the agent that omitting optional params is safe and expected. This raises the score to 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action: 'return the bookkeeping checks that apply to it, grouped by dimension,' with clear output contents. It also distinguishes this tool from siblings like score_books_health and explain_books_check by explicitly saying it never says books are correct, only what to look at.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit usage guidance: 'START HERE for a books health check' and 'Call it immediately with partial facts rather than interviewing the user first.' It also provides exclusions: 'This never says whether books are correct' and mentions cash-basis businesses get a shorter list, which helps the agent decide when to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_compliance_obligationsList compliance obligationsARead-onlyIdempotentInspect
START HERE. Given whatever you know about a company, return the federal and state compliance obligations that apply, ordered by urgency with overdue items first, plus the specific questions that unblock the rest. Call it immediately with partial facts rather than interviewing the user first: it is designed for incomplete input, and every fact you are missing comes back as one precise question instead of the dozen generic ones you would otherwise ask. Then answer those and call it again. Defaults to a short actionable view; relay it roughly as written rather than expanding it into an essay, and pass detail='full' only when the user asks for evidence tiers and remediation detail. Covers US federal plus CA, CO, DE, FL, GA, IL, MA, NJ, NY, PA, TX and WA; every other state is reported honestly as not yet covered. This never tells you whether you are compliant, only what to check.
| Name | Required | Description | Default |
|---|---|---|---|
| detail | No | Default 'brief': what applies, when it is due, and the questions, with no evidence or remediation detail. Around a third the length of 'summary' and it is what a founder can actually act on. 'summary' adds one check and one next step per item. 'full' adds every evidence tier, innocent explanation and remediation step and runs to several thousand words, so use it only when the user asks for that depth. Prefer staying brief and calling get_evidence_recipe or explain_obligation for the one or two items that actually matter. | |
| answers | No | Answers to the yes/no threshold questions this tool asks, keyed by the threshold id shown in the question, e.g. { "foreign-accounts-over-10k": false, "ny-sales-over-500k": false }. Without this the same questions repeat forever and the obligations behind them can never resolve. Omit a key you genuinely do not know. | |
| tax_year | No | Tax year under review. Defaults to last calendar year. | |
| sells_saas | No | Whether the product is SaaS. New York taxes SaaS as prewritten software. | |
| entity_type | No | Legal entity type. | |
| company_name | No | Optional, and the only identifying thing this tool records. Median stores it to see which companies use this tool and may follow up. It changes nothing about the answer, so omit it if the user has not agreed to share it. Send the legal entity name only, never an address, EIN or anything from a document. | |
| revenue_band | No | Gross annual revenue as a band, never an exact figure: 'pre-revenue', 'under-250k', '250k-1m', '1m-5m' or 'over-5m'. Drives revenue-scaled thresholds such as state franchise tax minimums and economic nexus. Leave unset if unknown; the tool asks for it rather than guessing. | |
| sales_states | No | States with customers. | |
| foreign_owned | No | Whether any non-US person owns 25% or more. Leave unset if genuinely unknown. | |
| payroll_model | No | Critical. An employer of record holds state registrations under its own entity, so the answer flips entirely on this. | |
| presence_type | No | Per-state presence, e.g. { NY: 'virtual-mailbox' }. | |
| registered_in | No | Whether the company is already registered to do business in a state, e.g. { NY: false }. Drives obligations that only begin at registration. | |
| formation_date | No | ISO date the entity was formed. | |
| employee_states | No | States where W-2 employees work. | |
| fiscal_year_end | No | MM-DD, e.g. '12-31'. Defaults to 12-31. | |
| formation_state | No | Two-letter state of formation, e.g. 'DE'. | |
| extensions_filed | No | Whether a tax extension was filed, keyed 'US' for federal and by state code, e.g. { US: true, NY: false }. OMIT a key you are unsure about: no deadline is asserted for an unknown key, because assuming an extension tells a late filer they have months in hand. | |
| operating_states | No | States where the company operates. | |
| payroll_provider | No | The provider's name only, e.g. 'Deel', 'Gusto', 'Rippling'. Not account numbers, not employee details, not anything else. | |
| contractor_states | No | States where 1099 contractors work. | |
| formation_platform | No | Decides whether a bundled first-year registered agent explains a missing fee. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and idempotentHint, but the description adds substantial behavioral context beyond that: it never tells whether the company is compliant, it honestly reports uncovered states, it returns precise questions for missing facts, and it asks the agent to relay the short view roughly as written. These are valuable operational constraints not present in the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but densely informative, with no filler. It front-loads the critical 'START HERE' directive and core behavior, then layers usage policy, coverage limits, and disclaimers. Given 21 parameters and a complex iterative workflow, the length is justified and well organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description provides a reasonable account of what the tool returns: compliance obligations, ordering, questions, and a default short actionable view. It also covers state coverage and the compliance-check disclaimer. It could specify the exact return structure, but for a text-generation tool of this complexity, the description is adequately complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The tool description goes beyond the schema by explaining the workflow connection between the answers parameter and the repeated-question loop, and by clarifying when to use detail='full' versus the default. It doesn't enumerate parameters, but the schema already does that thoroughly; the description adds meaningful usage context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb and resource: 'return the federal and state compliance obligations that apply, ordered by urgency with overdue items first, plus the specific questions that unblock the rest.' It also clearly distinguishes itself from siblings by positioning itself as the 'START HERE' entry point and pointing to explain_obligation and get_evidence_recipe for follow-up detail.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance: call immediately with partial facts rather than interviewing the user, designed for incomplete input, then answer the returned questions and call again. It also states when not to use it (pass detail='full' only when the user asks) and names alternatives (get_evidence_recipe, explain_obligation) for selective deep-dives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_servicesList servicesARead-onlyIdempotentInspect
List the services Median offers (daily bookkeeping, tax filing, R&D tax credits, CFO advisory) with a short summary and page link for each.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the tool is read-only, non-open-world, and idempotent, so the description doesn't need to restate those. It adds value by specifying the exact content of the response (four service categories, short summary, page link per service), which is beyond the annotations and helps the agent anticipate the output without an output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence that front-loads the essential purpose. Every word counts—it lists the services and the output detail (summary and link) without redundancy. There is no fluff or irrelevant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (zero parameters, no output schema), the description is complete. It fully explains what the tool returns and names the specific services, making it actionable for an agent. The annotations cover safety and determinism, so no additional behavioral warnings are needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has no parameters, so the description carries the full burden of explaining what the tool does and what the response contains. It does this well, mentioning the types of services and the structure (summary + link), which is more than sufficient for a zero-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'list' and the specific resource 'services offered by Median', enumerating the exact categories of services and what the result includes (summary and page link). This distinguishes it from sibling tools like get_pricing or get_company_overview.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by specifying what services are listed, but it does not explicitly state when to prefer this tool over alternatives or when not to use it. However, a zero-parameter tool with a very specific purpose (listing services) inherently separates itself from siblings; the lack of explicit alternatives is acceptable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
score_books_healthScore books healthARead-onlyIdempotentInspect
Score a books health check from the statuses you actually evidenced. Returns the weighted score, the band, a per-dimension breakdown and the findings ranked worst first. Two rules a plain average does not give you: a confirmed critical caps the grade, and the band is WITHHELD entirely when too little was evidenced, because a flattering number off three checks is the part people quote. Report the withheld state rather than supplying a band yourself.
| Name | Required | Description | Default |
|---|---|---|---|
| basis | No | Cash or accrual, printed at the top of most profit and loss reports. This matters more than anything else here: several revenue checks do not apply on a cash basis. | |
| ledger | No | Which ledger the books live in. Visible in the header of any export. | |
| results | Yes | One entry per check you actually worked. Send check ids and statuses only: never figures, balances, account numbers or names. | |
| has_raised | No | Whether the business has raised on a SAFE, note or priced round, or taken a loan. | |
| processors | No | Payment processors that hold a balance before paying out, e.g. ['stripe','paypal']. Names only. Drives the single most-missed check. | |
| bills_ahead | No | Whether anything is billed ahead of delivery: subscriptions, retainers, annual plans. | |
| company_name | No | Optional, and the only identifying thing this tool records. Median stores it to see which businesses use this tool and may follow up. It changes nothing about the answer, so omit it if the user has not agreed to share it. Send the business name only, never a balance, an account number or anything copied out of a document. | |
| payroll_model | No | How people are paid. | |
| has_contractors | No | Whether 1099 contractors are paid. | |
| charges_sales_tax | No | Whether sales tax is charged anywhere. | |
| needs_segment_reporting | No | Whether any part of the business needs separate reporting: locations, programs, product lines, more than one entity. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations declare readOnlyHint=true and idempotentHint=true, so the agent knows this is a safe read operation. The description adds significant behavioral context beyond that: it reveals that a confirmed critical caps the grade, that the band is withheld when evidence is insufficient, and that the tool returns a ranked findings list. It also warns the agent not to fabricate a band ('Report the withheld state rather than supplying a band yourself'), which is critical for correct use. No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise but packed with essential information: it states the purpose, outputs, two critical rules, and a key usage caveat in just three sentences. It front-loads the purpose and output list, then adds behavioral rules. Every sentence earns its place, with no redundancy or filler. Ideal length for a tool with moderate complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (11 parameters, multiple rules, no output schema), the description is complete: it explains the purpose, output structure, special scoring rules (critical cap and withholding), and the necessity of reporting withheld state. The schema covers parameter details, so the description doesn't need to repeat them. The absence of an output schema is compensated by the description's clear list of return elements. The only minor gap is that it doesn't explicitly state input validation rules (e.g., minItems), but the schema already handles that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, with all 11 parameters documented with rich descriptions in the schema. The tool description itself does not add further parameter-specific semantics, but it does instruct the agent to send only check ids and statuses, never figures or balances, which reinforces the 'results' parameter. Given high schema coverage, the baseline of 3 is appropriate; the description doesn't need to do more, but it doesn't go beyond the schema either.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Score a books health check') and clearly states the resource and what the tool does—it computes a weighted health score from evidenced statuses. It distinguishes itself from sibling tools like list_books_checks and explain_books_check by detailing the output (weighted score, band, per-dimension breakdown, findings ranked worst first) and adding unique rules. The title 'Score books health' is directly expanded with actionable detail, making it clear and non-tautological.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells the agent when to use this tool: when scoring a health check from actually evidenced statuses. It also provides crucial when-not guidance by stating 'Report the withheld state rather than supplying a band yourself,' preventing misuse. While it doesn't name alternative tools directly, it implies that this tool is for computing a score rather than explaining checks or listing obligations, which is sufficient given the sibling context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
- Changed
get_blog_post1 field changed- changed
Input schema / properties / slug / descriptionPrevious value: -"Post slug, e.g. 'daily-close-for-startups'."New value: +"Post slug, e.g. 'how-to-switch-bookkeeping-services'."
2 tool updates
- Changed
get_confounders1 field changed- added
Input schema / properties / revenue_band / descriptionAdded value: +"Gross annual revenue as a band, never an exact figure: 'pre-revenue', 'under-250k', '250k-1m', '1m-5m' or 'over-5m'. Drives revenue-scaled thresholds such as state franchise tax minimums and economic nexus. Leave unset if unknown; the tool asks for it rather than guessing."
- Changed
list_compliance_obligations1 field changed- added
Input schema / properties / revenue_band / descriptionAdded value: +"Gross annual revenue as a band, never an exact figure: 'pre-revenue', 'under-250k', '250k-1m', '1m-5m' or 'over-5m'. Drives revenue-scaled thresholds such as state franchise tax minimums and economic nexus. Leave unset if unknown; the tool asks for it rather than guessing."
12 tool updates
- First observed
explain_books_check - First observed
explain_obligation - First observed
get_blog_post - First observed
get_company_overview - First observed
get_confounders - First observed
get_evidence_recipe - First observed
get_pricing - First observed
list_blog_posts - First observed
list_books_checks - First observed
list_compliance_obligations - First observed
list_services - First observed
score_books_health
Related MCP Connectors
US small-biz & foreign-LLC filing deadlines: FinCEN BOI, IRS 5472, franchise tax, 1099 (sourced).
US LLC filing fees, deadlines, and name checks for all 50 states — from official sources.
Generate small-business compliance calendars and renewal checklists.
Check micro-entity company accounts: raw figures in, validated balance sheet and deadlines out.
Related MCP Servers
- FlicenseNot gradedqualityCmaintenanceHelps small business owners classify business documents and messages, extract key information like dates and amounts, generate action checklists, check missing documents, and draft business messages for accountants, clients, and employees.-
- AlicenseBqualityAmaintenance39 tax tools for US individual taxpayers — federal/state tax calculations, credits, deductions, retirement strategies, audit risk, and tax planning. All calculations run locally, no data leaves the machine. Supports TY2024 and TY2025 (One Big Beautiful Bill Act).44323 npm12MIT
- AlicenseNot gradedqualityBmaintenanceCurrent, source-cited US federal tax constants and freelancer calculators for tax year 2026, including the July 1 mid-year mileage change. Every response carries its IRS/SSA primary source and a last-verified date; refuses rather than guesses.MIT
- FlicenseNot gradedqualityDmaintenanceAutomates comprehensive accounting workflows including bookkeeping, tax planning, payroll processing, sales tax compliance, and client management. Integrates with QuickBooks and processes financial documents with AI-powered transaction categorization and compliance monitoring.1-
Glama MCP Gateway
Add one secure layer between your agents and this server.