proofread-mcp
proofread-mcp is an MCP server that checks case citations in drafts against proofread.law's registers of ~10M US court opinions and 1.1M Swiss decisions, and suggests leading cases.
Check citations in prose (
check_citations): paste text (up to 2M chars); returns coverage statement, per-tier counts, flagged rows, citation count and a report id. Optionaldeep: truealso judges whether each cited opinion supports the sentence it is cited for.Check citations in a file (
check_document): same result for a local.pdf,.docx,.txtor.mdfile up to 10 MB (.docxneeds a paid plan).Look up one citation (
resolve_citation): answers whether a case sits at a citation (found, ambiguous, not in register, cannot verify, known citation, or no citation recognised), plus that volume's coverage.Look up a list of citations (
resolve_citations): up to 500 citation strings in one call, one line per citation in input order.See what's covered (
coverage): the coverage statement and storage notice;jurisdiction: "ch"gives the Swiss register, courts held and index share.Render a report (
render_report): markdown diligence report from a report id or full report JSON, with every row.Suggest Swiss leading cases (
suggest_cases): from a Swiss federal statute article (e.g.Art. 41 OR) or a paragraph citing one, get a ranked list of leading BGE cases with field of law, quoted passage, practice changes and links; filters for domain, language andk.Suggest US cases (beta) (
suggest_cases): from a sentence stating a rule of law, get up to 3 cases whose text states it, with the matched passage, how the case binds a named court, and opinion links.Save and manage briefs (opt-in):
save_brief,list_briefs,get_brief,update_brief(edit-and-recheck loop, keeps last 20 versions) anddelete_brief; stored encrypted in the user's proofread.law account and requires an API key.Account and billing:
sign_upcreates an account and API key (shown once, adopted for the session);billing_linkreturns a Stripe Checkout link for payg, solo or firm plans.Limits to know: cannot resolve Westlaw/Lexis identifiers, check statutes, regulations or secondary sources, or say whether a case is still good law; a red row means "check this," never "this case does not exist."
Generates Stripe Checkout links for upgrading a proofread.law account to a paid plan (pay as you go, solo, or firm).
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@proofread-mcpCheck the case citations in this brief before we file."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
proofread-mcp
An MCP server for proofread.law. It lets Claude Desktop, Claude Code, Cursor and the OpenAI Agents SDK check case citations before a draft is filed: about 10 million US court opinions, and 1.1 million Swiss decisions (BGE/ATF/DTF, Federal Supreme Court dockets, the federal courts and all 26 cantons). The jurisdiction is detected from the draft, and a Swiss report comes back in the draft's language: German, French or Italian.
proofread.law checks each citation against an open register of about 10 million court opinions (CourtListener bulk data). Every result says what was checked, what was found, and what the register cannot see.
What it does
Fourteen tools:
Tool | Input | What comes back |
| text, | The coverage statement, counts per tier, one line per row that needs a human (red first, then orange, then deep-check rows to review), the number of citations found, a report id |
| path to a | The same, for a file on disk |
| one citation string | The register's answer for that citation: found (case, court, date, parallel citations, link), ambiguous (candidates), not in the register, cannot verify, known citation, or no citation recognised; the coverage of that volume; the coverage statement |
| a list of up to 500 citation strings | Counts by status, one line per citation in input order, the coverage statement |
|
| The coverage statement and the storage notice; |
| a report id from a previous check, or the full report JSON | A markdown diligence report with every row |
| a Swiss federal statute article or a paragraph that cites one, or a sentence of US law; | Swiss: the leading Federal Supreme Court cases (BGE) cited with the article, to read, in one ranked list, each labelled with its field of law and given with the quoted passage, any later change of practice, the decision's link and a link to check the citation; in the query's language. US (beta): up to 3 cases whose own text states the sentence, each with the passage the model matched, the paragraphs around it, how it binds the court you name, the opinion's link and a link to check the citation |
| text, | Saves the brief to the user's account (opt-in, stored encrypted) and checks it: the brief id, then the same compact result as |
| none | The saved briefs: id, title, when saved and last checked, counts per tier, number of versions |
| id, | The saved text, the latest report and the versions kept; with |
| id, | Saves the edited text as a new version and re-checks it: the flags resolved since the previous version, the new flags (each with its headline), the unchanged count, then every row that still needs attention. |
| id | Permanently deletes the brief and its versions |
| the account owner's email, a name for the agent | A proofread.law account and an API key (shown once); the server uses it for the rest of the session |
|
| A Stripe Checkout link for the account owner; needs an API key |
The compact result of a check is capped at about 12,000 characters; when a long brief has more flagged rows than fit, the text says how many were left out and render_report has them all.
Saved briefs (opt-in)
Nothing is saved unless save_brief or update_brief is called; check_citations and check_document never save anything. A saved brief is stored encrypted in the user's proofread.law account until delete_brief removes it, so the brief tools need an API key (any PROOFREAD_API_KEY, or one from sign_up); without one they answer with a message and send nothing. A server with storage switched off answers storage_off, and the tools say so.
The loop for fixing flagged citations: save_brief once, fix a flagged row in the text (the citation, the case name, the quotation, or take the citation out), call update_brief with the id and the whole edited text, and read what changed: which flags were resolved (flagged before, not flagged now), which are new, and the rows that still need attention. The last 20 versions are kept; get_brief lists them and reads an earlier one. Ids are integers in the API and strings in the tools.
check_citations and check_document read prose: they compare the case name and any quotation with the register. resolve_citation and resolve_citations look the citation string up in the register (the /v1/resolve API) and tell you which case sits there; they do not compare it with the name you have.
Tiers, in the words the tools use:
Tier | Word | Meaning |
red | check this | The register holds something concrete that disagrees: a different case at that citation, a volume or page that does not match, quoted words not in the opinion |
orange | cannot verify | Nothing to check against: a Westlaw or Lexis identifier, a volume newer than the register, a reporter the register holds only in part. Not evidence either way |
green | found | The citation resolves to a case whose caption matches |
white | deep check | With |
What it cannot do: resolve Westlaw (WL) or Lexis identifiers, check statutes, regulations or secondary sources, or say whether a case is still good law. Those limits are stated in every tool description and in the coverage statement that comes with every result.
Case suggestions (Switzerland)
suggest_cases takes a Swiss federal statute article (Art. 41 OR, art. 41 CO, Art. 8 ZGB, art. 9 Cst.; lowercase and dotless forms such as art 41 or are read too) or a paragraph that cites articles, and lists the leading Federal Supreme Court cases (BGE) the court cites with it. Federal acts only; cantonal law is not covered. There is no free-text search: a query without an article answers with a message that no article was recognised. The answer is one ranked list in the measured order, each case labelled with its field of law; the domain filter keeps the same order within one field, so a filtered row keeps its overall rank (Rang 7, Rang 9, ...) and a filter note says which field is shown. Each row has the citation, date, field, the decision's language, its rank, the quoted passage (regeste or consideration), any later change of practice (a changed precedent is listed with its flag, never dropped), the decision at the court's site and a link that opens the check on proofread.law. Every string comes back in the query's language (German, French or Italian; lang: "en" for English), and the tool prints it as it is. A suggestion is a case to read: it has not been checked against your sentence, and a row without a flag is not evidence that its practice still holds. To check a citation, use check_citations.
Arguments: query (required, up to 20,000 characters; over 1,500 it is sent as POST), domain (all, civil, criminal, public or social; default all), lang (de, fr, it or en; default the query's language), k (number of rows, 1 to 50; default 10). With a large k, rows past about 20,000 characters of text keep their header line and any practice flag and leave out the passage and links, which stay in the structured result (with the counts per field). During the trial phase it needs a paid-plan or trial API key; otherwise the API answers plan_required and the tool says where to upgrade.
Example call, suggest_cases with {"query": "Art. 41 OR, Art. 97 OR", "k": 3} (dev instance, 2026-09-26; rows 2 and 3 cut):
Gelesen als: Art. 41 OR, Art. 97 OR
1. BGE 146 IV 76, 13.11.2019, Strafrecht, FR, Rang 1
Das Bundesgericht zitiert diesen Entscheid zusammen mit Art. 41 OR; der Entscheid selbst nennt den Artikel nicht. Aus der Regeste: «a) Art. 110 Abs. 1 StGB; Art. 118, 121 Abs. 1 und 382 Abs. 1 StPO; ...»
Entscheid öffnen: https://search.bger.ch/ext/eurospider/live/de/php/clir/http/index.php?highlight_docid=atf%3A%2F%2F146-IV-76%3Ade&lang=de&type=show_document
Zitat prüfen: https://proofread.law/?cite=Art.%2041%20OR%3B%20BGE%20146%20IV%2076#check
...
Die Vorschläge sind publizierte Leitentscheide (BGE). Sie sind danach gereiht, wie oft das Bundesgericht sie zusammen mit dem Artikel zitiert, [...] Ein Vorschlag ist ein Entscheid zum Lesen: Ob er Ihre Aussage stützt, prüft diese Liste nicht, und ein Entscheid ohne Hinweis ist kein Beleg dafür, dass seine Praxis weiter gilt.With "domain": "civil" the list starts with the filter note and keeps the overall ranks:
Gelesen als: Art. 41 OR
Filter: Zivilrecht. Die Reihenfolge ist dieselbe wie in der ganzen Liste.
1. BGE 132 III 122, 13.09.2005, Zivilrecht, FR, Rang 7
Regeste: «Rechtmässigkeit von im Arbeitskampf eingesetzten Mitteln (Art. 28 BV; Art. 41 und 357a OR). ...»
...A filter that leaves no rows answers with a message saying so. A later change of practice appears right under the row it concerns, for example Praxisänderung durch BGE 145 III 1, möglicherweise nur teilweise: <link> followed by the regeste sentence that marks it.
Case suggestions (United States, beta)
suggest_cases also takes a sentence from a US draft that states a rule of law (one sentence, or a short paragraph up to 1,500 characters) and lists up to 3 cases whose own text states it. A model reads the opinion of the court in each of the top 10 candidates, and a case is shown only when the model finds a passage in it that states the sentence. Each row has the reference, the court, the year, how the case stands to the court you name (binding here, same circuit, persuasive), any later history the citator found (reversed in part, superseded by statute), the passage the model matched with the paragraphs before and after it (each cut to about 600 characters), a link to the opinion on CourtListener and a link that opens the check on proofread.law. Cases that a later court overruled or reversed are left out, and a note names them.
This is a beta. Every answer starts with a line that says so, with the measured numbers: on the site's own index, of 94 cases shown for 50 sentences from filed briefs, 73 (78%) support the sentence and 4 do not, as read blind by two model readers, not practising lawyers (details). Read the passage before you cite the case. A suggestion is a case to read; to check a citation, use check_citations.
Arguments: court (optional, up to 200 characters) is the court the brief is filed in: a federal court of appeals (9th Cir., Ninth Circuit, ca9), a district court (N.D. Cal., S.D.N.Y., cand), or a state or territory name or postal code (California, CA) for the federal court there. Cases that bind that court come first; without a court, Supreme Court cases come first. State courts are not read yet; a court the API does not read comes back as an error that says what it reads. jurisdiction is auto by default: a Swiss statute article, or German, French or Italian text, is Swiss, and anything else is a US sentence; us or ch sets it. domain, lang and k apply to Swiss answers only, and a US answer is in English.
Input that states no rule of law (a question, a citation or case name, a heading, a statement about the record or a party's argument, or more than 1,500 characters) gets no list. The answer says what to paste instead and gives an example sentence, and it costs nothing. A US answer that ran the model counts as one deep-checked citation, not a resolve. When the day's budget for model calls is spent, US suggestions are back the next day (UTC).
Example call, suggest_cases with a sentence on the plausibility pleading standard and "court": "N.D. Cal." (dev instance, 2026-09-26; the passage and the paragraphs around it shortened, rows 2 and 3 cut):
Beta: Each case is shown with the passage we matched; read it before you cite the case. Measured on the site's own index: 78% of shown cases support the sentence, and 4 of 94 did not (details on https://proofread.law/measurements#suggestions).
Ordered for N.D. Cal. (Ninth Circuit): binding cases first.
1. Bell Atlantic Corp. v. Twombly, 550 U.S. 544 (2007), Supreme Court, 2007, Binding here (Supreme Court)
The passage we matched: "Here, in contrast, we do not require heightened fact pleading of specifics, but only enough facts to state a claim to relief that is plausible on its face. Because the plaintiffs here have not nudged ..."
Before: "Plaintiffs say that our analysis runs counter to Swierkiewicz, 534 U. S., at 508 , which held that “a complaint in an employment discrimination lawsuit [need] n..."
After: "The judgment of the Court of Appeals for the Second Circuit is reversed, and the case is remanded for further proceedings consistent with this opinion...."
Read the opinion: https://www.courtlistener.com/opinion/145730/bell-atlantic-corp-v-twombly/
Check this citation: https://proofread.law/?cite=Bell%20Atlantic%20Corp.%20v.%20Twombly%2C%20550%20U.S.%20544%20%282007%29#check
2. Ashcroft v. Iqbal, 556 U.S. 662 (2009), Supreme Court, 2009, Binding here (Supreme Court)
...
3. Greg Landers v. Quality Communications, Inc., 771 F.3d 638 (9th Cir. 2014), Court of Appeals for the Ninth Circuit, 2014, Binding here
...
How we build this list: we search the sentences in which federal appellate courts and the Supreme Court cite a case, and the cases' own text. [...] We leave out cases that a later court overruled or reversed.A question gets the beta line, then the API's message and an example:
This reads as a question. Paste the rule itself, as a sentence, and we will look for cases that state it.
For example: "A complaint must contain sufficient factual matter, accepted as true, to state a claim to relief that is plausible on its face."Related MCP server: CourtListener Citation Validation MCP Server
Install
Needs Node 20 or newer. No install step is required; npx fetches proofread-mcp from npm.
Claude Desktop (extension bundle)
Download proofread-mcp-<version>.mcpb from the latest GitHub release and open it with Claude Desktop (double-click, or Settings, Extensions, Install Extension). The bundle carries the server and its dependencies and uses the Node runtime that ships with Claude Desktop, so nothing else has to be installed. The only setting is the optional API key, stored by Claude Desktop as a sensitive value and passed to the server as PROOFREAD_API_KEY.
To build the bundle yourself: scripts/build-mcpb.sh (needs npx @anthropic-ai/mcpb) writes build/proofread-mcp-<version>.mcpb from manifest.json, icon.png and the npm package contents.
Claude Desktop (config file)
Edit claude_desktop_config.json (Settings, Developer, Edit Config):
{
"mcpServers": {
"proofread": {
"command": "npx",
"args": ["-y", "proofread-mcp"],
"env": {
"PROOFREAD_API_KEY": "pl_..."
}
}
}
}Leave out env to use the free tier.
Claude Code
claude mcp add proofread -- npx -y proofread-mcp
# with an API key:
claude mcp add proofread -e PROOFREAD_API_KEY=pl_... -- npx -y proofread-mcpCursor
Settings, MCP, Add new global MCP server, or write .cursor/mcp.json in the project:
{
"mcpServers": {
"proofread": {
"command": "npx",
"args": ["-y", "proofread-mcp"],
"env": { "PROOFREAD_API_KEY": "pl_..." }
}
}
}OpenAI Agents SDK (over HTTP)
Start the server with the HTTP transport:
PROOFREAD_API_KEY=pl_... npx proofread-mcp --http --port 3333
# MCP endpoint: http://127.0.0.1:3333/mcp health: http://127.0.0.1:3333/healthThen connect from the Agents SDK:
from agents import Agent, Runner
from agents.mcp import MCPServerStreamableHttp
# The SDK's default read timeout is 5 s. A check of a long brief takes up to 6 s and a deep check 1 to 2 s per citation,
# so give the session up to 15 minutes (the API's own deep-check limit).
async with MCPServerStreamableHttp(params={"url": "http://127.0.0.1:3333/mcp"}, client_session_timeout_seconds=900) as proofread:
agent = Agent(name="Drafting assistant", instructions="Check every case citation before you rely on it.", mcp_servers=[proofread])
result = await Runner.run(agent, "Check the citations in this paragraph: ...")import { Agent, run, MCPServerStreamableHttp } from "@openai/agents";
const proofread = new MCPServerStreamableHttp({ url: "http://127.0.0.1:3333/mcp", name: "proofread" });
await proofread.connect();
try {
const agent = new Agent({ name: "Drafting assistant", mcpServers: [proofread] });
const result = await run(agent, "Check the citations in this paragraph: ...");
} finally {
await proofread.close();
}The HTTP server binds to 127.0.0.1 by default and is single-tenant by design: it has no authentication of its own, report ids are shared across sessions, and a key from sign_up is adopted by the whole process. Sessions that stay idle for 30 minutes are closed. To expose it on a network use --host 0.0.0.0 and put it behind something that adds authentication.
Any MCP client
stdio: run proofread-mcp. Streamable HTTP: run proofread-mcp --http --port 3333 and point the client at /mcp.
Environment variables
Variable | Default | Meaning |
| unset | An API key ( |
|
| Base URL, for a self-hosted or test instance |
Accounts and keys
An agent can open an account itself: sign_up posts the owner's real email and a name to POST /agent/signup and gets a key back, shown once. The server uses that key for the rest of the session; put it in PROOFREAD_API_KEY to keep it. The owner receives one confirmation email. When a quota is used up (a tool answers "needs the ... plan" or "monthly allowance used"), billing_link returns a Stripe Checkout link for the owner; nothing is charged until they pay. Plans and prices: proofread.law/pricing. The onboarding text the API publishes for agents is at proofread.law/agent/onboarding.md.
Free tier
Per month, per IP address without a key or per account with a free-tier key:
Tools | Quota |
| 20 checks, of which 3 may be deep checks |
| each runs a default check and counts as one of those checks; a title alone is not counted. Each plan has a cap on saved briefs |
| 1,000 resolves (each citation in a list counts as one) |
| paid plans and trials only during the trial phase; a Swiss answer counts as one resolve (a query without an article costs nothing); a US answer that ran the model counts as one deep-checked citation (input that states no rule of law costs nothing) |
| free, not counted |
There is also a limit of 20 requests an hour per IP (more on paid plans). When a limit is reached the tool returns a plain message with the retry time or the upgrade link; nothing is thrown at the protocol level.
.docx upload and unlimited checks need a paid plan. See proofread.law/pricing.
Privacy Policy
The full policy is at proofread.law/privacy. What applies to this server:
What is collected. The text or file you check leaves your machine and goes to proofread.law's own server (
PROOFREAD_API, defaulthttps://proofread.law), over HTTPS, in the request that checks it.save_briefandupdate_briefsend the brief's text and title the same way;suggest_casessends the query (an article, a sentence, or the paragraph you give it) and the court, if you name one.sign_upsends the account owner's email address and the agent name you give it. Nothing else is sent: no conversation history, no other files, no telemetry.Use and storage on proofread.law. A checked input is processed in memory and discarded when the report is returned; no copy is written to disk. The exception is opt-in: a brief saved with
save_brieforupdate_briefis stored encrypted in the user's account, with its reports and earlier versions, untildelete_briefdeletes it (permanently). No other tool saves anything. The log line per request carries the kind of input, its size, the number of citations, the tier counts and the time taken, never a citation, a party name or a word of text. With an account, proofread.law keeps the email address, the plan, a monthly count of checks and a hash of each API key.Third parties. A citation string the register cannot resolve may be looked up in CourtListener's citation API (the citation string only, never prose). Deep check (
deep: true) is opt-in: the clause before each citation (up to 700 characters) and the cited opinion go to the model judge. A USsuggest_casesquery (the beta) goes to the same model judge, with the opinions of the candidate cases. Those two are the only cases in which words you send leave proofread.law. Sign-in links go through Resend; payments through Stripe, which sees the card and proofread.law does not. Requests pass through Cloudflare's edge in transit.Retention. Inputs and reports: none, except saved briefs, which are kept until the user deletes them. Account records: until the owner asks for deletion at privacy@proofread.law.
This server. It stores nothing on disk, saved briefs included (they live in the proofread.law account, not here). It keeps the last 50 reports in memory so
render_reportcan be called with a short id; they are gone when the process exits. The API key lives in your MCP client's configuration (Claude Desktop stores the extension's key as a sensitive setting); a key fromsign_upis held in memory for the session and shown to you once so you can store it. It is sent only toPROOFREAD_API, as anAuthorization: Bearerheader. If a deep-check stream stops before every citation was judged, the tool answers with an error that says how many were checked; it never presents a partial deep check as a finished one.Contact. Data Alchemy Labs, privacy@proofread.law.
The coverage caveat
Every result starts with the coverage statement, for example:
Checked against 10.1 M cases (CourtListener bulk data 2026-06-30, last refreshed 2026-09-19); federal appellate 2019 to 2023 is 10 to 15% incomplete; Westlaw/Lexis identifiers are not resolvable; statutes, regulations and secondary sources are not checked.
Read it. A citation that is not in the register is a register fact with a coverage qualifier, not proof that the case does not exist. A red row says "check this"; the tools say what was checked and what was found, never that a case is invented.
Example
Input:
Title VII forbids discrimination because of sexual orientation. Bostock v. Clayton County, 509 U.S. 644 (2020).
check_citations returns (real output, review instance, 2026-09-20):
Coverage: Checked against 10.1 M cases (CourtListener bulk data 2026-06-30, last refreshed 2026-09-19); federal appellate 2019 to 2023 is 10 to 15% incomplete; Westlaw/Lexis identifiers are not resolvable; statutes, regulations and secondary sources are not checked.
Summary: 1 citation, 1 row. Check this (red): 1. Cannot verify (orange): 0. Found (green): 0.
Flagged rows:
- CHECK THIS: 509 U.S. 644 (Bostock v. Clayton County). Register has Bostock v. Clayton County at 590 U.S. 644. Check the volume. In the register, 509 U.S. 644 is Shaw v. Reno. The case named in the document exists; this citation does not point to it. Register: https://www.courtlistener.com/opinion/4760997/bostock-v-clayton-county/
Found: 0 rows resolved to a case in the register.
Report id: r_1ff74051 (give it to render_report for a markdown report). Elapsed: 0.03 s.
Storage: Nothing you submit is stored. The document is processed in memory and discarded when this report is returned; only counts (citations, tiers, timing) are logged, never text.Development
npm install
npm run build # tsc -> dist/
npm test # vitest, mocked fetch, no network
LIVE=1 npm test -- test/live.test.ts # three live calls against proofread.law (counts against the free tier)
node scripts/smoke-stdio.mjs # spawn the stdio server, initialize, tools/list, four live tool calls
node scripts/smoke-stdio.mjs --offline # the same without networkLayout: src/client.ts is the typed HTTP client (/verify, /render, /api/coverage, /v1/resolve single and batch, /v1/briefs, /v1/suggest), src/format.ts the compact formatter for checks, src/resolve_format.ts the one for register answers, src/briefs_format.ts the one for saved briefs (/v1/briefs), src/suggest_format.ts the one for case suggestions (/v1/suggest), src/tools/<name>.ts one file per tool, src/server.ts registers them, src/cli.ts picks the transport. The remaining register routes (/v1/extract, /v1/case/{id}, /v1/coverage per reporter) slot in the same way: one method on the client, one file under src/tools/.
Publishing
See RELEASE.md: npm, the MCP Registry (server.json is in the repository), the Claude Desktop extension bundle (manifest.json, scripts/build-mcpb.sh, attached to each GitHub release) and Anthropic's connector directory.
License
MIT, Data Alchemy Labs.
Available Tools
14 toolsbilling_linkGet a checkout link for a paid planAIdempotent
Get a Stripe Checkout link that upgrades this API key's account to a paid plan (POST /agent/checkout-link). Use it when a check answers 'needs the ... plan' (402) or 'monthly allowance used' (429), or when the user asks to upgrade. Needs an API key (PROOFREAD_API_KEY, or one from sign_up). Plans: payg = pay as you go (no monthly fee; per check and per deep-checked citation), solo = monthly, firm = monthly with more keys; current prices at https://proofread.law/pricing. Give the link to the account owner to open in a browser; the plan is live within a minute of payment. No charge happens until the owner pays. An account that already has a subscription answers 409 (the owner changes plans in the billing portal at https://proofread.law/account).
| Name | Required | Description | Default |
|---|---|---|---|
| plan | No | payg (pay as you go), solo or firm. | payg |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, it discloses that no charge happens until the owner pays, that the plan activates within a minute, that an API key is required, and that an existing subscription yields 409. This adds meaningful behavioral context without contradicting the structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action, then efficiently covers usage triggers, prerequisites, plan details, and behavioral caveats. Every sentence contributes distinct, decision-relevant information; there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with no output schema, the description covers the essential context: when to call it, what it returns (a checkout link), authentication needs, plan semantics, payment timing, and the error case for existing subscriptions. An agent has what it needs to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single plan parameter is already enumerated in the schema, but the description adds real business meaning: payg is per-check with no monthly fee, solo and firm are monthly, and firm includes more keys. It also points to current pricing, going well beyond the schema's bare enum.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: "Get a Stripe Checkout link that upgrades this API key's account to a paid plan," and even includes the endpoint. It clearly distinguishes this billing tool from the sibling checking/reporting tools like check_citations and coverage.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit triggers: use it when a check returns 402 or 429, or when the user asks to upgrade. It also explains when not to use it: an account with an existing subscription returns 409 and should use the billing portal instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_citationsCheck legal citations in textARead-onlyIdempotent
Check every case citation in a text against proofread.law's register of about 10 million US court opinions (CourtListener bulk data). Use it on a draft brief, memo, letter or any prose that cites cases, before the citations are relied on. Returns the coverage statement, counts per tier, one line per row that needs a human (red = check this: the register holds something concrete that disagrees, such as a different case at that citation or quoted words not in the opinion; orange = cannot verify: nothing to check against, such as a Westlaw/Lexis identifier or a volume newer than the register), the number of citations found, and a report id for render_report. Cannot: resolve Westlaw (WL) or Lexis identifiers, check statutes, regulations or secondary sources, or say whether a case is still good law. A red row means 'check this', never 'this case does not exist'; an orange row means the register has nothing to check against, which is not evidence either way. deep=true also asks, for each found citation, whether the opinion supports the sentence it is cited for (white rows). It is slower (1 to 2 s per citation), opt-in because the clause before each citation is sent to a model judge, limited to 3 per month on the free tier, and its answers are a review queue, not a verdict. Free tier: 20 checks a month. Without an API key the free tier applies per IP address; a key from sign_up (free tier) identifies the account, and a paid-plan key lifts the limits.
| Name | Required | Description | Default |
|---|---|---|---|
| deep | No | Also check whether each cited opinion supports the sentence it is cited for. Slower, opt-in, 3 per month on the free tier. | |
| text | Yes | The text to check, as written (paragraphs, footnotes, a whole brief). Pasted text is fine. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnly, idempotent), the description discloses the return format (coverage statement, counts per tier, rows per tier, report id), the semantics of red/orange rows, the fact that deep mode sends clauses to a model judge, and rate limits (free tier 20/month, deep 3/month). This goes well beyond the annotations and gives the agent a clear behavioral model.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but front-loaded: purpose in the first sentence, usage in the second, then returns, limitations, result interpretation, deep mode, and limits. Each sentence adds distinct information with no redundancy, though it is long. It could be trimmed slightly, but the detail is mostly necessary for correct use.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully specifies what the tool returns (coverage statement, counts per tier, one line per row needing human, number of citations, report id) and how to interpret red/orange rows. It also covers limitations, deep mode, and rate limits, making it complete for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so parameters are already documented. The description adds context for deep=true: it generates white rows, takes 1-2 s per citation, sends clauses to a model judge, and produces a review queue, not a verdict. For text, it repeats the schema's 'paragraphs, footnotes, a whole brief' but adds context about draft briefs and memos; the added value is moderate but meaningful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Check every case citation in a text against proofread.law's register of about 10 million US court opinions (CourtListener bulk data).' It also lists what the tool cannot do ('Cannot: resolve Westlaw (WL) or Lexis identifiers, check statutes, regulations or secondary sources...'), which clearly distinguishes it from sibling resolve_citation(s) and check_document tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to use it: 'Use it on a draft brief, memo, letter or any prose that cites cases, before the citations are relied on.' It also states exclusions via the 'Cannot' list, and describes deep=true as an optional mode, so an agent knows when not to enable it. The interpretation guidance for red/orange rows further clarifies usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_documentCheck legal citations in a fileARead-onlyIdempotent
Check every case citation in a document on disk (PDF, DOCX, TXT or Markdown, up to 10 MB) against proofread.law's register of about 10 million US court opinions. The file is read here and uploaded to proofread.law, which extracts the text in memory, checks it and discards it. Returns the same compact result as check_citations: coverage statement, counts per tier, one line per red (check this) or orange (cannot verify) row, the number of citations found, and a report id for render_report. Scanned PDFs without a text layer, encrypted PDFs and legacy .doc files cannot be read; .docx needs a paid plan. Cannot: resolve Westlaw (WL) or Lexis identifiers, check statutes, regulations or secondary sources, or say whether a case is still good law. A red row means 'check this', never 'this case does not exist'; an orange row means the register has nothing to check against, which is not evidence either way. deep=true also asks, for each found citation, whether the opinion supports the sentence it is cited for (white rows). It is slower (1 to 2 s per citation), opt-in because the clause before each citation is sent to a model judge, limited to 3 per month on the free tier, and its answers are a review queue, not a verdict. Without an API key the free tier applies per IP address; a key from sign_up (free tier) identifies the account, and a paid-plan key lifts the limits.
| Name | Required | Description | Default |
|---|---|---|---|
| deep | No | Also check whether each cited opinion supports the sentence it is cited for. Slower, opt-in, 3 per month on the free tier. | |
| path | Yes | Absolute path to a .pdf, .docx, .txt or .md file on this machine. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool readOnly and idempotent, and the description adds valuable behavior beyond that: the file is uploaded and discarded, red rows mean 'check this' not 'case does not exist', orange rows mean 'cannot verify', and the deep=true mode sends the clause to a model judge but returns a review queue, not a verdict. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average but densely packed with essential constraints and disclaimers. It front-loads the core action and then efficiently layers limitations, result semantics, deep mode, and API-key behavior. There is minor duplication with the schema's deep parameter description, which keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description specifies the return shape (same compact result as check_citations: coverage statement, tier counts, one line per red/orange row, citation count, report id). It also covers file constraints, semantic meaning of each result color, deep-mode caveats, and auth/rate-limit behavior. Nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already covers both parameters with high quality and 100% coveragechers, so the baseline is 3. The description adds meaningful extra semantics: path must point to PDF/DOCX/TXT/MD up to 10 MB, excludes scanned/encrypted/.doc files, explains .docx plan requirement, and elaborates what deep=true does beyond the schema's 'slower, opt-in, 3 per month'. This is more than compensation, so 4 is justified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the verb ('Check'), the resource ('every case citation in a document on disk'), and the matching authority (proofread.law's register of ~10 million opinions). It clearly distinguishes itself from the sibling check_citations by being file-based, while still referencing check_citations for result-shape continuity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear operational context: what file types are supported, size limit, paid-plan requirement for .docx, and explicit exclusions ('Cannot: resolve Westlaw or Lexis identifiers, check statutes, regulations or secondary sources, or say whether a case is still good law'). It lacks an explicit 'use sibling X instead when...' recommendation, so it misses the top score by a small margin.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
coverageWhat proofread.law checks againstARead-onlyIdempotent
The coverage statement (which opinions the register holds, its date, its known gaps, what is not checked: Westlaw/Lexis identifiers, statutes, regulations, secondary sources) and the storage notice (a checked text is not stored; only a brief the user saves with save_brief is kept, encrypted in their account). Pass jurisdiction 'ch' for the Swiss register instead (BGE/ATF/DTF, Federal Supreme Court dockets, the federal courts and the 26 cantons), which also lists the courts held and the share of the live index each covers. Call it when a user asks what the check covers, how current it is, or what happens to their text. Free, not counted as a check.
| Name | Required | Description | Default |
|---|---|---|---|
| jurisdiction | No | Which register to describe. Default 'us'. 'ch' for Swiss case law. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and destructiveHint=false. The description adds significant context beyond these: it details what is not checked (Westlaw/Lexis identifiers, statutes, regulations, secondary sources), clarifies the storage notice (text not stored, only saved briefs encrypted), and notes the tool is free and not counted as a check. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured. It front-loads the core purpose (coverage statement and storage notice), then elaborates on the jurisdiction parameter, then gives usage triggers. While it is a bit long, every sentence contributes necessary information—no fluff. The structure aids readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one optional parameter and no output schema, the description is fully self-sufficient. It covers what the tool does, when to call it, parameter semantics, behavioral details (storage, free usage), and jurisdiction variants. An agent has everything needed to invoke it correctly without additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already describes the 'jurisdiction' parameter with a default and options. The description adds valuable context: it explains what the Swiss register includes (BGE/ATF/DTF, Federal Supreme Court dockets, federal courts, 26 cantons) and that it lists courts held and share of live index. This goes beyond the schema's simple enum description, enriching the parameter's meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear purpose: it returns the coverage statement and storage notice. It explicitly lists what the coverage includes (opinions, date, gaps, exclusions) and mentions the jurisdiction option, making it distinct from any sibling tool that deals with checking or briefs. No ambiguity about what it does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit trigger conditions: 'Call it when a user asks what the check covers, how current it is, or what happens to their text.' It also explains the jurisdiction parameter for Swiss case law. It does not explicitly say when not to use it, but the conditions are clear enough for an agent to decide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_briefPermanently delete a saved briefADestructiveIdempotent
Permanently delete a brief and all its saved versions from the user's proofread.law account. The deletion is permanent: it cannot be undone and the text cannot be recovered afterwards. Only call it when the user asks to delete that brief; if the id is in any doubt, confirm it with list_briefs first. Deleting runs no check. Saving is opt-in: nothing is saved unless save_brief or update_brief is called (check_citations and check_document never save anything). A saved brief is stored encrypted in the user's proofread.law account until delete_brief removes it, and needs an API key (PROOFREAD_API_KEY, or one from sign_up).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The id of the brief to delete, from save_brief or list_briefs. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, but the description adds crucial context: deletion is permanent and unrecoverable, no check is run before deletion, and saved briefs are stored encrypted until removed. It also explains the API key requirement. This goes well beyond the annotations and fully discloses the destructive behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the most important fact (permanent deletion). Every sentence adds meaningful context: permanence, when to call, no-check behavior, opt-in saving, encryption, and API key. Slightly long but each clause earns its place; no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive single-parameter tool with no output schema, the description covers everything an agent needs: what gets deleted, permanence, confirmation step, no-check behavior, storage/encryption context, and authentication requirement. The sibling list and schema complete the picture.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single parameter 'id' is already described in the schema. The description adds value by explaining the id comes from save_brief or list_briefs and that it should be confirmed if in doubt. Since the schema already covers the parameter well, a 4 is appropriate rather than 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Permanently delete'), a specific resource ('a brief and all its saved versions'), and the account context ('user's proofread.law account'). It clearly distinguishes from siblings like save_brief, update_brief, and list_briefs by emphasizing permanence and deletion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Only call it when the user asks to delete that brief' and instructs to confirm with list_briefs if the id is in doubt. It also clarifies that saving is opt-in and that check_citations/check_document never save, which prevents misuse. This is strong when-to-use guidance with an alternative named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_briefGet a saved brief and its latest reportARead-onlyIdempotent
Get one brief saved in the user's proofread.law account: its saved text, the latest report (coverage statement, counts per tier, the rows that need attention, a report id for render_report) and the versions kept. Use it to pick up where the user left off, or to get the exact saved text before editing it for update_brief. include_text=false returns the report without the text (a long brief is long). Pass version to read an earlier version's text and counts instead of the latest. It only reads: this tool saves nothing and runs no check. Saving is opt-in: nothing is saved unless save_brief or update_brief is called (check_citations and check_document never save anything). A saved brief is stored encrypted in the user's proofread.law account until delete_brief removes it, and needs an API key (PROOFREAD_API_KEY, or one from sign_up).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The brief id from save_brief or list_briefs. | |
| version | No | An earlier version number from the versions list; omit for the latest. | |
| include_text | No | Return the saved text (default true). false returns only the report. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint=true, idempotentHint=true, destructiveHint=false), the description adds specific behavioral context: it only reads, saves nothing, runs no check, and explicitly names the tools that do save (save_brief, update_brief). It further discloses that saved briefs are stored encrypted until delete_brief removes them and that an API key is needed. This goes well beyond the annotation booleans.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but information-dense; every sentence delivers actionable detail (what the report contains, the read-only guarantee, opt-in saving, encryption, API key). It front-loads the core purpose in the first sentence and then layers usage, constraints, and caveats logically. A slight trim could tighten it, but the verbosity is justified given the tool's scope.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with 3 parameters and no output schema, this description covers all necessary facets: return contents, parameter usage, side effects (none), storage/encryption, authentication, and relationships to sibling tools. Nothing an agent needs to invoke it correctly is missing, and the annotations already handle safety hints, so the description is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents all parameters thoroughly (coverage 100%), so the baseline is 3. The description adds practical semantics: id is sourced from save_brief or list_briefs, version refers to an earlier entry in the versions list, and include_text=false is recommended for long briefs. This enriches the schema's descriptions with real-world usage context, pushing it to a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Get') and resource ('one brief saved in the user's proofread.law account'), then enumerates exactly what is returned (saved text, report details, versions). It explicitly differentiates from siblings like save_brief, update_brief, delete_brief, and list_briefs by describing its read-only retrieval role and mentioning render_report and update_brief as downstream consumers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It tells the agent when to use it ('to pick up where the user left off' or 'to get the exact saved text before editing it for update_brief'), and when not to (it saves nothing, runs no check, and saving is opt-in only via save_brief or update_brief). It also clarifies the use of include_text and version parameters with practical guidance ('a long brief is long'), and mentions the required API key.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_briefsList the saved briefsARead-onlyIdempotent
List the briefs saved in the user's proofread.law account: id, title, when each was last saved and checked, the number of citations, the counts per tier from its latest check, and the number of versions kept. Use it to find a brief's id when the user refers to a draft by name, before get_brief, update_brief or delete_brief. It only reads: this tool saves nothing and runs no check. Saving is opt-in: nothing is saved unless save_brief or update_brief is called (check_citations and check_document never save anything). A saved brief is stored encrypted in the user's proofread.law account until delete_brief removes it, and needs an API key (PROOFREAD_API_KEY, or one from sign_up).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral context by explicitly stating that the tool 'only reads: this tool saves nothing and runs no check,' reinforcing the read-only nature and clarifying that it does not trigger any analysis. It also discloses that saving is opt-in and never happens implicitly, which is crucial for agent decision-making. This goes beyond the annotations by clarifying the side-effect-free guarantee in operational terms. Minor gap: it doesn't mention pagination or sorting, but for a zero-parameter tool, the return is likely a complete list, so this is acceptable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is moderately sized, covering multiple aspects: what is returned, when to use it, and the non-saving guarantee. It is appropriately front-loaded with the primary purpose and returned fields, and only later explains the saving semantics, which is secondary. The sentences are information-dense but not excessive; they earn their place by adding unique behavioral and usage context. A slight deduction because the save-semantics explanation could be condensed, but it is relevant for avoiding misuse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (zero parameters, no output schema, no nested objects), the description is complete. It tells the agent what the tool does, what it returns, when to use it, and what side effects (none) it has. The annotations cover safety, and the description provides all the operational details needed for an agent to call it correctly. The only missing piece might be an explicit statement about the return format (e.g., 'array of objects'), but the description's enumeration of fields implies the structure. For this complexity level, nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema is effectively empty, so there is nothing to document. The baseline for zero-parameter tools is 4, as the description doesn't need to add parameter details. The description appropriately focuses on the return value and usage context, which is what the agent needs to know. It does not repeat any schema information because there is none. This is a well-handled case where the description compensates for the lack of parameter structure by being clear about outputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: listing saved briefs and their metadata. It specifies the exact resource ('briefs saved in the user's proofread.law account') and enumerates the returned fields (id, title, timestamps, citation counts, tier counts, versions). It distinguishes itself from siblings by explicitly stating 'Use it to find a brief's id... before get_brief, update_brief or delete_brief,' which differentiates it from get_brief (which retrieves a single brief) and other operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use guidance: 'Use it to find a brief's id when the user refers to a draft by name, before get_brief, update_brief or delete_brief.' It also clarifies when not to use it relative to saving operations: 'Saving is opt-in: nothing is saved unless save_brief or update_brief is called (check_citations and check_document never save anything).' This helps the agent decide between list_briefs and save/update tools, and provides clear context for the typical workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
render_reportRender a check as a markdown reportARead-onlyIdempotent
Turn a finished check into a markdown diligence report: header, coverage and storage notices, a summary table sorted check-this, cannot-verify, support, found, and a detail block per flagged row with the register evidence. Use it when the user wants a report to keep or attach to the file, or when the compact result was cut short. Pass the report_id returned by check_citations or check_document (ids live in this server's memory until it exits), or the full report JSON from proofread.law's /verify endpoint. Free, not counted.
| Name | Required | Description | Default |
|---|---|---|---|
| report | No | A full report JSON as returned by POST /verify, if you have one instead of an id. | |
| report_id | No | The 'Report id' from a previous check_citations or check_document result. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark it read-only, idempotent, and non-destructive. The description adds meaningful behavioral context: report IDs live in server memory only until exit, the tool is free and uncounted, and it expects a 'finished' check rather than an arbitrary input.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences each earn their place: purpose and output structure, usage condition and input sources, then a terse free/uncounted note. The content is front-loaded with the core verb and resource before details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by detailing what the report contains. It also covers both input alternatives, the memory constraint, the finished-check prerequisite, and cost behavior, leaving no obvious gap for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both parameters already have clear descriptions. The description adds useful cross-tool context by tying report_id to check_citations/check_document results and clarifying that either report_id or full report JSON can be supplied, though it does not add substantial new syntax-level detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb phrase, 'Turn a finished check into a markdown diligence report,' and enumerates the report's exact composition: header, coverage/storage notices, sorted summary table, and per-row detail blocks. This makes the tool's role as a rendering step distinct from sibling checking and resolution tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use the tool: when the user wants a report to keep or attach, or when the compact result was cut short. It does not explicitly state when not to use it or name alternatives, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolve_citationLook up one citation in the registerARead-onlyIdempotent
Look up a single case citation (for example '590 U.S. 644', or 'Bostock v. Clayton County, 590 U.S. 644 (2020)') in proofread.law's register and answer: is there a case at this citation, which one (name, court, date, parallel citations, link), and how complete the register is for that volume. This is a register lookup of the citation, not a comparison with the case name you have: if the case it returns is not the one you expected, the citation points elsewhere. Statuses: found; ambiguous (several entries, candidates listed); not in the register (a register fact with a coverage qualifier, never proof that the case does not exist); cannot verify (a Westlaw/Lexis identifier, or a volume the register cannot see yet); known citation (other opinions cite it, the opinion itself is not held); no citation recognised. Use it when one citation is in doubt; use check_citations for prose, and resolve_citations for a list. Cannot: resolve Westlaw (WL) or Lexis identifiers, check statutes, regulations or secondary sources, or say whether a case is still good law. A red row means 'check this', never 'this case does not exist'; an orange row means the register has nothing to check against, which is not evidence either way. Counts against the resolve quota (1,000 a month free), not the check quota. Without an API key the free tier applies per IP address; a key from sign_up (free tier) identifies the account, and a paid-plan key lifts the limits.
| Name | Required | Description | Default |
|---|---|---|---|
| citation | Yes | One citation string. A case name and year around it are fine; only the reporter citation is resolved. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, idempotentHint, and non-destructive. The description substantially adds behavioral context: enumerates statuses (found, ambiguous, not in register, cannot verify, known citation, no citation recognised), clarifies that a 'not in register' result is 'never proof that the case does not exist,' and explains red/orange row semantics. It also discloses quota impact and API-key behavior. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence earns its place. It front-loads the core purpose, then lists statuses, usage guidance, exclusions, and quota notes in a logical order. No redundancy or filler; each clause adds distinct information that an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries full responsibility for explaining return semantics, and it does so comprehensively: statuses, coverage qualifiers, red/orange row meanings, quota limits, and API-key behavior. It covers error cases and ambiguity handling. Nothing an agent needs to call this correctly or interpret the result is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (the citation parameter is described as 'One citation string'), so baseline is 3. The description adds meaningful nuance: 'A case name and year around it are fine; only the reporter citation is resolved,' which clarifies acceptable input formats and what will be parsed. This goes beyond the schema and helps agents format calls correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('look up'), a specific resource ('a single case citation in proofread.law's register'), and precisely defines the output: whether a case exists, its details, and register completeness. It explicitly distinguishes from siblings by naming check_citations (for prose) and resolve_citations (for a list), and lists what it cannot do (Westlaw/Lexis IDs, statutes, good-law status). An agent can unambiguously identify when to use this tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use: 'Use it when one citation is in doubt' and directly names alternatives: 'use check_citations for prose, and resolve_citations for a list.' It also lists exclusions (cannot resolve WL/Lexis, check statutes, etc.), giving clear routing guidance. No inference is required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolve_citationsLook up a list of citations in the registerARead-onlyIdempotent
Look up up to 500 case citation strings in proofread.law's register in one call and get one line per citation, in input order: found (the case, court, date, link), ambiguous, not in the register (a register fact with a coverage qualifier, never proof that the case does not exist), cannot verify (Westlaw/Lexis identifier, or a volume the register cannot see yet), known citation, or no citation recognised. Use it for a table of authorities or any list of citations you already have; use check_citations for prose (it also checks names and quotations). Cannot: resolve Westlaw (WL) or Lexis identifiers, check statutes, regulations or secondary sources, or say whether a case is still good law. A red row means 'check this', never 'this case does not exist'; an orange row means the register has nothing to check against, which is not evidence either way. Each citation counts against the resolve quota (1,000 a month free), not the check quota. Without an API key the free tier applies per IP address; a key from sign_up (free tier) identifies the account, and a paid-plan key lifts the limits.
| Name | Required | Description | Default |
|---|---|---|---|
| cites | Yes | Citation strings, one per entry, up to 500. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool as read-only, idempotent, and open-world, and the description goes further: it explains per-line outcome categories, red/orange row meanings, the open-world caveat, quota counting, and authentication behavior. There is no contradiction with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded and every subsequent clause carries operational meaning: outcomes, caveats, usage split, quota, and auth. Despite its length, the description packs high-density guidance with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter lookup tool with a rich but unspecified output, the description covers the return shape, per-row semantics, error interpretation, limits, quota, and authentication. There is no output schema, so this level of detail is necessary and sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already fully documents cites as up to 500 strings, so the baseline is 3. The description adds that these are case citation strings, clarifies that Westlaw or Lexis identifiers cannot be resolved, and ties each entry to a returned line in input order, giving the agent more usable constraints than the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action ('Look up up to 500 case citation strings') and names the resource (proofread.law's register). It also distinguishes list lookup from check_citations, so an agent can select the right tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states exactly when to use the tool ('for a table of authorities or any list of citations you already have') and when to use check_citations instead. It also lists unsupported inputs and quota/API-key conditions, removing ambiguity about prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_briefSave a brief to the account and check itA
Save a brief (or any draft that cites cases) to the user's proofread.law account and check its citations in the same call. Use it when the user asks to keep a draft and come back to it, or to start the edit-and-recheck loop: save once, then after each round of edits call update_brief with the id. Returns the brief id, the coverage statement, counts per tier, one line per row that needs attention (red = check this: the register holds something concrete that disagrees; orange = cannot verify: nothing to check against, not evidence either way), and a report id for render_report. Only call it when the user wants the draft saved; to check without saving, use check_citations. Each call saves a new brief and counts as one check. Cannot: resolve Westlaw (WL) or Lexis identifiers, check statutes, regulations or secondary sources, or say whether a case is still good law. A red row means 'check this', never 'this case does not exist'; an orange row means the register has nothing to check against, which is not evidence either way. Saving is opt-in: nothing is saved unless save_brief or update_brief is called (check_citations and check_document never save anything). A saved brief is stored encrypted in the user's proofread.law account until delete_brief removes it, and needs an API key (PROOFREAD_API_KEY, or one from sign_up).
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | The full text of the brief, as written (paragraphs, footnotes, the whole draft). | |
| title | No | A name the user will recognise in list_briefs, e.g. 'Motion to dismiss, Smith v. Jones'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond annotations by explaining non-idempotency (each call saves a new brief), opt-in saving, encryption, API key requirement, and the precise meaning of red/orange rows. It also lists limitations (cannot resolve WL/Lexis, statutes, etc.) and clarifies red means 'check this', not 'case does not exist'. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but each sentence carries distinct value: purpose, usage, return format, limitations, semantics, opt-in, storage, and auth. It front-loads the core purpose and flows logically. While it could be tightened, the density is justified by the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description thoroughly explains the return values (brief id, coverage statement, tier counts, attention rows with red/orange meaning, report id) and includes auth, storage, opt-in, and limitations. Nothing critical is missing for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already fully documents both parameters (100% coverage) with clear descriptions. The tool description adds no additional parameter-level semantics; it focuses on behavior and usage. Baseline 3 is appropriate given the schema's completeness.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb (save), resource (brief to account), and an additional action (check citations). It distinguishes from siblings by explicitly naming check_citations as the alternative for checking without saving, leaving no ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit when-to-use instructions: when the user wants to keep a draft or start the edit-and-recheck loop, and when-not: only when saving is desired, with the alternative tool named. It also maps the workflow with update_brief for subsequent edits.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sign_upCreate a proofread.law account and API keyA
Open a proofread.law account for its owner and get an API key, in one call (POST /agent/signup). Use it when the user wants their own quota instead of the anonymous per-IP free tier, or before billing_link. The email must be the account OWNER's real inbox (placeholder domains are rejected); the owner receives one confirmation email and nothing else. The key is shown once: this server adopts it for the rest of this session, and the user should put it in PROOFREAD_API_KEY in their MCP configuration so it survives a restart. Never send the key anywhere but proofread.law. Calling again with the same address while the account is unconfirmed and unpaid rotates the key; a confirmed or paying account answers 409 (the owner manages keys at https://proofread.law/account). Free tier per account: 20 checks, 3 deep checks, 1,000 resolves a month; 5 sign-ups an hour per client.
| Name | Required | Description | Default |
|---|---|---|---|
| Yes | The account owner's real email address. Ask the user for it; do not invent one. | ||
| agent_name | Yes | A name for this agent or client (1 to 64 characters: letters, digits, spaces, dots, hyphens, underscores), e.g. 'Claude Desktop'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations only indicate readOnlyHint=false, idempotent=false, destructive=false, but the description discloses key rotation on repeated signup, the 409 on confirmed/paying accounts, session adoption of the key, the need to persist it in PROOFREAD_API_KEY, rate limits, and quota limits. This goes well beyond annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence carries necessary operational detail—purpose, usage trigger, email requirement, key handling, repeat-call behavior, quotas, and rate limit. Dense but not bloated, with key actions front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, it explains the critical outcome (key shown once), error condition (409), post-call persistence steps, and account limits. An agent has everything needed to invoke it correctly and handle the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds real constraints: email must be the account owner's real inbox and placeholder domains are rejected, which is more than the schema's phrasing. Agent_name gets no additional meaning, but overall parameter semantics are enhanced.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Open a proofread.law account for its owner and get an API key, in one call'. It also names the endpoint and explicitly differentiates from siblings by tying its use to quota ownership and billing_link sequencing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use: 'when the user wants their own quota instead of the anonymous per-IP free tier, or before billing_link'. It also implies when not to use by describing 409 for confirmed/paying accounts and warns never to send the key elsewhere, giving clear routing and safety guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
suggest_casesSuggest leading Swiss cases for a statute articleARead-onlyIdempotent
Swiss law only. Give a Swiss federal statute article ('Art. 41 OR', 'art. 41 CO', 'Art. 8 ZGB', 'art. 9 Cst.') or a paragraph that cites one, and get the leading Federal Supreme Court cases (BGE/ATF/DTF) the court cites with that article, as cases to read. The query needs a statute article: there is no free-text search, and a query without one answers that no article was recognised. Federal acts only; cantonal law is not covered. The answer is one ranked list, each case labelled with its field of law; the domain filter keeps the same order within one field. Each row has the citation, date, field, rank, the quoted passage (regeste or consideration), the decision's link, a link to check the citation on proofread.law, and any later change of practice (a changed precedent is listed with its flag, never left out; a row without a flag is not evidence that its practice still holds). A suggestion has not been checked against the user's sentence; to check a citation, use check_citations. The answer comes back in the query's language (German, French or Italian; lang='en' for English). During the trial phase it needs a paid-plan or trial API key; each answered query counts as one resolve.
| Name | Required | Description | Default |
|---|---|---|---|
| k | No | Number of rows, 1 to 50. Default 10. | |
| lang | No | The answer's language: de, fr, it or en. Default: the query's language. | |
| query | Yes | A Swiss federal statute article, e.g. 'Art. 41 OR' or 'art. 9 Cst.', or a paragraph that cites one (up to 20,000 characters). | |
| domain | No | Only this field of law (civil, criminal, public or social), in the same order as the full list. Default all. | all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and idempotentHint, so the description correctly doesn't repeat those. It adds rich behavioral detail: the ranked-list output structure, per-row fields (citation, date, field, rank, quoted passage, links), the flag semantics for changed precedents ('a row without a flag is not evidence that its practice still holds'), the language selection rule, and the billing note. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average, but every sentence carries a distinct piece of information: scope, input requirement, output structure, caveats, language, and billing. It is front-loaded with the core function and then flows logically through constraints and output details. The length is justified by the tool's complexity, so 4 is appropriate rather than 5 for slight over-brevity in some clauses.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with four parameters and no output schema, the description is exceptionally complete. It covers the input format, the output fields, the ordering behavior, the language default, the trial-phase key requirement, and the relationship to check_citations. An agent has everything needed to call it correctly and interpret the response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds value by clarifying that the query must be a statute article (or a paragraph citing one) and that free-text search is not supported, which is not in the schema. It also explains the domain filter keeps order within a field, which is not in the schema. These additions justify a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb-resource pair: 'get the leading Federal Supreme Court cases (BGE/ATF/DTF) the court cites with that article'. It explicitly differentiates from the sibling check_citations by stating that suggestions are not checked against the user's sentence and pointing to check_citations for that. No ambiguity remains about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states when to use it (you have a Swiss federal statute article) and when not (cantonal law, free-text queries). It explicitly says 'a query without one answers that no article was recognised' and names the alternative tool check_citations for verification. It also notes the trial-phase API-key requirement, which is a practical usage constraint.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_briefSave an edited brief and re-check itA
Save an edited version of a brief saved in the user's proofread.law account and re-check it. This is the loop for fixing flagged citations: edit the text (the citation, case name or quotation a row points at, or take the citation out), call update_brief with the whole edited text, and read the changes: which flags were resolved (flagged before, not flagged now), which flags are new, how many rows are unchanged, and then every row that still needs attention. Repeat until every remaining row has been reviewed by the user. Send the full text, not a diff or an excerpt: it becomes the latest version, and the previous version is kept (get_brief lists the versions). recheck=true re-runs the check on the saved text without editing it, e.g. after the register was updated. A changed text or a recheck counts as one check and keeps a new version; a title alone renames the brief without a check and is not counted. Cannot: resolve Westlaw (WL) or Lexis identifiers, check statutes, regulations or secondary sources, or say whether a case is still good law. A red row means 'check this', never 'this case does not exist'; an orange row means the register has nothing to check against, which is not evidence either way. Saving is opt-in: nothing is saved unless save_brief or update_brief is called (check_citations and check_document never save anything). A saved brief is stored encrypted in the user's proofread.law account until delete_brief removes it, and needs an API key (PROOFREAD_API_KEY, or one from sign_up).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The brief id from save_brief or list_briefs. | |
| text | No | The whole edited text of the brief. It replaces the latest version; the previous one is kept. | |
| title | No | A new name for the brief. | |
| recheck | No | Re-run the check on the saved text without editing it, e.g. after the register was updated. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only state readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false. The description adds substantial behavior not in annotations: it creates a new version and keeps the previous one, saving is opt-in and encrypted, check counts, red/orange row semantics, and the caveat that 'red row means check this, never this case does not exist'. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Description is long but every sentence earns its place. It front-loads the core purpose and loop, then covers edge cases (recheck, title-only, versions, limitations, red/orange semantics, opt-in saving, encryption, API key). It is densely packed with actionable information and avoids redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (4 params, no output schema, many behavioral nuances), the description is remarkably complete. It explains the full workflow, versioning behavior, check counting, limitations, row semantics, saving policy, and security. An agent has everything needed to call it correctly and interpret results despite no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds value by clarifying that 'text' must be the whole edited text (not a diff or excerpt), that it replaces the latest version, and that recheck=true re-runs without editing. It also emphasizes sending full text, which the schema doesn't explicitly state. This goes beyond the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Save an edited version') and resource ('brief'), and immediately positions it as the loop for fixing flagged citations, distinguishing it from siblings like check_citations and save_brief. It clearly tells the agent what it is for and how it differs from related tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance (the loop for fixing flagged citations, repeat until user has reviewed every row), when not to use it (Cannot resolve Westlaw/Lexis identifiers, check statutes/regulations/secondary sources), and when to use recheck. It also distinguishes from save_brief (opt-in saving) and check_citations (never saves).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.3.0- Added
suggest_cases
5 tool updates
v0.2.0- Added
delete_brief - Added
get_brief - Added
list_briefs - Added
save_brief - Added
update_brief
4 tool updates
v0.1.4- Added
billing_link - Changed
check_document1 field changed- changed
Input schema / properties / path / descriptionPrevious value: -"Absolute path to a .pdf, .docx or .txt file on this machine."New value: +"Absolute path to a .pdf, .docx, .txt or .md file on this machine."
- Changed
coverage1 field changed- added
Input schema / properties / jurisdictionAdded value: +{ + "description": "Which register to describe. Default 'us'. 'ch' for Swiss case law.", + "enum": [ + "us", + "ch" + ], + "type": "string" +}
- Added
sign_up
6 tool updates
v0.1.0- First observed
check_citations - First observed
check_document - First observed
coverage - First observed
render_report - First observed
resolve_citation - First observed
resolve_citations
TDQS
Scored across 14 tools
Tools are mostly distinct: check_citations vs check_document differ by input type (text vs file), resolve_citation vs resolve_citations differ by count, and suggest_cases offers a different domain. The cluster of citation tools could be confused, but descriptions clearly separate prose checking, list resolution, single lookup, and Swiss suggestions. Brief CRUD and account tools are unambiguous.
Most tools follow a verb_noun pattern (check_citations, resolve_citations, list_briefs, save_brief, render_report), with a few exceptions like 'coverage' (noun) and 'sign_up' (phrasal verb). The pattern is predictable and readable, with only minor deviations.
14 tools is well within the ideal range for this domain. Each tool has a clear purpose: checking, resolving, saving, managing, reporting, and account setup. No redundancy or bloat; the count matches the feature set.
The tool set covers the full lifecycle: checking citations in text and files, resolving single and batch citations, suggesting Swiss cases, managing saved briefs (CRUD), generating reports, and handling account quotas. Minor gaps exist, like no direct tool to check statutes or regulations, but those are explicitly out of scope, so the surface is appropriately complete.
Maintenance
Related MCP Connectors
Verify legal citations, case treatment, quotes and whole briefs against 10.7M U.S. opinions
Resolve, search and verify legal citations against the official sources, with provenance.
Check if a US legal citation exists, names its case and has a valid pin cite; or find a case. Free.
Search U.S. case law, fetch opinions, and ask matter-aware legal questions over your documents.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables US case law search, citation parsing, practice management via Clio, and federal court filings through PACER.-
- AlicenseNot gradedqualityDmaintenanceValidates legal citations against the CourtListener database to detect hallucinated citations in legal documents.5MIT
- AlicenseNot gradedqualityDmaintenanceEnables Claude Code to index and semantically search through PDFs, code, and documents with exact citations and zero hallucinations.MIT
- FlicenseAqualityDmaintenanceEnables Claude Desktop to search the CanLII Canadian legal database and retrieve the full text of matching legal documents.1-