canli-validation-mcp
Summary: This MCP server gives AI agents fifteen tools for validating whether a backtest's Sharpe ratio is real or just the luckiest of many variants tried, plus receipt verification, service status, and SEC company financial history.
Get a free key (
get_key) — issues and holds an in-memory API key for the session when neitherCANLI_KEYnor local mode is set.Deflated Sharpe (
validate_deflated_sharpe) — tests whether a Sharpe survives the number of variants tried; takes seven statistics or a raw return series.Overfitting (CSCV) (
validate_overfitting) — probability that picking the best of several backtested variants was overfitting, given every variant's returns matrix.Data-snooping tests (
validate_reality_check) — Hansen's SPA, White's Reality Check and Romano-Wolf StepM on a matrix of all variants' returns, optionally versus a benchmark.Paper evidence (
validate_paper_evidence) — checks a paper or simulated record against thecanli.paper-evidence.v0schema, with JSON pointers for each failure.Breadth ceiling (
validate_breadth) — highest book Sharpe reachable from sleeves of a given quality/correlation, sleeves needed for a target.Minimum track record length (
validate_track_record) — observations and years needed for a Sharpe to beat a benchmark.Minimum backtest length (
validate_backtest_length) — years before the best of N trials isn't expected to hit a target Sharpe by luck.Haircut Sharpe (
validate_haircut_sharpe) — Harvey-Liu multiple-testing haircut via Bonferroni, independent tests, Holm and BHY.Luck-equivalent trials (
validate_luck_trials) — how many skill-less strategies a search needed for its best Sharpe to arise by luck.One-call audit (
audit_backtest) — runs deflated Sharpe, track-record and (given variants) overfitting on one series, reading returns from files in local mode.Receipts (
get_receipt,verify_receipt) — fetch a stored verdict and verify its Ed25519 signature and content hashes offline.Service status (
service_status) — checks the API is up and quota constants, useful after a timeout.SEC company financial history (
company_financial_history) — lists a company's us-gaap concepts by CIK or ticker, or returns observations newest-first with filing accession, form, filed date, unit and source SHA-256.Every result carries limits — each answer includes verbatim boundary sentences saying what it does not establish, plus a signed, reproducible receipt.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@canli-validation-mcpvalidate my backtest's deflated Sharpe ratio"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
canli-validation-mcp
Is your best backtest real, or just the luckiest of the variants you tried? This MCP server lets Claude, Cursor or any MCP client answer that with the standard corrections: the deflated Sharpe ratio, the CSCV probability of backtest overfitting, data-snooping tests on every variant tried (Hansen's SPA, White's Reality Check, Romano-Wolf StepM), the minimum track record length, the haircut Sharpe ratio, and luck-equivalent trials. Free and MIT-licensed.
Quick start
# Claude Code, computed on your machine: no key, nothing sent anywhere
claude mcp add canli-local --env CANLI_LOCAL=1 -- npx -y canli-validation-mcp
# the same tools with signed, stored receipts from the free API (a free key is issued on first use)
claude mcp add canli -- npx -y canli-validation-mcp
# nothing to install: the hosted endpoint
claude mcp add --transport http canli https://canlicapital.com/mcpThen ask, for example: "I tried 229 variants and kept the best: an annualised Sharpe of
1.5 over 730 daily returns (365 a year), skew -0.5, kurtosis 5, and the variants'
Sharpe ratios spread by 0.57. Is it real?" The assistant calls validate_deflated_sharpe: the
best of 229 skill-less variants would reach 1.60 by luck alone, so the probability that the
Sharpe is above zero falls from 98.1% to 44.4% once the search is counted. Claude Desktop and any
other stdio client (Cursor, VS Code) run the same npx command; see "Claude Desktop" and
"Generic stdio client" below.
Related MCP server: Equibles Agent Terminal
What it does
An MCP (Model Context Protocol) server over canlicapital.com's free, keyed validation API. It gives a coding agent fifteen tools: issue a free key, run the nine validators (deflated Sharpe, CSCV overfitting, data-snooping tests (Hansen's SPA, White's Reality Check and Romano-Wolf StepM), paper-evidence conformance, breadth ceiling, minimum track record length, minimum backtest length, haircut Sharpe ratio, luck-equivalent trials), audit one backtest with three of them in a single call, fetch a stored receipt, verify a receipt's signature offline, read service status, and read a company's reported financial history from SEC filings. Every validation result carries, beside the number, the sentences that say what it does not establish and the receipt that records it, success or error, so the agent cannot see a number without its limits.
This package is published to npm as canli-validation-mcp.
Also listed on the official MCP Registry (io.github.arhancanli/canli-validation-mcp) and
cursor.directory. Run it with npx, no install step, as shown below. A local checkout is only needed to develop or
test this package itself; see "Local checkout" near the bottom.
This README describes the version in package.json. Unversioned npx runs npm's latest
release; npx -y canli-validation-mcp@<version> pins one.
How well agents use it
A fixed benchmark gives models these tools and scores whether they pick the right one and return
the right answer, against ground truth computed from the same checked code. Results for three
models, with every run recorded, are in
bench/agent/README.md.
What the API is (and is not)
The engine is the product. The service runs your submitted numbers through the same honesty
arithmetic canlicapital.com's own paper record runs on itself and hands back a verdict anyone can
recompute from the receipt. It signs every receipt, does not accept market data, does not
grade a strategy, and never saw your data source, its costs, or any lookahead in how a series was
built. See docs/superpowers/specs/2026-09-05-developer-key-validation-api-design.md in the main
repository for the full design.
Tools
Tool | Calls | Key required |
|
| no |
|
| yes |
|
| yes |
|
| yes |
|
| yes |
|
| yes |
|
| yes |
|
| yes |
|
| yes |
|
| yes |
| the deflated Sharpe, track record and, with | yes |
|
| no |
|
| no |
|
| no |
|
| no |
validate_deflated_sharpe accepts exactly one of two input shapes, never a mix of both:
the seven contract fields:
observed_sharpe_annualized,observations,periods_per_year,skew,non_excess_kurtosis,effective_independent_trials,cross_trial_sharpe_sd_annualized;or a return series plus the trials behind it:
returns,periods_per_year,effective_independent_trials,cross_trial_sharpe_sd_annualized.
Sending fields from both shapes, or from neither, is rejected before any request leaves the
process; see src/schemas.mjs.
validate_reality_check takes the returns of every variant the search tried, one row per period
and one column per variant, as matrix or as matrix_file (a CSV or JSON on your machine), and an
optional benchmark series; with none, variants are tested against zero. It runs three tests on one
seeded stationary bootstrap:
Hansen's SPA: the chance that the best variant's studentized excess return is this good if no variant has an edge, with lower and upper bounds (
spa.p_value, the consistent p-value, is the headline, andmonte_carlo_seits sampling error);White's Reality Check: the same question without studentizing, so one volatile variant can dominate it;
Romano and Wolf's StepM: which variants beat the benchmark, with the familywise error held at
alpha.
A result reproduces exactly from its seed. On three fixed-seed cases the p-values agree with
Python's arch 8.0 and with a numpy transcription of Hansen's formulas within Monte Carlo error, and
the StepM sets match arch's (js/snooping-core.test.js). Only the variants sent are counted: a
search that tried more than it sends makes luck look smaller than it was.
company_financial_history is different from the other tools: it reads the public company
reference at canlicapital.com, not the validation API. Give it a CIK (1 to 10 digits) to list a
company's available financial histories, or a CIK and a us-gaap concept such as Revenues to
get the observations, newest first (limit defaults to 40, maximum 200). Every observation
keeps its filing accession, form, filed date and unit, and the result carries the SHA-256 of the
original SEC response and the record's own boundary sentence: these are accounting values as
reported to the SEC, not market prices, returns or a recommendation. Companies and concepts
outside the current release return an error with the available concepts listed.
Auditing a backtest in one call
audit_backtest takes one strategy's return series, the number of variants tried and their Sharpe
dispersion, and optionally every variant's returns. It runs validate_deflated_sharpe on the
series, then validate_track_record on the Sharpe, skew and kurtosis that check derived, then,
with variants, validate_overfitting. Each check is exactly what its own tool returns, with its
own receipt; the boundary sentences they share are stated once. A check that refuses (a Sharpe
that cannot beat the benchmark has no minimum track record) is reported as that check's error.
The audit adds no grade of its own. Through the API it uses one validation per check.
When the server runs on your machine, returns_file and variants_file take the path of the
backtest's output instead of the numbers: a CSV (comma, semicolon or tab separated, with or without
a header; date and label columns are ignored, an unnamed or counting index column is skipped and
reported) or a JSON array. returns_column picks the column when there are several. Only the
numbers are read; the hosted endpoint refuses file paths. An agent copying a long series into a
call can drop values, and the copy costs tokens; in our agent benchmark, reading the file instead
took three audit questions from 4 of 9 to 8 of 9 answered correctly on gpt-5.4-mini.
Prompts, resources and structured results
Clients that show MCP prompts offer two guided workflows: validate_backtest (deflated Sharpe, then
overfitting, then the track record needed, reported with what each number does not establish) and
track_record_needed. Two resources can be read: canli://limits, the boundary sentences every
result carries, and canli://sources, the papers behind each validator and how each is checked
against them. Every tool result carries its envelope both as text and as structuredContent.
Compact context (0.3.0)
An agent pays for every token a tool returns, including whitespace it never reads. Since 0.3.0
every result is minified JSON, and a company history returns its observations as one columns
header and one row per observation, with a unit shared by every row stated once. No field is
dropped: every boundary sentence, accession number, form, filed date and source hash is still in
the result, and columns + rows rebuild each observation exactly.
Measured with bench/token_cost.py on live records (tokenizer: tiktoken o200k_base; other
tokenizers give different absolute counts), 20 observations each:
Record | 0.2.0 (indented) | minified | 0.3.0 (minified, columnar) |
Apple, StockholdersEquity | 2,214 | 1,560 (−29.5%) | 1,060 (−52.1%) |
Microsoft, CashAndCashEquivalentsAtCarryingValue | 2,231 | 1,574 (−29.4%) | 1,065 (−52.3%) |
The validation tools gain the minification only: the live service_status envelope measured 593
tokens indented and 475 minified (−19.9%); their boundary sentences are kept word for word. These
are measurements of these results, not a claim about any other server.
Breaking change from 0.2.0: history.observations is now {unit?, columns, rows} instead of
an array of objects.
Configuration
Variable | Default | Meaning |
|
| Where the API lives. Point it at a preview deployment for testing. |
| unset | A key already issued from |
| unset |
|
| all | Which tools to list: a comma-separated choice of |
| unset |
|
If CANLI_KEY is not set and local mode is off, call get_key once per session before the validators. The key it
returns lives only in this process's memory for the life of the session; it is not written to
disk.
HTTP failures and API error envelopes are marked as MCP tool errors while preserving the complete JSON envelope. A successful validation with a negative verdict remains a normal result. Requests have a 30-second deadline covering headers and body, reject redirects, and are never retried automatically. A timeout may occur after the service has processed a request; check service status before deciding to submit again. Non-JSON response bodies and raw network errors are omitted from tool errors.
Install
No install step. npx fetches the published package on first run, so every client config below
just spawns npx -y canli-validation-mcp. See "Local checkout" near the bottom to develop or test
this package itself instead of running the published one.
Hosted endpoint (no install)
The same tools are served at https://canlicapital.com/mcp over MCP Streamable HTTP, for clients
that connect to a URL instead of spawning a process (Claude.ai connectors, ChatGPT, Cursor's remote
servers). Nothing to install, and no Node.js on your machine.
claude mcp add --transport http canli https://canlicapital.com/mcpWithout a key, requests run under a shared anonymous key, so the daily validation quota is shared by every hosted caller. When that shared quota is used up for the day, validations are still answered, computed by the same code on the hosted endpoint, but without a stored receipt; the result says so. For your own quota, issue a free key (see /developers) and send it as a header:
claude mcp add --transport http canli https://canlicapital.com/mcp --header "Authorization: Bearer $CANLI_KEY"Add ?toolsets= to the URL to list only some tools (see "Toolsets" below), for example
https://canlicapital.com/mcp?toolsets=company.
The endpoint is stateless. On it, get_key issues nothing and says which key is in use, because a
key issued there would not reach the next request. A malformed Authorization header is refused
rather than replaced with the shared key.
Claude Desktop
Add to claude_desktop_config.json (Settings, Developer, Edit Config):
{
"mcpServers": {
"canli": {
"command": "npx",
"args": ["-y", "canli-validation-mcp"]
}
}
}Restart Claude Desktop afterward. Add an "env" object with CANLI_API_BASE to point this at a
preview deployment instead of the default.
Claude Code
claude mcp add canli -- npx -y canli-validation-mcpRun claude mcp list to confirm it is registered, and claude mcp remove canli to remove it.
Private local mode
Set CANLI_LOCAL=1 and the eight validators run on your machine: nothing about the series you
submit is sent to canlicapital.com, no key is needed, and no receipt is stored. The computation is
the API's own, shipped byte for byte in src/local (a test fails if it drifts), so a local result
equals the hosted one; it names no receipt id because none was made.
claude mcp add canli-local --env CANLI_LOCAL=1 -- npx -y canli-validation-mcpget_receipt, service_status and company_financial_history still read from canlicapital.com;
they send no series. In the Claude Desktop extension this is the "Private local mode" setting.
Generic stdio client
Any MCP client that can spawn a process and speak stdio will work. Using the official SDK directly, from Node:
import { Client } from "@modelcontextprotocol/sdk/client/index.js";
import { StdioClientTransport } from "@modelcontextprotocol/sdk/client/stdio.js";
const transport = new StdioClientTransport({
command: "npx",
args: ["-y", "canli-validation-mcp"],
env: { ...process.env, CANLI_API_BASE: "https://canlicapital.com" },
});
const client = new Client({ name: "my-agent", version: "0.1.0" });
await client.connect(transport);
const { tools } = await client.listTools();
console.log(tools.map((t) => t.name));
const keyResult = await client.callTool({ name: "get_key", arguments: { label: "my-agent" } });
if (keyResult.isError) throw new Error("Key setup failed; inspect the error privately.");
// The session retains the issued key. Avoid printing its envelope into logs.
const result = await client.callTool({
name: "validate_deflated_sharpe",
arguments: {
returns: [0.004, -0.002, 0.007, 0.001, -0.003, 0.005, 0.002, -0.001],
periods_per_year: 252,
effective_independent_trials: 30,
cross_trial_sharpe_sd_annualized: 0.5,
},
});
console.log(result.content[0].text); // the answer, its limits and its receipt
await client.close();What a result does not establish (boundary language)
Every envelope this server returns carries these sentences, verbatim, from the API itself
(api/_lib/limits.js):
This verdict is about the series exactly as submitted. The service never saw the data source, its costs, survivorship, or any lookahead in how the series was built.
A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.
The receipt is content-hashed and reproducible from the open-source core it names. It is not signed.
Quotas: 1000 validations per key per UTC day, 5 keys per client per UTC day, 1048576 bytes per request, 20000 observations per series, 200 variants per matrix.
Each tool's description also states one of these sentences, so an agent sees the boundary before it calls the tool, not only after.
Compact results (tokens)
A validation result is the answer (data), the sentences above except the quota line, and the
receipt's id and URL; an error keeps its error. The rest of the API envelope (schema, endpoint,
timestamps, claim and capital class, the human page, the source-file hashes and the quota line)
describes the service rather than the answer, and an agent pays for every token of it on every
call. It stays in the stored receipt, which get_receipt returns in full, and in service_status.
On a breadth result this is about half the text. Set CANLI_FULL_ENVELOPE=1 to receive every field.
Toolsets (tokens)
A client sends the model the whole tool list on every turn, and it is most of each turn's prompt:
a validation result is a few hundred tokens, the list of all fifteen tools several thousand. A
client that needs one kind of tool can list only that kind, with CANLI_TOOLSETS (stdio) or
?toolsets= (hosted endpoint). The default is every tool.
toolset | tools |
|
|
|
|
|
|
|
|
Measured with bench/tool_list_tokens.py (tokenizer: tiktoken o200k_base; other models'
tokenizers give different absolute counts), in the shape an OpenAI-style client sends the list:
CANLI_TOOLSETS | tools | tokens per turn | of all |
| 15 | 4,194 | 100% |
| 11 | 3,526 | 84% |
| 2 | 312 | 7% |
| 1 | 273 | 7% |
| 1 | 89 | 2% |
Providers cache a tool list that is identical from turn to turn and bill the cached part at a
fraction of the price (test/tool-list-stable.test.mjs keeps each list byte-stable); a smaller list
costs less either way.
Local checkout
Only needed to develop or test this package itself, not to run the published one.
cd mcp
npm ci
node src/server.mjsPoint a client at the checkout instead of npm by spawning node /absolute/path/to/meridian/mcp/src/server.mjs
in place of npx -y canli-validation-mcp in any config above.
Testing
npm test
npm run test:packageRuns node --test over test/*.test.mjs: schema round-trips against the API's own OpenAPI and
manifest examples, one success, error-envelope and quota-429 case per keyed tool (a fake fetch
stands in for the network), a check that every tool description carries a limits sentence, a
check that those sentences have not drifted from api/_lib/limits.js, a check that no shipped
file contains an em dash, and one test that spawns the actual server binary and performs a real
MCP tools/list and callTool handshake over stdio against a local HTTP stub, so the wiring is
proven rather than assumed.
test:package creates the actual npm tarball, checks its exact file list and license,
installs it in a temporary consumer directory, and runs the stdio test against that
installed entry. Dependency installation contacts npm; tool calls use only the local
HTTP stub. It neither publishes a package nor issues a production API key.
Dependencies
Only @modelcontextprotocol/server (pinned exact; the MCP SDK's server package, which itself
depends only on @modelcontextprotocol/core and zod) and zod (pinned exact): four packages in
the installed tree, where 0.7 and earlier pulled in about 95 through @modelcontextprotocol/sdk.
No other runtime dependency is added, and nothing in this package touches the site's root package.json,
.vercelignore, api/, scripts/, js/, or public/.
Available Tools
15 toolsaudit_backtestAudit a backtestA
One-call audit of a strategy's returns: deflated Sharpe, minimum track record and, with every variant's returns, the probability of backtest overfitting, each the matching validator's result with its own receipt. Point returns_file at the backtest's CSV or JSON instead of pasting long series. Prefer it to calling the validators one by one; one validation per check. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.
| Name | Required | Description | Default |
|---|---|---|---|
| returns | No | Periodic returns as fractions (0.01 = 1%), oldest first; replaces the Sharpe, observations, skew and kurtosis. | |
| n_splits | No | Even number of blocks, at least 2; default 16. | |
| variants | No | Optional returns of every variant tried (this one included), one row per period, one column per variant; adds the overfitting check. | |
| confidence | No | Between 0 and 1; default 0.95. | |
| returns_file | No | Path to a CSV or JSON of the returns on this machine (not on the hosted endpoint), instead of returns. | |
| variants_file | No | Path to a CSV or JSON with one numeric column per variant, instead of variants. | |
| returns_column | No | Column name or 1-based position, when returns_file has several numeric columns. | |
| periods_per_year | Yes | Periods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly. | |
| benchmark_sharpe_annualized | No | Annualized Sharpe to beat; default 0. | |
| effective_independent_trials | Yes | Independent variants tried before choosing this one. | |
| cross_trial_sharpe_sd_annualized | Yes | Standard deviation of annualized Sharpe across those trials. |
Output Schema
| Name | Required | Description |
|---|---|---|
| note | No | |
| error | No | |
| checks | No | |
| limits | No | |
| not_run | No | |
| readings | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and openWorldHint=true, and the description adds context consistent with that: each check returns the matching validator's result with its own receipt, and only one validation runs per check. It also discloses an interpretation limit ('not admission to anything and is not a forecast'), which is genuine behavioral guidance. It stops short of saying what is written or persisted, hence 4 rather than 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the aggregate purpose before the input and caveat details. It is dense and one clause ('each the matching validator's result with its own receipt') is grammatically awkward, but no sentence is filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and with 11 parameters at full schema coverage the description's job is mostly orientation. It covers the aggregation scope, the file-input alternative, the one-check-per-validation constraint, and the interpretation caveat, leaving only minor gaps such as failure behavior when returns_file is unreadable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds real meaning on top of it: returns_file is 'not on the hosted endpoint,' so the agent learns the path is resolved on the local machine, and it clarifies that supplying variants is what adds the overfitting check. That is useful beyond the schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('One-call audit of a strategy's returns') and enumerates exactly what it computes: deflated Sharpe, minimum track record, and, when variants are supplied, the probability of backtest overfitting. This maps directly onto the sibling validators (validate_deflated_sharpe, validate_track_record, validate_overfitting) so an agent can tell it apart from them without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes between alternatives: 'Prefer it to calling the validators one by one; one validation per check' and 'Point returns_file at the backtest's CSV or JSON instead of pasting long series.' Both the aggregate-vs-individual choice and the file-vs-inline choice are named with their selecting conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
company_financial_historyCompany financial history (SEC)ARead-only
SEC-reported financial history for one company in the canlicapital.com reference, by cik or ticker: without a concept, the histories available; with one, observations newest first with accession, form, filed date and unit, plus the source's SHA-256. For point-in-time values use canli-fundamentals-mcp. Public company accounting reference, not market prices, returns, an investment recommendation, or ALPHAC performance. Validate a separately constructed return series with the validation API; accounting values are not returns.
| Name | Required | Description | Default |
|---|---|---|---|
| cik | No | SEC CIK; send cik or ticker. | |
| limit | No | Most observations, newest first; default 40. | |
| ticker | No | Ticker such as AAPL; send ticker or cik. | |
| concept | No | us-gaap concept such as Assets; omit to list them. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | |
| source | No | |
| company | No | |
| history | No | |
| histories | No | |
| claim_boundary | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds genuine context beyond that: the response ordering (newest first), returned fields, the source SHA-256 for provenance, and the strong caveat that values are accounting figures rather than returns. It omits auth/rate-limit behavior, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, but the tail is padded with stacked disclaimers ('not market prices, returns, an investment recommendation, or ALPHAC performance' and a separate return-validation sentence). The intended meaning survives, but several clauses could be trimmed without loss.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be enumerated, and the description still covers the conditional response shape, ordering, provenance hash, and scope limits. For a read-only, 4-param tool this is near-complete; only auth/rate-limit context is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds real conditional semantics not in the schema: omitting 'concept' returns the list of available histories, while supplying it returns observations — behavior the schema names alone don't convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('SEC-reported financial history for one company ... by cik or ticker') and bounds the domain clearly (accounting reference, not market prices/returns). It does not name a sibling tool directly, but it points at external products and the validation API, so an agent can still separate it from the validate_* family.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives conditional usage ('without a concept, the histories available; with one, observations newest first') and explicit when-not routing ('For point-in-time values use canli-fundamentals-mcp', 'Validate a separately constructed return series with the validation API; accounting values are not returns'). Alternatives are named, but they are mostly outside the sibling set, so it stops short of a full routing rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_keyGet a free validation keyA
Issue a free validation key for this session. Rarely needed: the first validation issues one itself unless CANLI_KEY or local mode is set, and the read tools need none. Quotas: 1000 validations per key per UTC day, 5 keys per client per UTC day, 1048576 bytes per validation request, 1024 bytes per key revocation request, 20000 observations per series, 200 variants per matrix.
| Name | Required | Description | Default |
|---|---|---|---|
| label | No | Name for the key. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| note | No | |
| error | No | |
| key_source | No | |
| key_present | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only establish that this is an open-world, non-idempotent, non-destructive write. The description adds substantial operational context the agent cannot get elsewhere: per-day key and validation quotas, request and revocation byte limits, and series/matrix cardinality caps. This is exactly the kind of rate-limit and quota disclosure that earns full credit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose and the 'rarely needed' caveat before the quota list, so the most decision-relevant information comes first. The quota sentence is long and includes limits for adjacent endpoints (key revocation, series, matrices) that are not this tool's concern, which is mild scope creep.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail is not required, and the description fully covers prerequisites, when to skip the call, and quotas. It never says whether calling it twice produces the same or additional keys, though idempotentHint=false partially covers that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and there is a single optional 'label' parameter already documented in the schema. The description adds no meaning beyond it (no naming conventions, uniqueness, or default behavior), so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Issue a free validation key for this session'), which immediately distinguishes it from the validate_* siblings that consume such keys. The word 'free' and 'for this session' scopes the artifact being produced.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says the tool is 'rarely needed' and names the conditions under which it is unnecessary: the first validation self-issues a key unless CANLI_KEY or local mode is set, and read tools need none. That is genuine when-not-to-use guidance, not just a when.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_receiptGet a receiptARead-only
Fetch a stored verdict by receipt id to re-read it. No key; verify_receipt checks it is genuine. The receipt is content-hashed, reproducible from the open-source core it names, and signed with Ed25519 by a key published at https://canlicapital.com/.well-known/canli-receipt-keys.json.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Receipt id from a validation result. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| error | No | |
| limits | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this readOnly, non-destructive and openWorld, so the safety profile is covered. The description adds genuinely useful behavior beyond that: no key required, the receipt is content-hashed and reproducible from the open-source core, and it is Ed25519-signed by a key at a published well-known URL. Missing only error/absence behavior for an unknown id.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly packed sentences, front-loaded with the core action and immediately followed by the sibling disambiguation and trust model. No filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and annotations carry the safety profile. For a single-param read tool, the description supplies the routing, auth, and verification-model context an agent needs to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'id' param is fully documented in the schema (pattern plus provenance note). The description adds only the phrase 'by receipt id', so it neither compensates nor detracts; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Fetch) and resource (a stored verdict by receipt id) plus the intent (to re-read it), and explicitly differentiates itself from the sibling verify_receipt. An agent can distinguish it from the validate_* and verify_* tools without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It routes the agent: use this to re-read a stored receipt, use verify_receipt to check genuineness. It also states 'No key', clarifying the auth posture for invocation. It stops short of enumerating when-not-to-use (e.g., what to do if the id is unknown or the receipt is missing).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
service_statusService statusARead-only
Whether the validation API is up, with its quotas; check after a timeout before resubmitting. No key. This verdict is about the series exactly as submitted. The service never saw the data source, its costs, survivorship, or any lookahead in how the series was built.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| error | No | |
| limits | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so safety is covered. The description adds genuinely new context: 'No key' (no authentication required) and that quotas are returned, plus a caveat about the scope of the verdict.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and usage trigger are front-loaded well, but one of the three sentences is off-topic for a status endpoint and does not earn its place, diluting an otherwise tight definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and no input params need documenting. The remaining gap is that the trailing 'series as submitted' caveat is irrelevant to a service-status call and could mislead an agent about what this tool reports.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the baseline of 4 applies; there are no inputs whose semantics need explaining.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening clause ('Whether the validation API is up, with its quotas') states a specific resource and check, which distinguishes it from the validate_* siblings. However, the trailing sentences about 'the series exactly as submitted' and what 'the service never saw' belong to a validation-verdict tool, not a health check, and muddy what the tool actually is.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'check after a timeout before resubmitting' gives an explicit trigger condition for calling it, which is exactly the kind of context an agent needs. It does not name sibling alternatives, but the trigger is clear enough to route correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_backtest_lengthMinimum backtest lengthA
Minimum backtest length (years) before the best of N independent trials is not expected to reach a target Sharpe by luck; with backtest_years, the most trials those years allow. For planning a search; once it has a result, use validate_deflated_sharpe. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.
| Name | Required | Description | Default |
|---|---|---|---|
| backtest_years | No | Backtest length in years, for the most trials it allows. | |
| target_sharpe_annualized | No | In-sample annualized Sharpe you would call a discovery; default 1. | |
| effective_independent_trials | No | Independent trials tried (backtests, parameter sets, ideas). |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| note | No | |
| error | No | |
| limits | No | |
| receipt | No | |
| computed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, openWorldHint=true, destructiveHint=false, which is an odd profile for what the description frames as a pure planning calculation; the description does not reconcile that (no side effects, no persisted state, no external calls disclosed). It does add genuine interpretive context that outputs are not admissions or forecasts, which goes beyond the annotations. With annotations already present, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and the sibling handoff are front-loaded, which is good, but the opening is a single long clause-stacked sentence that is hard to parse on first read. The closing sentence about thresholds not being admission is useful but borders on editorial and is not clearly tied to what this tool returns.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the description covers purpose, planning-time usage, and the handoff tool. For a stateless computation with fully documented parameters, that is close to sufficient; only the reconciliation of the readOnlyHint=false annotation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and all three parameters are documented in the schema with bounds and defaults. The description restates backtest_years ('with backtest_years, the most trials those years allow') without adding syntax or format meaning beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific quantity (minimum backtest length in years) and the condition under which it holds (best of N independent trials not reaching target Sharpe by luck), which is far more specific than the title alone. It is somewhat syntactically dense and reads like a textbook definition rather than a plain 'what this does', and it does not directly contrast itself with siblings like validate_luck_trials.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the setting explicitly ('For planning a search') and names the alternative with the condition that selects it ('once it has a result, use validate_deflated_sharpe'). It also adds a when-not caveat: a deflated Sharpe or overfitting probability past any threshold is not admission to anything, so it should not be used as a decision gate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_breadthValidate breadth ceilingA
Book Sharpe ceiling from adding sleeves of this quality and correlation, the Sharpe at a sleeve count, and the sleeves a target needs. For portfolio construction; it validates no single strategy. This verdict is about the series exactly as submitted. The service never saw the data source, its costs, survivorship, or any lookahead in how the series was built.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | Target book Sharpe, for the sleeves it needs. | |
| sleeves | No | Sleeve count, for that book's Sharpe. | |
| sleeve_sharpe | Yes | Annualized Sharpe of one sleeve. | |
| average_pairwise_correlation | Yes | Average correlation between sleeves, -1 to 1. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| note | No | |
| error | No | |
| limits | No | |
| receipt | No | |
| computed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses meaningful scope limits: the verdict applies only to the series as submitted and the service never inspected data source, costs, survivorship, or lookahead. That tells the agent the result is a purely mathematical check. The readOnlyHint=false/openWorldHint=true annotations are neither repeated nor contradicted; the description adds real context rather than echoing them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is reasonably short (four sentences) and front-loads the outputs, but the first sentence is dense and awkwardly constructed, forcing re-reading to parse the three-output mapping. The disclaimer sentences earn their place, but the core sentence could be clearer.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a non-trivial portfolio-math tool, the description covers purpose, applicability, the three outputs' drivers, and the critical caveat about what data was never inspected. An output schema exists, so return-value explanation is not required. The only gap is the absence of explicit guidance on when to prefer this over a sibling validator.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds the parameter-to-output relationship: sleeve_sharpe plus correlation determine the ceiling, sleeve count determines that book's Sharpe, and target determines the sleeve count needed. That pairing is not obvious from the schema alone, so it earns above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the three concrete outputs (book Sharpe ceiling, Sharpe at a sleeve count, sleeves required for a target), which is a specific computational purpose rather than a restatement of the name. It also differentiates itself from the sibling single-strategy validators with 'For portfolio construction; it validates no single strategy.' The opening sentence is grammatically elliptical ('Book Sharpe ceiling from adding sleeves...') but the intent is recoverable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one usage signal — 'For portfolio construction; it validates no single strategy' — which implicitly routes the agent away from the sleeve-less validate_* siblings (validate_deflated_sharpe, validate_haircut_sharpe, etc.). However, it never names an alternative or states a when-not condition, so the selection guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_deflated_sharpeValidate deflated SharpeA
Deflated Sharpe ratio: the probability (0 to 1) that the selected strategy's Sharpe beats the best that luck gives across the variants tried, with the probabilistic Sharpe and that luck benchmark. Send the seven statistics or a return series. With every variant's returns use validate_overfitting; luck as a trial count, validate_luck_trials; a multiple-testing haircut, validate_haircut_sharpe. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.
| Name | Required | Description | Default |
|---|---|---|---|
| skew | No | Skewness of returns; 0 if Normal. | |
| returns | No | Periodic returns as fractions (0.01 = 1%), oldest first; replaces the Sharpe, observations, skew and kurtosis. | |
| observations | No | Number of return observations. | |
| periods_per_year | No | Periods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly. | |
| non_excess_kurtosis | No | Kurtosis, not excess kurtosis; 3 if Normal. | |
| observed_sharpe_annualized | No | Annualized Sharpe as observed. | |
| effective_independent_trials | No | Independent variants tried before choosing this one. | |
| cross_trial_sharpe_sd_annualized | No | Standard deviation of annualized Sharpe across those trials. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| note | No | |
| error | No | |
| limits | No | |
| receipt | No | |
| computed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, openWorldHint=true and destructiveHint=false, so safety signalling is already covered. The description still adds real context beyond that: it explains the two valid input modes (statistics vs. return series) and warns that the resulting probability is not admission to anything and not a forecast. It does not mention auth, cost, or rate limits, but for a local statistical computation that gap is small.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The whole definition fits in three dense sentences with no filler; the routing alternatives are front-loaded mid-paragraph and the interpretive caveat closes it. It is slightly run-on and could be split, but every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation. For a tool with 8 optional parameters and no required inputs, the description covers purpose, the two input modes, sibling alternatives, and interpretation limits. It could go further on which combination of statistics is minimally sufficient, but nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so baseline would be 3, but the description adds the either/or input mode ('the seven statistics or a return series') that the schema only implies via the 'replaces' note on returns. This clarifies mutual exclusivity across the 8 all-optional parameters, which is meaningful routing information beyond the field-level descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states precisely what the tool computes: the probability (0-1) that the selected strategy's Sharpe beats the best luck produces across the variants tried, plus the probabilistic Sharpe and luck benchmark. It also names the sibling tools it is not (validate_overfitting, validate_luck_trials, validate_haircut_sharpe), so the agent can separate it from adjacent validators. It falls short of 5 only because it opens with a concept definition rather than a direct verb+resource statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit routing: send the seven statistics or a return series; use validate_overfitting when every variant's returns are available; validate_luck_trials when luck is expressed as a trial count; validate_haircut_sharpe for a multiple-testing haircut. The alternative tools and the conditions selecting them are spelled out rather than left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_haircut_sharpeHaircut Sharpe ratioA
Haircut Sharpe for multiple testing (Harvey and Liu 2015): the Sharpe a single test would have needed, by Bonferroni and independent tests, and with the other tests' Sharpes, Holm and BHY. For the probability the Sharpe is real, use validate_deflated_sharpe. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.
| Name | Required | Description | Default |
|---|---|---|---|
| tests | No | Tests run, this one included; gives the Bonferroni and independent-test haircuts. | |
| observations | Yes | Return observations behind the Sharpe. | |
| autocorrelation | No | Lag-1 autocorrelation of returns, -1 to 1; default 0. Corrects the annualized Sharpe (Lo 2002). | |
| periods_per_year | Yes | Periods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly. | |
| observed_sharpe_annualized | Yes | Annualized Sharpe as observed. | |
| other_sharpe_ratios_annualized | No | Annualized Sharpes of the other tests over the same observations; adds Holm and BHY. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| note | No | |
| error | No | |
| limits | No | |
| receipt | No | |
| computed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (destructiveHint=false, openWorldHint=true), so the description's burden is lighter. It adds interpretive behavior ('not admission to anything and is not a forecast') and clarifies which inputs drive which outputs. It does not state that the call is a pure stateless computation or whether anything is persisted, which matters given readOnlyHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the method and citation, no filler. The second sentence packs two distinct ideas (sibling routing and threshold caveat) and is grammatically dense, but every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained; the description still covers method provenance, the routing alternative, and a misuse caveat. It leaves open how this differs operationally from related siblings like validate_luck_trials and validate_overfitting, which is the main remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all six parameters including worked examples (252 daily, 365 crypto). The description only loosely links inputs to outputs ('by Bonferroni and independent tests', 'with the other tests' Sharpes, Holm and BHY'), which is baseline-level added value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific computation (haircut Sharpe for multiple testing, Harvey and Liu 2015) and enumerates exactly what it returns: Bonferroni/independent-test haircuts and Holm/BHY adjustments. It also names the sibling it is not (validate_deflated_sharpe), so an agent can route without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives one explicit routing rule: 'For the probability the Sharpe is real, use validate_deflated_sharpe.' That is a real alternative-plus-condition. It does not cover when to prefer this over other siblings such as validate_overfitting or validate_luck_trials, so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_luck_trialsLuck-equivalent trialsA
How many skill-less strategies a search would need for its best to reach this Sharpe by luck (Monte Carlo), and with a trial count, the chance it did. States luck as the best of N random tries; for the probability the Sharpe is real, use validate_deflated_sharpe. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.
| Name | Required | Description | Default |
|---|---|---|---|
| skew | No | Skewness; below -0.5 the reading warns the counts are too generous. | |
| observations | Yes | Return observations behind the Sharpe. | |
| autocorrelation | No | Lag-1 autocorrelation of returns, -1 to 1; default 0. Corrects the annualized Sharpe (Lo 2002). | |
| periods_per_year | Yes | Periods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly. | |
| observed_sharpe_annualized | Yes | Annualized Sharpe as observed. | |
| effective_independent_trials | No | Independent trials tried; adds the chance the best reached this Sharpe by luck. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| note | No | |
| error | No | |
| limits | No | |
| receipt | No | |
| computed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations supply the safety profile (destructiveHint=false, openWorldHint=true), and the description adds genuine extra context: the method is Monte Carlo and luck is defined as the best of N random tries. It also warns that a threshold crossing is neither admission nor forecast, which is useful interpretive guidance. It says nothing about why readOnlyHint is false for what reads as a pure computation, leaving an unexplained annotation/label mismatch.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core computation before the sibling pointer and the caveat. Every sentence carries content, though the opening sentence is syntactically heavy and could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and annotations cover the safety profile. Purpose, sibling routing, and interpretive limits are all present; the main residual gap is the unaddressed readOnlyHint=false for a statistical computation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameter meanings, defaults (autocorrelation default 0), and bounds are already documented; baseline is 3. The description only implies effective_independent_trials by calling it 'a trial count' and does not add syntax or semantics beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific computation: how many skill-less strategies a search would need for its best to reach the observed Sharpe by luck, and (given a trial count) the probability of that. It names the sibling it is not (validate_deflated_sharpe), so the agent can separate the two. The phrasing is dense and inverted, which costs it a point but the substance is specific and non-tautological.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly routes the agent: 'for the probability the Sharpe is real, use validate_deflated_sharpe', and the clause 'with a trial count, the chance it did' signals when to supply effective_independent_trials. The closing caveat about thresholds also frames how to read the result. There is no explicit when-not, but the alternative pointer is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_overfittingValidate overfitting (CSCV)A
Probability of backtest overfitting (0 to 1) by CSCV: how often the in-sample best variant falls below the out-of-sample median. Needs every variant's returns (periods by variants); with summary statistics only, use validate_deflated_sharpe. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Sampling seed; default 42. | |
| matrix | Yes | Returns as fractions, one row per period, one column per variant. | |
| n_splits | No | Even number of blocks, at least 2; default 16. | |
| max_combinations | No | Most splits evaluated, up to 2000 (default). |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| note | No | |
| error | No | |
| limits | No | |
| receipt | No | |
| computed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnly=false, openWorld=true, destructive=false, which says little about a statistical computation. The description compensates with the required input shape and a meaningful interpretive caveat (threshold crossings are neither admission nor forecast), which is real behavioral guidance beyond structured fields. It stops short of runtime traits like cost, determinism, or failure modes when the matrix is malformed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the metric and method, then requirements, then sibling routing, then the interpretation caveat. Three dense sentences with no filler, each carrying distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description covers inputs, the alternative tool, and result interpretation for a non-trivial statistical procedure. It could say more about assumptions (e.g. minimum period/variant counts) or determinism, but nothing essential to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents matrix, seed, n_splits, and max_combinations. The description's 'periods by variants' restates the matrix schema description rather than adding syntax or constraints. Baseline 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific outcome (probability of backtest overfitting, 0-1) and the method (CSCV), then defines the metric operationally as how often the in-sample best variant falls below the out-of-sample median. It explicitly distinguishes itself from validate_deflated_sharpe, so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the input condition that selects this tool (every variant's returns required) and names the alternative for the other case (summary statistics only -> validate_deflated_sharpe). This is an explicit when/alternatives pairing, not implied usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_paper_evidenceValidate paper evidenceA
Whether a paper or simulated performance record meets canli.paper-evidence.v0, with a JSON pointer per failure. Checks structure and required disclosures, not whether the returns are good. This verdict is about the series exactly as submitted. The service never saw the data source, its costs, survivorship, or any lookahead in how the series was built.
| Name | Required | Description | Default |
|---|---|---|---|
| record | Yes | A canli.paper-evidence.v0 record. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| note | No | |
| error | No | |
| limits | No | |
| receipt | No | |
| computed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=false, openWorldHint=true, destructiveHint=false, so the description usefully adds that the verdict concerns the submitted series only and that the service never saw the data source, costs, survivorship, or lookahead. This is exactly the kind of limitation an agent needs. It stops short of stating permission requirements or whether calling it has side effects (notably, readOnlyHint=false is left unexplained).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and output, and each sentence is short. The final two sentences overlap somewhat, both reassuring about what the verdict does not cover, which is mild redundancy rather than waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the description correctly delimits what the verdict does and does not assert. Given the single nested object parameter, this is nearly complete; only the missing sibling routing keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
One parameter with 100% schema description coverage, so the schema already carries the semantics. The description's mention of canli.paper-evidence.v0 simply restates the schema's own description and adds no new syntax or format detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource (validates a paper/simulated performance record against canli.paper-evidence.v0) and states the output shape (JSON pointer per failure). It does not, however, differentiate itself from the many sibling validate_* tools, leaving the agent to infer which validator applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clarifies scope by negation ('checks structure and required disclosures, not whether the returns are good'), which implies when the tool is appropriate. But it never names an alternative or states a positive trigger condition versus siblings like validate_track_record or validate_backtest.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_reality_checkData-snooping tests (SPA, Reality Check, StepM)A
Data-snooping tests on every variant a search tried: Hansen's SPA p-value that the best beat the benchmark only by luck, White's Reality Check, and the variants Romano-Wolf StepM finds better. Send all variants tried, not only the winners. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.
| Name | Required | Description | Default |
|---|---|---|---|
| reps | No | Bootstrap draws; default 2000. | |
| seed | No | Sampling seed; default 42. | |
| alpha | No | Familywise error for StepM; default 0.05. | |
| matrix | No | Returns of every variant the search tried, one row per period, one column per variant. | |
| benchmark | No | Benchmark return per period; default zero. | |
| matrix_file | No | Path to a CSV or JSON with one numeric column per variant, instead of matrix. | |
| block_length | No | Mean bootstrap block in periods; default round(n^(1/3)). |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| note | No | |
| error | No | |
| limits | No | |
| receipt | No | |
| computed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry the safety profile (destructiveHint=false, openWorldHint=true), and the description adds a genuine interpretive caveat beyond them: a deflated Sharpe or overfitting probability above/below a threshold "is not admission to anything and is not a forecast." That tells the agent the output is diagnostic, not a decision — useful context the annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with what the tool computes and then the input rule. Dense but each clause earns its place; minor run-on structure in the second sentence keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained; all seven parameters are fully described in the schema, and the description supplies the key input expectation (include every variant tried). For a stateless computation tool this is essentially complete, missing only explicit sibling routing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all seven parameters including defaults and bounds, and the baseline is 3. The description only reinforces the matrix requirement ("every variant a search tried"), which the schema already states, so it adds little beyond the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific resource (data-snooping tests) plus the three concrete tests it runs (SPA p-value, White's Reality Check, Romano-Wolf StepM), so the agent knows exactly what computation happens. It only implicitly separates itself from siblings like validate_deflated_sharpe and validate_overfitting via the closing caveat rather than stating routing outright.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Send all variants tried, not only the winners" is an explicit, actionable usage directive that tells the agent what input to gather, which most siblings don't provide. It stops short of naming when to pick this over validate_deflated_sharpe or validate_overfitting, so it is clear context rather than full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_track_recordMinimum track record lengthA
Minimum track record length (observations and years) for an observed Sharpe to beat a benchmark at a confidence level; with observations, the record's probabilistic Sharpe so far. For live or paper records; to size a backtest for its trials, use validate_backtest_length. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.
| Name | Required | Description | Default |
|---|---|---|---|
| skew | Yes | Skewness of returns; 0 if Normal. | |
| confidence | No | Between 0 and 1; default 0.95. | |
| observations | No | Record length so far, for its probabilistic Sharpe. | |
| periods_per_year | Yes | Periods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly. | |
| non_excess_kurtosis | Yes | Kurtosis, not excess kurtosis; 3 if Normal. | |
| observed_sharpe_annualized | Yes | Annualized Sharpe as observed. | |
| benchmark_sharpe_annualized | No | Annualized Sharpe to beat; default 0. |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | |
| note | No | |
| error | No | |
| limits | No | |
| receipt | No | |
| computed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations (readOnlyHint=false, destructiveHint=false, openWorldHint=true) are somewhat at odds with an obvious pure-calculation tool, and the description neither confirms nor contradicts them. It does add interpretive context ('not admission to anything and is not a forecast') and discloses that output includes observations/years, but says nothing about precision, numerical caveats, or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core computation and then the sibling routing. The final disclaimer sentence is useful but slightly tangential and the first sentence is dense with nested clauses, costing a little clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and full schema coverage, the description does not need to explain return values, and it covers purpose, scope, and routing. It is complete enough to invoke correctly, though it never explains what the periodic/year conversion or confidence inputs imply about results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already defines every parameter, including defaults (benchmark 0, confidence 0.95). The description names 'observations', 'confidence level', and benchmark only in passing and adds no format or edge-case semantics beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific computation (minimum track record length, in observations and years) on a named resource (an observed Sharpe vs a benchmark) and explicitly names the sibling it is not: validate_backtest_length. An agent can distinguish it from the several validate_* siblings without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when ('For live or paper records') and an explicit alternative with its selecting condition ('to size a backtest for its trials, use validate_backtest_length'). It also warns what the tool is not for, which is exactly the routing guidance an agent needs among near-identical validate_* siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_receiptVerify a receiptARead-only
Verify a receipt offline: its Ed25519 signature against the bundled canlicapital.com key, its output hash and its id. Send an id to fetch it first, or the receipt itself. The receipt is content-hashed, reproducible from the open-source core it names, and signed with Ed25519 by a key published at https://canlicapital.com/.well-known/canli-receipt-keys.json.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Receipt id from a validation result. | |
| receipt | No | A receipt as get_receipt returns it, to verify without fetching. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | No | |
| valid | No | |
| checks | No | |
| key_id | No | |
| meaning | No | |
| receipt_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true and openWorldHint=true, and the description goes beyond them by disclosing that verification happens offline, that the trust anchor is a specific bundled key, and that the key is published at a well-known URL. It stops short of describing failure modes or what makes verification fail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense, front-loaded sentences that lead with the action and the checks performed; the reproducibility and key-publication details are useful but slightly burden the second sentence. Nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need no explanation, and the annotations carry the safety profile. The description covers the cryptographic scope and both input paths, leaving only edge-case behavior (e.g., malformed receipts, network failure when fetching by id) unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, and the description adds genuine semantics: passing an id triggers an internal fetch while passing a receipt object verifies without fetching. That behavioral distinction between the two optional parameters is not evident from the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (verify) and resource (receipt) and enumerates exactly what is checked: the Ed25519 signature, the output hash, and the id. The phrase 'Send an id to fetch it first' implicitly differentiates it from the sibling get_receipt, which only retrieves.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly spells out the two mutually exclusive input modes ('Send an id to fetch it first, or the receipt itself'), which is real usage guidance for an agent choosing how to call it. It does not, however, state when verification is unnecessary or how it relates to validate_* siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
15 tool updates
v0.10.1- Changed
audit_backtest10 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Input schema / properties / cross_trial_sharpe_sd_annualized / descriptionPrevious value: -"Standard deviation of the annualized Sharpe across those trials."New value: +"Standard deviation of annualized Sharpe across those trials." - changed
Input schema / properties / periods_per_year / descriptionPrevious value: -"Observations per year: 252 daily, 365 daily crypto, 52 weekly, 12 monthly."New value: +"Periods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly." - changed
Input schema / properties / returns / descriptionPrevious value: -"Periodic returns as fractions (0.01 is 1%), oldest first; replaces the Sharpe, observations, skew and kurtosis fields."New value: +"Periodic returns as fractions (0.01 = 1%), oldest first; replaces the Sharpe, observations, skew and kurtosis." - changed
Input schema / properties / returns_column / descriptionPrevious value: -"Header name or 1-based position of the returns column when returns_file has several numeric columns."New value: +"Column name or 1-based position, when returns_file has several numeric columns." - changed
Input schema / properties / returns_file / descriptionPrevious value: -"Path to a CSV or JSON file of the returns on the machine running this server, instead of returns. Not available on the hosted endpoint."New value: +"Path to a CSV or JSON of the returns on this machine (not on the hosted endpoint), instead of returns." - changed
Input schema / properties / variants / descriptionPrevious value: -"Optional returns of every variant tried, this one included, as fractions: one row per period, one column per variant. Adds the overfitting check."New value: +"Optional returns of every variant tried (this one included), one row per period, one column per variant; adds the overfitting check." - added
Input schema / properties / variants / items / maxItemsAdded value: +200 - changed
Input schema / properties / variants_file / descriptionPrevious value: -"Path to a CSV or JSON file of every variant's returns (one numeric column per variant), instead of variants."New value: +"Path to a CSV or JSON with one numeric column per variant, instead of variants." - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": {}, + "description": "checks holds each check's full validate_ result by name; readings one sentence per check; not_run the checks skipped and why.", + "properties": { + "checks": { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + "error": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + }, + "limits": { + "items": { + "type": "string" + }, + "type": "array" + }, + "not_run": { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + "note": { + "type": "string" + }, + "readings": { + "additionalProperties": {}, + "properties": {}, + "type": "object" + } + }, + "type": "object" +}
- Changed
company_financial_history2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": {}, + "description": "Without a concept, histories lists what is available; with one, history holds the observations newest first as columns and rows, each with its filing. Values are as reported to the SEC.", + "properties": { + "claim_boundary": { + "type": "string" + }, + "company": { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + "error": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + }, + "histories": { + "items": {}, + "type": "array" + }, + "history": { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + "source": { + "additionalProperties": {}, + "properties": {}, + "type": "object" + } + }, + "type": "object" +}
- Changed
get_key2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": {}, + "description": "key_source says which key the session uses; data.key is set when a new key was issued.", + "properties": { + "data": { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + "error": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + }, + "key_present": { + "type": "boolean" + }, + "key_source": { + "type": "string" + }, + "note": { + "type": "string" + } + }, + "type": "object" +}
- Changed
get_receipt2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": {}, + "description": "data is the stored receipt: validator, input, output, their hashes, source hashes and signature.", + "properties": { + "data": { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + "error": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + }, + "limits": { + "items": { + "type": "string" + }, + "type": "array" + } + }, + "type": "object" +}
- Changed
service_status2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": {}, + "description": "data.store_reachable, data.quotas and, when available, data.usage for this key.", + "properties": { + "data": { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + "error": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + }, + "limits": { + "items": { + "type": "string" + }, + "type": "array" + } + }, + "type": "object" +}
- Changed
validate_backtest_length5 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Input schema / properties / backtest_years / descriptionPrevious value: -"Length of the backtest in years, to get the most independent trials it allows."New value: +"Backtest length in years, for the most trials it allows." - changed
Input schema / properties / effective_independent_trials / descriptionPrevious value: -"Independent trials (backtests, parameter sets, ideas) tried; gives the minimum backtest length."New value: +"Independent trials tried (backtests, parameter sets, ideas)." - changed
Input schema / properties / target_sharpe_annualized / descriptionPrevious value: -"In-sample annualized Sharpe you would take as a discovery; default 1."New value: +"In-sample annualized Sharpe you would call a discovery; default 1." - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": {}, + "description": "data.result holds the statistics (data.plain_reading states them); limits say what the result does not establish; receipt ({id, url}) is the signed record, null when none was stored; error ({code, message}) is set only on refusal.", + "properties": { + "computed": { + "type": "string" + }, + "data": { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + "error": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + }, + "limits": { + "items": { + "type": "string" + }, + "type": "array" + }, + "note": { + "type": "string" + }, + "receipt": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + } + }, + "type": "object" +}
- Changed
validate_breadth4 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Input schema / properties / sleeves / descriptionPrevious value: -"Sleeve count, to get that book's Sharpe."New value: +"Sleeve count, for that book's Sharpe." - changed
Input schema / properties / target / descriptionPrevious value: -"Target book Sharpe, to get the sleeves it needs."New value: +"Target book Sharpe, for the sleeves it needs." - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": {}, + "description": "data.result holds the statistics (data.plain_reading states them); limits say what the result does not establish; receipt ({id, url}) is the signed record, null when none was stored; error ({code, message}) is set only on refusal.", + "properties": { + "computed": { + "type": "string" + }, + "data": { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + "error": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + }, + "limits": { + "items": { + "type": "string" + }, + "type": "array" + }, + "note": { + "type": "string" + }, + "receipt": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + } + }, + "type": "object" +}
- Changed
validate_deflated_sharpe6 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Input schema / properties / cross_trial_sharpe_sd_annualized / descriptionPrevious value: -"Standard deviation of the annualized Sharpe across those trials."New value: +"Standard deviation of annualized Sharpe across those trials." - changed
Input schema / properties / periods_per_year / descriptionPrevious value: -"Observations per year: 252 daily, 365 daily crypto, 52 weekly, 12 monthly."New value: +"Periods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly." - changed
Input schema / properties / returns / descriptionPrevious value: -"Periodic returns as fractions (0.01 is 1%), oldest first; replaces the Sharpe, observations, skew and kurtosis fields."New value: +"Periodic returns as fractions (0.01 = 1%), oldest first; replaces the Sharpe, observations, skew and kurtosis." - changed
Input schema / properties / skew / descriptionPrevious value: -"Skewness of the returns; below -0.5 the reading warns that the counts are too generous."New value: +"Skewness of returns; 0 if Normal." - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": {}, + "description": "data.result holds the statistics (data.plain_reading states them); limits say what the result does not establish; receipt ({id, url}) is the signed record, null when none was stored; error ({code, message}) is set only on refusal.", + "properties": { + "computed": { + "type": "string" + }, + "data": { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + "error": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + }, + "limits": { + "items": { + "type": "string" + }, + "type": "array" + }, + "note": { + "type": "string" + }, + "receipt": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + } + }, + "type": "object" +}
- Changed
validate_haircut_sharpe7 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Input schema / properties / autocorrelation / descriptionPrevious value: -"First-order autocorrelation of the returns, -1 to 1; default 0. Corrects the annualized Sharpe as Lo (2002)."New value: +"Lag-1 autocorrelation of returns, -1 to 1; default 0. Corrects the annualized Sharpe (Lo 2002)." - changed
Input schema / properties / observations / descriptionPrevious value: -"Number of return observations behind the Sharpe ratio."New value: +"Return observations behind the Sharpe." - changed
Input schema / properties / other_sharpe_ratios_annualized / descriptionPrevious value: -"Annualized Sharpe ratios of the other tests, over the same observations; adds the Holm and BHY haircuts."New value: +"Annualized Sharpes of the other tests over the same observations; adds Holm and BHY." - changed
Input schema / properties / periods_per_year / descriptionPrevious value: -"Observations per year: 252 daily, 365 daily crypto, 52 weekly, 12 monthly."New value: +"Periods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly." - changed
Input schema / properties / tests / descriptionPrevious value: -"Total tests run, this one included; gives the Bonferroni and independent-test haircuts."New value: +"Tests run, this one included; gives the Bonferroni and independent-test haircuts." - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": {}, + "description": "data.result holds the statistics (data.plain_reading states them); limits say what the result does not establish; receipt ({id, url}) is the signed record, null when none was stored; error ({code, message}) is set only on refusal.", + "properties": { + "computed": { + "type": "string" + }, + "data": { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + "error": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + }, + "limits": { + "items": { + "type": "string" + }, + "type": "array" + }, + "note": { + "type": "string" + }, + "receipt": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + } + }, + "type": "object" +}
- Changed
validate_luck_trials7 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Input schema / properties / autocorrelation / descriptionPrevious value: -"First-order autocorrelation of the returns, -1 to 1; default 0. Corrects the annualized Sharpe as Lo (2002)."New value: +"Lag-1 autocorrelation of returns, -1 to 1; default 0. Corrects the annualized Sharpe (Lo 2002)." - changed
Input schema / properties / effective_independent_trials / descriptionPrevious value: -"Independent trials tried; adds the chance that the best of them reached this Sharpe by luck."New value: +"Independent trials tried; adds the chance the best reached this Sharpe by luck." - changed
Input schema / properties / observations / descriptionPrevious value: -"Number of return observations behind the Sharpe ratio."New value: +"Return observations behind the Sharpe." - changed
Input schema / properties / periods_per_year / descriptionPrevious value: -"Observations per year: 252 daily, 365 daily crypto, 52 weekly, 12 monthly."New value: +"Periods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly." - changed
Input schema / properties / skew / descriptionPrevious value: -"Skewness of the returns; below -0.5 the reading warns that the counts are too generous."New value: +"Skewness; below -0.5 the reading warns the counts are too generous." - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": {}, + "description": "data.result holds the statistics (data.plain_reading states them); limits say what the result does not establish; receipt ({id, url}) is the signed record, null when none was stored; error ({code, message}) is set only on refusal.", + "properties": { + "computed": { + "type": "string" + }, + "data": { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + "error": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + }, + "limits": { + "items": { + "type": "string" + }, + "type": "array" + }, + "note": { + "type": "string" + }, + "receipt": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + } + }, + "type": "object" +}
- Changed
validate_overfitting4 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Input schema / properties / matrix / descriptionPrevious value: -"Returns as fractions: one row per period, one column per variant."New value: +"Returns as fractions, one row per period, one column per variant." - added
Input schema / properties / matrix / items / maxItemsAdded value: +200 - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": {}, + "description": "data.result holds the statistics (data.plain_reading states them); limits say what the result does not establish; receipt ({id, url}) is the signed record, null when none was stored; error ({code, message}) is set only on refusal.", + "properties": { + "computed": { + "type": "string" + }, + "data": { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + "error": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + }, + "limits": { + "items": { + "type": "string" + }, + "type": "array" + }, + "note": { + "type": "string" + }, + "receipt": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + } + }, + "type": "object" +}
- Changed
validate_paper_evidence2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": {}, + "description": "data.result holds the statistics (data.plain_reading states them); limits say what the result does not establish; receipt ({id, url}) is the signed record, null when none was stored; error ({code, message}) is set only on refusal.", + "properties": { + "computed": { + "type": "string" + }, + "data": { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + "error": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + }, + "limits": { + "items": { + "type": "string" + }, + "type": "array" + }, + "note": { + "type": "string" + }, + "receipt": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + } + }, + "type": "object" +}
- Added
validate_reality_check - Changed
validate_track_record5 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Input schema / properties / observations / descriptionPrevious value: -"Record length so far, to get its probabilistic Sharpe."New value: +"Record length so far, for its probabilistic Sharpe." - changed
Input schema / properties / periods_per_year / descriptionPrevious value: -"Observations per year: 252 daily, 365 daily crypto, 52 weekly, 12 monthly."New value: +"Periods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly." - changed
Input schema / properties / skew / descriptionPrevious value: -"Skewness of the returns; below -0.5 the reading warns that the counts are too generous."New value: +"Skewness of returns; 0 if Normal." - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": {}, + "description": "data.result holds the statistics (data.plain_reading states them); limits say what the result does not establish; receipt ({id, url}) is the signed record, null when none was stored; error ({code, message}) is set only on refusal.", + "properties": { + "computed": { + "type": "string" + }, + "data": { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + "error": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + }, + "limits": { + "items": { + "type": "string" + }, + "type": "array" + }, + "note": { + "type": "string" + }, + "receipt": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + } + }, + "type": "object" +}
- Changed
verify_receipt3 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - changed
Input schema / properties / receipt / descriptionPrevious value: -"A receipt as get_receipt returns it (its data), to verify without fetching it."New value: +"A receipt as get_receipt returns it, to verify without fetching." - changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "additionalProperties": {}, + "description": "valid is true only when the signature, output hash and content id all check out; checks lists each.", + "properties": { + "checks": {}, + "error": { + "anyOf": [ + { + "additionalProperties": {}, + "properties": {}, + "type": "object" + }, + { + "type": "null" + } + ] + }, + "key_id": { + "type": [ + "string", + "null" + ] + }, + "meaning": { + "type": "string" + }, + "receipt_id": { + "type": [ + "string", + "null" + ] + }, + "valid": { + "type": "boolean" + } + }, + "type": "object" +}
13 tool updates
v0.6.0- Added
audit_backtest - Changed
company_financial_history4 fields changed- added
Input schema / properties / cik / descriptionAdded value: +"SEC CIK; send cik or ticker." - added
Input schema / properties / concept / descriptionAdded value: +"us-gaap concept such as Assets; omit to list them." - added
Input schema / properties / limit / descriptionAdded value: +"Most observations, newest first; default 40." - added
Input schema / properties / ticker / descriptionAdded value: +"Ticker such as AAPL; send ticker or cik."
- Changed
get_key1 field changed- added
Input schema / properties / label / descriptionAdded value: +"Name for the key."
- Changed
get_receipt1 field changed- added
Input schema / properties / id / descriptionAdded value: +"Receipt id from a validation result."
- Added
validate_backtest_length - Changed
validate_breadth4 fields changed- added
Input schema / properties / average_pairwise_correlation / descriptionAdded value: +"Average correlation between sleeves, -1 to 1." - added
Input schema / properties / sleeve_sharpe / descriptionAdded value: +"Annualized Sharpe of one sleeve." - added
Input schema / properties / sleeves / descriptionAdded value: +"Sleeve count, to get that book's Sharpe." - added
Input schema / properties / target / descriptionAdded value: +"Target book Sharpe, to get the sleeves it needs."
- Changed
validate_deflated_sharpe8 fields changed- added
Input schema / properties / cross_trial_sharpe_sd_annualized / descriptionAdded value: +"Standard deviation of the annualized Sharpe across those trials." - added
Input schema / properties / effective_independent_trials / descriptionAdded value: +"Independent variants tried before choosing this one." - added
Input schema / properties / non_excess_kurtosis / descriptionAdded value: +"Kurtosis, not excess kurtosis; 3 if Normal." - added
Input schema / properties / observations / descriptionAdded value: +"Number of return observations." - added
Input schema / properties / observed_sharpe_annualized / descriptionAdded value: +"Annualized Sharpe as observed." - added
Input schema / properties / periods_per_year / descriptionAdded value: +"Observations per year: 252 daily, 365 daily crypto, 52 weekly, 12 monthly." - added
Input schema / properties / returns / descriptionAdded value: +"Periodic returns as fractions (0.01 is 1%), oldest first; replaces the Sharpe, observations, skew and kurtosis fields." - added
Input schema / properties / skew / descriptionAdded value: +"Skewness of the returns; below -0.5 the reading warns that the counts are too generous."
- Added
validate_haircut_sharpe - Added
validate_luck_trials - Changed
validate_overfitting4 fields changed- added
Input schema / properties / matrix / descriptionAdded value: +"Returns as fractions: one row per period, one column per variant." - added
Input schema / properties / max_combinations / descriptionAdded value: +"Most splits evaluated, up to 2000 (default)." - added
Input schema / properties / n_splits / descriptionAdded value: +"Even number of blocks, at least 2; default 16." - added
Input schema / properties / seed / descriptionAdded value: +"Sampling seed; default 42."
- Changed
validate_paper_evidence1 field changed- added
Input schema / properties / record / descriptionAdded value: +"A canli.paper-evidence.v0 record."
- Changed
validate_track_record7 fields changed- added
Input schema / properties / benchmark_sharpe_annualized / descriptionAdded value: +"Annualized Sharpe to beat; default 0." - added
Input schema / properties / confidence / descriptionAdded value: +"Between 0 and 1; default 0.95." - added
Input schema / properties / non_excess_kurtosis / descriptionAdded value: +"Kurtosis, not excess kurtosis; 3 if Normal." - added
Input schema / properties / observations / descriptionAdded value: +"Record length so far, to get its probabilistic Sharpe." - added
Input schema / properties / observed_sharpe_annualized / descriptionAdded value: +"Annualized Sharpe as observed." - added
Input schema / properties / periods_per_year / descriptionAdded value: +"Observations per year: 252 daily, 365 daily crypto, 52 weekly, 12 monthly." - added
Input schema / properties / skew / descriptionAdded value: +"Skewness of the returns; below -0.5 the reading warns that the counts are too generous."
- Added
verify_receipt
9 tool updates
v0.5.0- First observed
company_financial_history - First observed
get_key - First observed
get_receipt - First observed
service_status - First observed
validate_breadth - First observed
validate_deflated_sharpe - First observed
validate_overfitting - First observed
validate_paper_evidence - First observed
validate_track_record
TDQS
Scored across 15 tools
The eight validate_* tools cover genuinely distinct statistical tests, and their descriptions explicitly cross-reference each other ('use validate_overfitting', 'use validate_deflated_sharpe'), which resolves most boundary cases. audit_backtest intentionally overlaps with the single-test validators, though the description tells the agent to prefer it, and company_financial_history is a mild domain outlier. Distinctions are clear enough that misselection is unlikely.
A strong validate_* prefix covers eight tools, and get_key/get_receipt/verify_receipt/audit_backtest follow a consistent verb_noun pattern. The deviations (service_status, company_financial_history) are noun-only but readable and domain-appropriate. No camelCase/snake_case mixing.
15 tools sits at the top of the ideal range but each validator targets a distinct test or lifecycle step, so the set is coherent rather than padded. The composite audit_backtest partially duplicates the individual validators, which is a small redundancy.
The surface covers the full validation lifecycle—key issuance, the family of statistical tests, an aggregate audit, receipt retrieval/verification, status, and SEC financial reference data. A key-revocation tool is implied by quota text but absent, and there is no batch-validate operation, minor gaps an agent can work around.
Maintenance
Related MCP Connectors
SEC filings and financial data for AI agents: 59 tools for statements, valuation and supply chains.
Agent-native SEC filing data: statements assembled, filings read and synthesized. No API key.
Real SEC, 13F, insider, congress & macro data your AI agent can cite. Hosted MCP, 24 tools.
Market regime, execution-cost, bar-QC and backtest-audit tools for agents. Pay per call via x402.
Related MCP Servers
- AlicenseBqualityAmaintenanceVerify a number before an agent asserts it — a Deflated Sharpe Ratio for backtest, plus eval-gap, subset-win, and judge-bias checks, with signed receipts anyone can verify offline.34MIT
- FlicenseNot gradedqualityFmaintenanceProvides financial data and market filings via Streamable HTTP tools. Enables AI agents to resolve equity entities, fetch SEC filings, compare institutional holders, and export finance receipts.-
- AlicenseNot gradedqualityBmaintenanceProvides tools to research crypto trading strategies via backtesting, walk-forward validation, and paper trading, with a deflated-Sharpe overfitting check. Enables natural-language-driven analysis and interpretation of strategy performance.3Apache 2.0
- AlicenseBqualityCmaintenanceEnables auditing and verification of algorithmic trading backtests from coding agents like Claude Code, Cursor, and Windsurf, including look-ahead bias detection, overfitting checks, and sealed audit proof verification.3MIT