Skip to main content
Glama

Server Details

Deflated Sharpe and PBO checks. A verdict is not admission to anything and is not a forecast.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
arhancanli/canlicapital
GitHub Stars
0

TDQS

A3.9/5.0

Scored across 15 tools

Disambiguation4/5

The statistical validators (deflated Sharpe, luck trials, haircut Sharpe, overfitting, reality check) are subtle and could be confused at a glance, but each description explicitly states when to use it and cross-references alternatives. audit_backtest further aggregates several validators, reducing misselection risk. Minor overlap remains, but boundaries are mostly clear.

Naming Consistency4/5

Most tools follow snake_case with clear prefixes: validate_* for statistical validators, get_*/verify_* for keys and receipts, and audit_backtest for orchestration. A few noun-phrase names (company_financial_history, service_status) are minor deviations but remain readable and consistent with the overall style.

Tool Count5/5

Fifteen tools is well-scoped for a validation API with multiple statistical methods, receipt verification, key issuance, and service status. Each validator addresses a distinct statistical question, and the audit tool aggregates them without making the count feel bloated.

Completeness4/5

The tool surface covers validation, receipt retrieval/verification, key issuance, and service status, with strong coverage of relevant statistical tests. However, there is no explicit key revocation or key-listing tool despite quotas mentioning revocation, and receipts appear immutable by design. These are minor gaps agents can likely work around.

Available Tools

15 tools
audit_backtestAudit a backtestAInspect

One-call audit of a strategy's returns: deflated Sharpe, minimum track record and, with every variant's returns, the probability of backtest overfitting, each the matching validator's result with its own receipt. Point returns_file at the backtest's CSV or JSON instead of pasting long series. Prefer it to calling the validators one by one; one validation per check. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.

ParametersJSON Schema
NameRequiredDescriptionDefault
returnsNoPeriodic returns as fractions (0.01 = 1%), oldest first; replaces the Sharpe, observations, skew and kurtosis.
n_splitsNoEven number of blocks, at least 2; default 16.
variantsNoOptional returns of every variant tried (this one included), one row per period, one column per variant; adds the overfitting check.
confidenceNoBetween 0 and 1; default 0.95.
returns_fileNoPath to a CSV or JSON of the returns on this machine (not on the hosted endpoint), instead of returns.
variants_fileNoPath to a CSV or JSON with one numeric column per variant, instead of variants.
returns_columnNoColumn name or 1-based position, when returns_file has several numeric columns.
periods_per_yearYesPeriods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly.
benchmark_sharpe_annualizedNoAnnualized Sharpe to beat; default 0.
effective_independent_trialsYesIndependent variants tried before choosing this one.
cross_trial_sharpe_sd_annualizedYesStandard deviation of annualized Sharpe across those trials.

Output Schema

ParametersJSON Schema
NameRequiredDescription
noteNo
errorNo
checksNo
limitsNo
not_runNo
readingsNo

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false and openWorldHint=true, and the description adds context consistent with that: each check returns the matching validator's result with its own receipt, and only one validation runs per check. It also discloses an interpretation limit ('not admission to anything and is not a forecast'), which is genuine behavioral guidance. It stops short of saying what is written or persisted, hence 4 rather than 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the aggregate purpose before the input and caveat details. It is dense and one clause ('each the matching validator's result with its own receipt') is grammatically awkward, but no sentence is filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and with 11 parameters at full schema coverage the description's job is mostly orientation. It covers the aggregation scope, the file-input alternative, the one-check-per-validation constraint, and the interpretation caveat, leaving only minor gaps such as failure behavior when returns_file is unreadable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds real meaning on top of it: returns_file is 'not on the hosted endpoint,' so the agent learns the path is resolved on the local machine, and it clarifies that supplying variants is what adds the overfitting check. That is useful beyond the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('One-call audit of a strategy's returns') and enumerates exactly what it computes: deflated Sharpe, minimum track record, and, when variants are supplied, the probability of backtest overfitting. This maps directly onto the sibling validators (validate_deflated_sharpe, validate_track_record, validate_overfitting) so an agent can tell it apart from them without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes between alternatives: 'Prefer it to calling the validators one by one; one validation per check' and 'Point returns_file at the backtest's CSV or JSON instead of pasting long series.' Both the aggregate-vs-individual choice and the file-vs-inline choice are named with their selecting conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

company_financial_historyCompany financial history (SEC)A
Read-only
Inspect

SEC-reported financial history for one company in the canlicapital.com reference, by cik or ticker: without a concept, the histories available; with one, observations newest first with accession, form, filed date and unit, plus the source's SHA-256. For point-in-time values use canli-fundamentals-mcp. Public company accounting reference, not market prices, returns, an investment recommendation, or ALPHAC performance. Validate a separately constructed return series with the validation API; accounting values are not returns.

ParametersJSON Schema
NameRequiredDescriptionDefault
cikNoSEC CIK; send cik or ticker.
limitNoMost observations, newest first; default 40.
tickerNoTicker such as AAPL; send ticker or cik.
conceptNous-gaap concept such as Assets; omit to list them.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
sourceNo
companyNo
historyNo
historiesNo
claim_boundaryNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds genuine context beyond that: the response ordering (newest first), returned fields, the source SHA-256 for provenance, and the strong caveat that values are accounting figures rather than returns. It omits auth/rate-limit behavior, so not a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, but the tail is padded with stacked disclaimers ('not market prices, returns, an investment recommendation, or ALPHAC performance' and a separate return-validation sentence). The intended meaning survives, but several clauses could be trimmed without loss.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be enumerated, and the description still covers the conditional response shape, ordering, provenance hash, and scope limits. For a read-only, 4-param tool this is near-complete; only auth/rate-limit context is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds real conditional semantics not in the schema: omitting 'concept' returns the list of available histories, while supplying it returns observations — behavior the schema names alone don't convey.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('SEC-reported financial history for one company ... by cik or ticker') and bounds the domain clearly (accounting reference, not market prices/returns). It does not name a sibling tool directly, but it points at external products and the validation API, so an agent can still separate it from the validate_* family.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives conditional usage ('without a concept, the histories available; with one, observations newest first') and explicit when-not routing ('For point-in-time values use canli-fundamentals-mcp', 'Validate a separately constructed return series with the validation API; accounting values are not returns'). Alternatives are named, but they are mostly outside the sibling set, so it stops short of a full routing rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_keyGet a free validation keyAInspect

Issue a free validation key for this session. Rarely needed: the first validation issues one itself unless CANLI_KEY or local mode is set, and the read tools need none. Quotas: 1000 validations per key per UTC day, 5 keys per client per UTC day, 1048576 bytes per validation request, 1024 bytes per key revocation request, 20000 observations per series, 200 variants per matrix.

ParametersJSON Schema
NameRequiredDescriptionDefault
labelNoName for the key.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
noteNo
errorNo
key_sourceNo
key_presentNo

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only establish that this is an open-world, non-idempotent, non-destructive write. The description adds substantial operational context the agent cannot get elsewhere: per-day key and validation quotas, request and revocation byte limits, and series/matrix cardinality caps. This is exactly the kind of rate-limit and quota disclosure that earns full credit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose and the 'rarely needed' caveat before the quota list, so the most decision-relevant information comes first. The quota sentence is long and includes limits for adjacent endpoints (key revocation, series, matrices) that are not this tool's concern, which is mild scope creep.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is not required, and the description fully covers prerequisites, when to skip the call, and quotas. It never says whether calling it twice produces the same or additional keys, though idempotentHint=false partially covers that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and there is a single optional 'label' parameter already documented in the schema. The description adds no meaning beyond it (no naming conventions, uniqueness, or default behavior), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Issue a free validation key for this session'), which immediately distinguishes it from the validate_* siblings that consume such keys. The word 'free' and 'for this session' scopes the artifact being produced.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says the tool is 'rarely needed' and names the conditions under which it is unnecessary: the first validation self-issues a key unless CANLI_KEY or local mode is set, and read tools need none. That is genuine when-not-to-use guidance, not just a when.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_receiptGet a receiptA
Read-only
Inspect

Fetch a stored verdict by receipt id to re-read it. No key; verify_receipt checks it is genuine. The receipt is content-hashed, reproducible from the open-source core it names, and signed with Ed25519 by a key published at https://canlicapital.com/.well-known/canli-receipt-keys.json.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesReceipt id from a validation result.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
errorNo
limitsNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this readOnly, non-destructive and openWorld, so the safety profile is covered. The description adds genuinely useful behavior beyond that: no key required, the receipt is content-hashed and reproducible from the open-source core, and it is Ed25519-signed by a key at a published well-known URL. Missing only error/absence behavior for an unknown id.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly packed sentences, front-loaded with the core action and immediately followed by the sibling disambiguation and trust model. No filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and annotations carry the safety profile. For a single-param read tool, the description supplies the routing, auth, and verification-model context an agent needs to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'id' param is fully documented in the schema (pattern plus provenance note). The description adds only the phrase 'by receipt id', so it neither compensates nor detracts; baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Fetch) and resource (a stored verdict by receipt id) plus the intent (to re-read it), and explicitly differentiates itself from the sibling verify_receipt. An agent can distinguish it from the validate_* and verify_* tools without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It routes the agent: use this to re-read a stored receipt, use verify_receipt to check genuineness. It also states 'No key', clarifying the auth posture for invocation. It stops short of enumerating when-not-to-use (e.g., what to do if the id is unknown or the receipt is missing).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

service_statusService statusA
Read-only
Inspect

Whether the validation API is up, with its quotas; check after a timeout before resubmitting. No key. This verdict is about the series exactly as submitted. The service never saw the data source, its costs, survivorship, or any lookahead in how the series was built.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
errorNo
limitsNo

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so safety is covered. The description adds genuinely new context: 'No key' (no authentication required) and that quotas are returned, plus a caveat about the scope of the verdict.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose and usage trigger are front-loaded well, but one of the three sentences is off-topic for a status endpoint and does not earn its place, diluting an otherwise tight definition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and no input params need documenting. The remaining gap is that the trailing 'series as submitted' caveat is irrelevant to a service-status call and could mislead an agent about what this tool reports.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Zero parameters, so the baseline of 4 applies; there are no inputs whose semantics need explaining.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening clause ('Whether the validation API is up, with its quotas') states a specific resource and check, which distinguishes it from the validate_* siblings. However, the trailing sentences about 'the series exactly as submitted' and what 'the service never saw' belong to a validation-verdict tool, not a health check, and muddy what the tool actually is.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'check after a timeout before resubmitting' gives an explicit trigger condition for calling it, which is exactly the kind of context an agent needs. It does not name sibling alternatives, but the trigger is clear enough to route correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_backtest_lengthMinimum backtest lengthAInspect

Minimum backtest length (years) before the best of N independent trials is not expected to reach a target Sharpe by luck; with backtest_years, the most trials those years allow. For planning a search; once it has a result, use validate_deflated_sharpe. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.

ParametersJSON Schema
NameRequiredDescriptionDefault
backtest_yearsNoBacktest length in years, for the most trials it allows.
target_sharpe_annualizedNoIn-sample annualized Sharpe you would call a discovery; default 1.
effective_independent_trialsNoIndependent trials tried (backtests, parameter sets, ideas).

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
noteNo
errorNo
limitsNo
receiptNo
computedNo

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, openWorldHint=true, destructiveHint=false, which is an odd profile for what the description frames as a pure planning calculation; the description does not reconcile that (no side effects, no persisted state, no external calls disclosed). It does add genuine interpretive context that outputs are not admissions or forecasts, which goes beyond the annotations. With annotations already present, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and the sibling handoff are front-loaded, which is good, but the opening is a single long clause-stacked sentence that is hard to parse on first read. The closing sentence about thresholds not being admission is useful but borders on editorial and is not clearly tied to what this tool returns.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, and the description covers purpose, planning-time usage, and the handoff tool. For a stateless computation with fully documented parameters, that is close to sufficient; only the reconciliation of the readOnlyHint=false annotation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and all three parameters are documented in the schema with bounds and defaults. The description restates backtest_years ('with backtest_years, the most trials those years allow') without adding syntax or format meaning beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific quantity (minimum backtest length in years) and the condition under which it holds (best of N independent trials not reaching target Sharpe by luck), which is far more specific than the title alone. It is somewhat syntactically dense and reads like a textbook definition rather than a plain 'what this does', and it does not directly contrast itself with siblings like validate_luck_trials.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the setting explicitly ('For planning a search') and names the alternative with the condition that selects it ('once it has a result, use validate_deflated_sharpe'). It also adds a when-not caveat: a deflated Sharpe or overfitting probability past any threshold is not admission to anything, so it should not be used as a decision gate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_breadthValidate breadth ceilingAInspect

Book Sharpe ceiling from adding sleeves of this quality and correlation, the Sharpe at a sleeve count, and the sleeves a target needs. For portfolio construction; it validates no single strategy. This verdict is about the series exactly as submitted. The service never saw the data source, its costs, survivorship, or any lookahead in how the series was built.

ParametersJSON Schema
NameRequiredDescriptionDefault
targetNoTarget book Sharpe, for the sleeves it needs.
sleevesNoSleeve count, for that book's Sharpe.
sleeve_sharpeYesAnnualized Sharpe of one sleeve.
average_pairwise_correlationYesAverage correlation between sleeves, -1 to 1.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
noteNo
errorNo
limitsNo
receiptNo
computedNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses meaningful scope limits: the verdict applies only to the series as submitted and the service never inspected data source, costs, survivorship, or lookahead. That tells the agent the result is a purely mathematical check. The readOnlyHint=false/openWorldHint=true annotations are neither repeated nor contradicted; the description adds real context rather than echoing them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is reasonably short (four sentences) and front-loads the outputs, but the first sentence is dense and awkwardly constructed, forcing re-reading to parse the three-output mapping. The disclaimer sentences earn their place, but the core sentence could be clearer.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a non-trivial portfolio-math tool, the description covers purpose, applicability, the three outputs' drivers, and the critical caveat about what data was never inspected. An output schema exists, so return-value explanation is not required. The only gap is the absence of explicit guidance on when to prefer this over a sibling validator.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds the parameter-to-output relationship: sleeve_sharpe plus correlation determine the ceiling, sleeve count determines that book's Sharpe, and target determines the sleeve count needed. That pairing is not obvious from the schema alone, so it earns above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the three concrete outputs (book Sharpe ceiling, Sharpe at a sleeve count, sleeves required for a target), which is a specific computational purpose rather than a restatement of the name. It also differentiates itself from the sibling single-strategy validators with 'For portfolio construction; it validates no single strategy.' The opening sentence is grammatically elliptical ('Book Sharpe ceiling from adding sleeves...') but the intent is recoverable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives one usage signal — 'For portfolio construction; it validates no single strategy' — which implicitly routes the agent away from the sleeve-less validate_* siblings (validate_deflated_sharpe, validate_haircut_sharpe, etc.). However, it never names an alternative or states a when-not condition, so the selection guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_deflated_sharpeValidate deflated SharpeAInspect

Deflated Sharpe ratio: the probability (0 to 1) that the selected strategy's Sharpe beats the best that luck gives across the variants tried, with the probabilistic Sharpe and that luck benchmark. Send the seven statistics or a return series. With every variant's returns use validate_overfitting; luck as a trial count, validate_luck_trials; a multiple-testing haircut, validate_haircut_sharpe. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.

ParametersJSON Schema
NameRequiredDescriptionDefault
skewNoSkewness of returns; 0 if Normal.
returnsNoPeriodic returns as fractions (0.01 = 1%), oldest first; replaces the Sharpe, observations, skew and kurtosis.
observationsNoNumber of return observations.
periods_per_yearNoPeriods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly.
non_excess_kurtosisNoKurtosis, not excess kurtosis; 3 if Normal.
observed_sharpe_annualizedNoAnnualized Sharpe as observed.
effective_independent_trialsNoIndependent variants tried before choosing this one.
cross_trial_sharpe_sd_annualizedNoStandard deviation of annualized Sharpe across those trials.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
noteNo
errorNo
limitsNo
receiptNo
computedNo

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=false, openWorldHint=true and destructiveHint=false, so safety signalling is already covered. The description still adds real context beyond that: it explains the two valid input modes (statistics vs. return series) and warns that the resulting probability is not admission to anything and not a forecast. It does not mention auth, cost, or rate limits, but for a local statistical computation that gap is small.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The whole definition fits in three dense sentences with no filler; the routing alternatives are front-loaded mid-paragraph and the interpretive caveat closes it. It is slightly run-on and could be split, but every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation. For a tool with 8 optional parameters and no required inputs, the description covers purpose, the two input modes, sibling alternatives, and interpretation limits. It could go further on which combination of statistics is minimally sufficient, but nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so baseline would be 3, but the description adds the either/or input mode ('the seven statistics or a return series') that the schema only implies via the 'replaces' note on returns. This clarifies mutual exclusivity across the 8 all-optional parameters, which is meaningful routing information beyond the field-level descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states precisely what the tool computes: the probability (0-1) that the selected strategy's Sharpe beats the best luck produces across the variants tried, plus the probabilistic Sharpe and luck benchmark. It also names the sibling tools it is not (validate_overfitting, validate_luck_trials, validate_haircut_sharpe), so the agent can separate it from adjacent validators. It falls short of 5 only because it opens with a concept definition rather than a direct verb+resource statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit routing: send the seven statistics or a return series; use validate_overfitting when every variant's returns are available; validate_luck_trials when luck is expressed as a trial count; validate_haircut_sharpe for a multiple-testing haircut. The alternative tools and the conditions selecting them are spelled out rather than left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_haircut_sharpeHaircut Sharpe ratioAInspect

Haircut Sharpe for multiple testing (Harvey and Liu 2015): the Sharpe a single test would have needed, by Bonferroni and independent tests, and with the other tests' Sharpes, Holm and BHY. For the probability the Sharpe is real, use validate_deflated_sharpe. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.

ParametersJSON Schema
NameRequiredDescriptionDefault
testsNoTests run, this one included; gives the Bonferroni and independent-test haircuts.
observationsYesReturn observations behind the Sharpe.
autocorrelationNoLag-1 autocorrelation of returns, -1 to 1; default 0. Corrects the annualized Sharpe (Lo 2002).
periods_per_yearYesPeriods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly.
observed_sharpe_annualizedYesAnnualized Sharpe as observed.
other_sharpe_ratios_annualizedNoAnnualized Sharpes of the other tests over the same observations; adds Holm and BHY.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
noteNo
errorNo
limitsNo
receiptNo
computedNo

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (destructiveHint=false, openWorldHint=true), so the description's burden is lighter. It adds interpretive behavior ('not admission to anything and is not a forecast') and clarifies which inputs drive which outputs. It does not state that the call is a pure stateless computation or whether anything is persisted, which matters given readOnlyHint=false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the method and citation, no filler. The second sentence packs two distinct ideas (sibling routing and threshold caveat) and is grammatically dense, but every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained; the description still covers method provenance, the routing alternative, and a misuse caveat. It leaves open how this differs operationally from related siblings like validate_luck_trials and validate_overfitting, which is the main remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all six parameters including worked examples (252 daily, 365 crypto). The description only loosely links inputs to outputs ('by Bonferroni and independent tests', 'with the other tests' Sharpes, Holm and BHY'), which is baseline-level added value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific computation (haircut Sharpe for multiple testing, Harvey and Liu 2015) and enumerates exactly what it returns: Bonferroni/independent-test haircuts and Holm/BHY adjustments. It also names the sibling it is not (validate_deflated_sharpe), so an agent can route without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives one explicit routing rule: 'For the probability the Sharpe is real, use validate_deflated_sharpe.' That is a real alternative-plus-condition. It does not cover when to prefer this over other siblings such as validate_overfitting or validate_luck_trials, so it stops short of full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_luck_trialsLuck-equivalent trialsAInspect

How many skill-less strategies a search would need for its best to reach this Sharpe by luck (Monte Carlo), and with a trial count, the chance it did. States luck as the best of N random tries; for the probability the Sharpe is real, use validate_deflated_sharpe. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.

ParametersJSON Schema
NameRequiredDescriptionDefault
skewNoSkewness; below -0.5 the reading warns the counts are too generous.
observationsYesReturn observations behind the Sharpe.
autocorrelationNoLag-1 autocorrelation of returns, -1 to 1; default 0. Corrects the annualized Sharpe (Lo 2002).
periods_per_yearYesPeriods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly.
observed_sharpe_annualizedYesAnnualized Sharpe as observed.
effective_independent_trialsNoIndependent trials tried; adds the chance the best reached this Sharpe by luck.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
noteNo
errorNo
limitsNo
receiptNo
computedNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations supply the safety profile (destructiveHint=false, openWorldHint=true), and the description adds genuine extra context: the method is Monte Carlo and luck is defined as the best of N random tries. It also warns that a threshold crossing is neither admission nor forecast, which is useful interpretive guidance. It says nothing about why readOnlyHint is false for what reads as a pure computation, leaving an unexplained annotation/label mismatch.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core computation before the sibling pointer and the caveat. Every sentence carries content, though the opening sentence is syntactically heavy and could be tightened.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and annotations cover the safety profile. Purpose, sibling routing, and interpretive limits are all present; the main residual gap is the unaddressed readOnlyHint=false for a statistical computation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameter meanings, defaults (autocorrelation default 0), and bounds are already documented; baseline is 3. The description only implies effective_independent_trials by calling it 'a trial count' and does not add syntax or semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific computation: how many skill-less strategies a search would need for its best to reach the observed Sharpe by luck, and (given a trial count) the probability of that. It names the sibling it is not (validate_deflated_sharpe), so the agent can separate the two. The phrasing is dense and inverted, which costs it a point but the substance is specific and non-tautological.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly routes the agent: 'for the probability the Sharpe is real, use validate_deflated_sharpe', and the clause 'with a trial count, the chance it did' signals when to supply effective_independent_trials. The closing caveat about thresholds also frames how to read the result. There is no explicit when-not, but the alternative pointer is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_overfittingValidate overfitting (CSCV)AInspect

Probability of backtest overfitting (0 to 1) by CSCV: how often the in-sample best variant falls below the out-of-sample median. Needs every variant's returns (periods by variants); with summary statistics only, use validate_deflated_sharpe. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.

ParametersJSON Schema
NameRequiredDescriptionDefault
seedNoSampling seed; default 42.
matrixYesReturns as fractions, one row per period, one column per variant.
n_splitsNoEven number of blocks, at least 2; default 16.
max_combinationsNoMost splits evaluated, up to 2000 (default).

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
noteNo
errorNo
limitsNo
receiptNo
computedNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnly=false, openWorld=true, destructive=false, which says little about a statistical computation. The description compensates with the required input shape and a meaningful interpretive caveat (threshold crossings are neither admission nor forecast), which is real behavioral guidance beyond structured fields. It stops short of runtime traits like cost, determinism, or failure modes when the matrix is malformed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the metric and method, then requirements, then sibling routing, then the interpretation caveat. Three dense sentences with no filler, each carrying distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers inputs, the alternative tool, and result interpretation for a non-trivial statistical procedure. It could say more about assumptions (e.g. minimum period/variant counts) or determinism, but nothing essential to correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents matrix, seed, n_splits, and max_combinations. The description's 'periods by variants' restates the matrix schema description rather than adding syntax or constraints. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific outcome (probability of backtest overfitting, 0-1) and the method (CSCV), then defines the metric operationally as how often the in-sample best variant falls below the out-of-sample median. It explicitly distinguishes itself from validate_deflated_sharpe, so an agent can route without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States the input condition that selects this tool (every variant's returns required) and names the alternative for the other case (summary statistics only -> validate_deflated_sharpe). This is an explicit when/alternatives pairing, not implied usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_paper_evidenceValidate paper evidenceAInspect

Whether a paper or simulated performance record meets canli.paper-evidence.v0, with a JSON pointer per failure. Checks structure and required disclosures, not whether the returns are good. This verdict is about the series exactly as submitted. The service never saw the data source, its costs, survivorship, or any lookahead in how the series was built.

ParametersJSON Schema
NameRequiredDescriptionDefault
recordYesA canli.paper-evidence.v0 record.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
noteNo
errorNo
limitsNo
receiptNo
computedNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false, openWorldHint=true, destructiveHint=false, so the description usefully adds that the verdict concerns the submitted series only and that the service never saw the data source, costs, survivorship, or lookahead. This is exactly the kind of limitation an agent needs. It stops short of stating permission requirements or whether calling it has side effects (notably, readOnlyHint=false is left unexplained).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and output, and each sentence is short. The final two sentences overlap somewhat, both reassuring about what the verdict does not cover, which is mild redundancy rather than waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description correctly delimits what the verdict does and does not assert. Given the single nested object parameter, this is nearly complete; only the missing sibling routing keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One parameter with 100% schema description coverage, so the schema already carries the semantics. The description's mention of canli.paper-evidence.v0 simply restates the schema's own description and adds no new syntax or format detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource (validates a paper/simulated performance record against canli.paper-evidence.v0) and states the output shape (JSON pointer per failure). It does not, however, differentiate itself from the many sibling validate_* tools, leaving the agent to infer which validator applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clarifies scope by negation ('checks structure and required disclosures, not whether the returns are good'), which implies when the tool is appropriate. But it never names an alternative or states a positive trigger condition versus siblings like validate_track_record or validate_backtest.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_reality_checkData-snooping tests (SPA, Reality Check, StepM)AInspect

Data-snooping tests on every variant a search tried: Hansen's SPA p-value that the best beat the benchmark only by luck, White's Reality Check, and the variants Romano-Wolf StepM finds better. Send all variants tried, not only the winners. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.

ParametersJSON Schema
NameRequiredDescriptionDefault
repsNoBootstrap draws; default 2000.
seedNoSampling seed; default 42.
alphaNoFamilywise error for StepM; default 0.05.
matrixNoReturns of every variant the search tried, one row per period, one column per variant.
benchmarkNoBenchmark return per period; default zero.
matrix_fileNoPath to a CSV or JSON with one numeric column per variant, instead of matrix.
block_lengthNoMean bootstrap block in periods; default round(n^(1/3)).

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
noteNo
errorNo
limitsNo
receiptNo
computedNo

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already carry the safety profile (destructiveHint=false, openWorldHint=true), and the description adds a genuine interpretive caveat beyond them: a deflated Sharpe or overfitting probability above/below a threshold "is not admission to anything and is not a forecast." That tells the agent the output is diagnostic, not a decision — useful context the annotations do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with what the tool computes and then the input rule. Dense but each clause earns its place; minor run-on structure in the second sentence keeps it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained; all seven parameters are fully described in the schema, and the description supplies the key input expectation (include every variant tried). For a stateless computation tool this is essentially complete, missing only explicit sibling routing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all seven parameters including defaults and bounds, and the baseline is 3. The description only reinforces the matrix requirement ("every variant a search tried"), which the schema already states, so it adds little beyond the structured fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific resource (data-snooping tests) plus the three concrete tests it runs (SPA p-value, White's Reality Check, Romano-Wolf StepM), so the agent knows exactly what computation happens. It only implicitly separates itself from siblings like validate_deflated_sharpe and validate_overfitting via the closing caveat rather than stating routing outright.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Send all variants tried, not only the winners" is an explicit, actionable usage directive that tells the agent what input to gather, which most siblings don't provide. It stops short of naming when to pick this over validate_deflated_sharpe or validate_overfitting, so it is clear context rather than full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

validate_track_recordMinimum track record lengthAInspect

Minimum track record length (observations and years) for an observed Sharpe to beat a benchmark at a confidence level; with observations, the record's probabilistic Sharpe so far. For live or paper records; to size a backtest for its trials, use validate_backtest_length. A deflated Sharpe or overfitting probability above or below any threshold is not admission to anything and is not a forecast.

ParametersJSON Schema
NameRequiredDescriptionDefault
skewYesSkewness of returns; 0 if Normal.
confidenceNoBetween 0 and 1; default 0.95.
observationsNoRecord length so far, for its probabilistic Sharpe.
periods_per_yearYesPeriods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly.
non_excess_kurtosisYesKurtosis, not excess kurtosis; 3 if Normal.
observed_sharpe_annualizedYesAnnualized Sharpe as observed.
benchmark_sharpe_annualizedNoAnnualized Sharpe to beat; default 0.

Output Schema

ParametersJSON Schema
NameRequiredDescription
dataNo
noteNo
errorNo
limitsNo
receiptNo
computedNo

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (readOnlyHint=false, destructiveHint=false, openWorldHint=true) are somewhat at odds with an obvious pure-calculation tool, and the description neither confirms nor contradicts them. It does add interpretive context ('not admission to anything and is not a forecast') and discloses that output includes observations/years, but says nothing about precision, numerical caveats, or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core computation and then the sibling routing. The final disclaimer sentence is useful but slightly tangential and the first sentence is dense with nested clauses, costing a little clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and full schema coverage, the description does not need to explain return values, and it covers purpose, scope, and routing. It is complete enough to invoke correctly, though it never explains what the periodic/year conversion or confidence inputs imply about results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already defines every parameter, including defaults (benchmark 0, confidence 0.95). The description names 'observations', 'confidence level', and benchmark only in passing and adds no format or edge-case semantics beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific computation (minimum track record length, in observations and years) on a named resource (an observed Sharpe vs a benchmark) and explicitly names the sibling it is not: validate_backtest_length. An agent can distinguish it from the several validate_* siblings without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when ('For live or paper records') and an explicit alternative with its selecting condition ('to size a backtest for its trials, use validate_backtest_length'). It also warns what the tool is not for, which is exactly the routing guidance an agent needs among near-identical validate_* siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_receiptVerify a receiptA
Read-only
Inspect

Verify a receipt offline: its Ed25519 signature against the bundled canlicapital.com key, its output hash and its id. Send an id to fetch it first, or the receipt itself. The receipt is content-hashed, reproducible from the open-source core it names, and signed with Ed25519 by a key published at https://canlicapital.com/.well-known/canli-receipt-keys.json.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoReceipt id from a validation result.
receiptNoA receipt as get_receipt returns it, to verify without fetching.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorNo
validNo
checksNo
key_idNo
meaningNo
receipt_idNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=true and openWorldHint=true, and the description goes beyond them by disclosing that verification happens offline, that the trust anchor is a specific bundled key, and that the key is published at a well-known URL. It stops short of describing failure modes or what makes verification fail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense, front-loaded sentences that lead with the action and the checks performed; the reproducibility and key-publication details are useful but slightly burden the second sentence. Nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no explanation, and the annotations carry the safety profile. The description covers the cryptographic scope and both input paths, leaving only edge-case behavior (e.g., malformed receipts, network failure when fetching by id) unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, and the description adds genuine semantics: passing an id triggers an internal fetch while passing a receipt object verifies without fetching. That behavioral distinction between the two optional parameters is not evident from the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (verify) and resource (receipt) and enumerates exactly what is checked: the Ed25519 signature, the output hash, and the id. The phrase 'Send an id to fetch it first' implicitly differentiates it from the sibling get_receipt, which only retrieves.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly spells out the two mutually exclusive input modes ('Send an id to fetch it first, or the receipt itself'), which is real usage guidance for an agent choosing how to call it. It does not, however, state when verification is unnecessary or how it relates to validate_* siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool update
    • Addedvalidate_reality_check
  2. 9 tool updates
    • Changedaudit_backtest7 fields changed
      • changedInput schema / properties / cross_trial_sharpe_sd_annualized / description
        Previous value: -"Standard deviation of the annualized Sharpe across those trials."New value: +"Standard deviation of annualized Sharpe across those trials."
      • changedInput schema / properties / periods_per_year / description
        Previous value: -"Observations per year: 252 daily, 365 daily crypto, 52 weekly, 12 monthly."New value: +"Periods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly."
      • changedInput schema / properties / returns / description
        Previous value: -"Periodic returns as fractions (0.01 is 1%), oldest first; replaces the Sharpe, observations, skew and kurtosis fields."New value: +"Periodic returns as fractions (0.01 = 1%), oldest first; replaces the Sharpe, observations, skew and kurtosis."
      • changedInput schema / properties / returns_column / description
        Previous value: -"Header name or 1-based position of the returns column when returns_file has several numeric columns."New value: +"Column name or 1-based position, when returns_file has several numeric columns."
      • changedInput schema / properties / returns_file / description
        Previous value: -"Path to a CSV or JSON file of the returns on the machine running this server, instead of returns. Not available on the hosted endpoint."New value: +"Path to a CSV or JSON of the returns on this machine (not on the hosted endpoint), instead of returns."
      • changedInput schema / properties / variants / description
        Previous value: -"Optional returns of every variant tried, this one included, as fractions: one row per period, one column per variant. Adds the overfitting check."New value: +"Optional returns of every variant tried (this one included), one row per period, one column per variant; adds the overfitting check."
      • changedInput schema / properties / variants_file / description
        Previous value: -"Path to a CSV or JSON file of every variant's returns (one numeric column per variant), instead of variants."New value: +"Path to a CSV or JSON with one numeric column per variant, instead of variants."
    • Changedvalidate_backtest_length3 fields changed
      • changedInput schema / properties / backtest_years / description
        Previous value: -"Length of the backtest in years, to get the most independent trials it allows."New value: +"Backtest length in years, for the most trials it allows."
      • changedInput schema / properties / effective_independent_trials / description
        Previous value: -"Independent trials (backtests, parameter sets, ideas) tried; gives the minimum backtest length."New value: +"Independent trials tried (backtests, parameter sets, ideas)."
      • changedInput schema / properties / target_sharpe_annualized / description
        Previous value: -"In-sample annualized Sharpe you would take as a discovery; default 1."New value: +"In-sample annualized Sharpe you would call a discovery; default 1."
    • Changedvalidate_breadth2 fields changed
      • changedInput schema / properties / sleeves / description
        Previous value: -"Sleeve count, to get that book's Sharpe."New value: +"Sleeve count, for that book's Sharpe."
      • changedInput schema / properties / target / description
        Previous value: -"Target book Sharpe, to get the sleeves it needs."New value: +"Target book Sharpe, for the sleeves it needs."
    • Changedvalidate_deflated_sharpe3 fields changed
      • changedInput schema / properties / cross_trial_sharpe_sd_annualized / description
        Previous value: -"Standard deviation of the annualized Sharpe across those trials."New value: +"Standard deviation of annualized Sharpe across those trials."
      • changedInput schema / properties / periods_per_year / description
        Previous value: -"Observations per year: 252 daily, 365 daily crypto, 52 weekly, 12 monthly."New value: +"Periods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly."
      • changedInput schema / properties / returns / description
        Previous value: -"Periodic returns as fractions (0.01 is 1%), oldest first; replaces the Sharpe, observations, skew and kurtosis fields."New value: +"Periodic returns as fractions (0.01 = 1%), oldest first; replaces the Sharpe, observations, skew and kurtosis."
    • Changedvalidate_haircut_sharpe5 fields changed
      • changedInput schema / properties / autocorrelation / description
        Previous value: -"First-order autocorrelation of the returns, -1 to 1; default 0. Corrects the annualized Sharpe as Lo (2002)."New value: +"Lag-1 autocorrelation of returns, -1 to 1; default 0. Corrects the annualized Sharpe (Lo 2002)."
      • changedInput schema / properties / observations / description
        Previous value: -"Number of return observations behind the Sharpe ratio."New value: +"Return observations behind the Sharpe."
      • changedInput schema / properties / other_sharpe_ratios_annualized / description
        Previous value: -"Annualized Sharpe ratios of the other tests, over the same observations; adds the Holm and BHY haircuts."New value: +"Annualized Sharpes of the other tests over the same observations; adds Holm and BHY."
      • changedInput schema / properties / periods_per_year / description
        Previous value: -"Observations per year: 252 daily, 365 daily crypto, 52 weekly, 12 monthly."New value: +"Periods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly."
      • changedInput schema / properties / tests / description
        Previous value: -"Total tests run, this one included; gives the Bonferroni and independent-test haircuts."New value: +"Tests run, this one included; gives the Bonferroni and independent-test haircuts."
    • Changedvalidate_luck_trials5 fields changed
      • changedInput schema / properties / autocorrelation / description
        Previous value: -"First-order autocorrelation of the returns, -1 to 1; default 0. Corrects the annualized Sharpe as Lo (2002)."New value: +"Lag-1 autocorrelation of returns, -1 to 1; default 0. Corrects the annualized Sharpe (Lo 2002)."
      • changedInput schema / properties / effective_independent_trials / description
        Previous value: -"Independent trials tried; adds the chance that the best of them reached this Sharpe by luck."New value: +"Independent trials tried; adds the chance the best reached this Sharpe by luck."
      • changedInput schema / properties / observations / description
        Previous value: -"Number of return observations behind the Sharpe ratio."New value: +"Return observations behind the Sharpe."
      • changedInput schema / properties / periods_per_year / description
        Previous value: -"Observations per year: 252 daily, 365 daily crypto, 52 weekly, 12 monthly."New value: +"Periods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly."
      • changedInput schema / properties / skew / description
        Previous value: -"Skewness of the returns; below -0.5 the reading warns that the counts are too generous."New value: +"Skewness; below -0.5 the reading warns the counts are too generous."
    • Changedvalidate_overfitting1 field changed
      • changedInput schema / properties / matrix / description
        Previous value: -"Returns as fractions: one row per period, one column per variant."New value: +"Returns as fractions, one row per period, one column per variant."
    • Changedvalidate_track_record2 fields changed
      • changedInput schema / properties / observations / description
        Previous value: -"Record length so far, to get its probabilistic Sharpe."New value: +"Record length so far, for its probabilistic Sharpe."
      • changedInput schema / properties / periods_per_year / description
        Previous value: -"Observations per year: 252 daily, 365 daily crypto, 52 weekly, 12 monthly."New value: +"Periods per year: 252 daily, 365 crypto, 52 weekly, 12 monthly."
    • Changedverify_receipt1 field changed
      • changedInput schema / properties / receipt / description
        Previous value: -"A receipt as get_receipt returns it (its data), to verify without fetching it."New value: +"A receipt as get_receipt returns it, to verify without fetching."
  3. 14 tool updates
    • First observedaudit_backtest
    • First observedcompany_financial_history
    • First observedget_key
    • First observedget_receipt
    • First observedservice_status
    • First observedvalidate_backtest_length
    • First observedvalidate_breadth
    • First observedvalidate_deflated_sharpe
    • First observedvalidate_haircut_sharpe
    • First observedvalidate_luck_trials
    • First observedvalidate_overfitting
    • First observedvalidate_paper_evidence
    • First observedvalidate_track_record
    • First observedverify_receipt

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    A
    maintenance
    Verify a number before an agent asserts it — a Deflated Sharpe Ratio for backtest, plus eval-gap, subset-win, and judge-bias checks, with signed receipts anyone can verify offline.
    34
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Most trading signals are noise. AlphaAssay puts them on trial — deflated Sharpe, out-of-sample, leakage forensics — and returns signed pass/fail verdicts anyone can verify. Methodology audits, not investment advice.
    17
    Apache 2.0
  • A
    license
    A
    quality
    A
    maintenance
    Market-intelligence MCP: 18 detection engines over 9,200+ instruments with calibrated uncertainty and outcome-verified provenance. Informational only, not financial advice.
    30
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Checks whether a trading backtest survives its own statistics: deflated Sharpe, multiple-testing correction against a best-of-N-noise benchmark, minimum track record length, and fill realism. Takes no market data and no API keys, and cannot recommend a trade — it only reports that a result is weaker than claimed or not yet provable.
    MIT
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.