P2predict-mcp
The right price is already in your data.
P2Predict turns your purchasing history into a price model your team can talk to. Ask what a part should cost, why, and how sure the answer is, in plain language, through the AI agent you already use.

Set up in one command:
pip install "p2predict[mcp]"Built for procurement and engineering teams in automotive, semiconductors, electronics, industrials, pharma, and chemicals.
The problem
Most of what a part costs is decided upstream, in a design review procurement never sees. An engineer tightens a tolerance beyond what the application needs, or locks the part to a single supplier, and the cost rides downstream with no number attached. By the time the BOM reaches procurement, the expensive decisions are frozen and all that is left to negotiate is the rounding.
Then the quote arrives. The supplier knows their cost to the cent; you have last year's PO and a few days to respond. Multiply that across every line you buy and the leaks look the same everywhere: premiums no one benchmarked, specs no one costed, a number you cannot defend when finance asks where the savings went.
The answers are already in your purchase history. Nobody has had the time to dig them out.
Related MCP server: AI Model Advisor MCP Server
What P2Predict does
P2Predict learns from your own purchasing data what drives the price of a part: supplier, material, size, spec, region. Then it gives your team a defensible target for any new or proposed part. Your category managers ask in plain English. The answer comes back grounded in what you have actually paid.
It runs on your machine, on your data, through any AI agent your team already uses. Nothing is uploaded. No vendor catalog, no cloud, no per-seat data-sharing.
What it is, and what it isn't
P2Predict is parametric price prediction. It learns the fundamental pricing structure in your historical buying data and benchmarks any part against it, the ones you are about to buy and the ones you already do. Being precise about that is the whole point, so here is the honest scope.
What it does
Learns from the prices you have actually paid and predicts what a similar part should cost.
Attributes that predicted price to each spec and to the supplier, so you can see what is moving the number.
Puts a calibrated likely-range on every estimate and flags where the data is too thin to trust.
Improves as you give it more of your own purchase history.
What it does not do
It is not a should-cost tool. It does not build a part up from raw material, labor, and machine time, and it cannot tell you a supplier's true cost or margin.
It only knows what your data has shown it. Ask about a part unlike anything in your history and it will widen the range or tell you to get a quote rather than guess.
The per-spec breakdown shows what is associated with price in your data. It is a read on your market, not a causal or engineering model of why a part costs what it does.
It does not invent data. No relevant history, no model.
The conversations it changes
This is where P2Predict earns its keep. Every one of these is a real question your team can now answer in seconds, with a number and the confidence behind it.
"Is this quote fair?"
Your category manager drops the quote on the agent.
Category manager: "Supplier quotes $14.20 for this part. What should it cost?"
P2Predict: "$12.40. Nine times in ten the real price lands between $10.80 and $13.90. This quote is running about 15% high."
Your category manager now knows exactly where they can push back, with real data behind it.
When the supplier pushes back: "Are you sure that's the right price?"
This is where most negotiations stall. Now you have an answer. Ask for the breakdown.
Category manager: "Why $12.40? Break it down."
P2Predict: "Supplier choice +$0.85, rush delivery +$1.20, tighter tolerance +$0.42, size +$0.40."
Now you argue the components line by line: "We agreed standard lead time. Take the $1.20 rush charge off and we're aligned." A line item is hard to wave away.
"What if we switch supplier?"
Hold the spec fixed, swap the supplier, read the delta.
Category manager: "What happens if we move this 16-cell pack monitor from Supplier A to Supplier B?"
P2Predict: "Down 37.7%, about $2.07 a unit, with the per-feature breakdown to back it up."
More targeted RFQs, and a faster sourcing decision. That number is your lever in the room.
In the design review: "Is this feature worth it?"
Engineering proposes a tighter tolerance. Before it gets locked in, price it.
Engineer: "We want to go from ±0.1mm to ±0.05mm."
Category manager to the agent: "What does that do to cost?"
P2Predict: "+$0.42 a unit, +18%, likely range $0.30 to $0.55."
Now the conversation is "is 18% worth this requirement?", a priced trade-off the room can settle on numbers.
In the cost-down workshop: "What is the design paying for that it doesn't need?"
Walk in with every spec priced. Which features carry real cost, which premiums are negotiable, where the design is paying for something the application never uses. Backed by your own data, with a confidence level on every finding.
RFQ triage: "Which of these 200 lines deserve a call?"
Drop the whole RFQ on the agent. Every line gets a target and a range. The eight to fifteen lines that fall outside their range are the ones worth a phone call. The rest are routine. Your team spends the afternoon on what actually moves the number.
It tells you how much to trust the number
Most tools hand you a number and walk away. P2Predict hands you the number and tells you how confident to be in it, per part, in dollars. That honesty is the whole point: a confident-but-wrong benchmark loses you credibility the moment a supplier checks it.
Three real parts, three different confidence ranges. The model is tight on the part it knows well and openly uncertain on the ones it doesn't. A narrow range means negotiate hard. A wide one means get a quote first. Your category manager always knows which.
A confidence range on every estimate. "$12.40, and nine times in ten the real price lands between $10.80 and $13.90."
An honest map of where the model is strong and where it is thin. P2Predict flags which parts of your category it can benchmark with confidence and which need a real quote, so nobody negotiates off a number the data can't support.
A reason for every number. Every estimate breaks down into what each spec and the supplier contribute, so you argue the components line by line.
See what actually drives the price
Point P2Predict at a category and it shows you the levers. These charts come straight out of the Battery Management ICs case study, built on public catalog data anyone can reproduce.
Supplier choice is the biggest lever on the board. Same single-cell chip, identical spec, sorted by who makes it:
The premium supplier is priced at roughly four times the value option for the same part. That is a number you take into a negotiation, backed by your own data.
Every estimate breaks down spec by spec. Ask why a part is priced the way it is and you get the full breakdown:

Package size, supplier premium, multi-cell architecture: each one in dollars, each one adding up exactly to the predicted price. This is what lets your category manager say "I know what I'm paying for, and here's the line I want to cut."
Complexity is priced, not assumed. Every package pin on this same chip adds cost, monotonically, from $2.31 at 8 pins to $4.88 at 48 pins:
And on the 30 parts held back from training, entirely unseen by the model, predictions land within roughly 16% of the actual price half the time, and within 73% nine times in ten. That's the model showing its work on parts it never saw, not asking you to trust it blind.
How it fits your stack
You don't use P2Predict; your agent does. It speaks to any AI agent through a standard connector — Claude, GPT, or a local model — so your category managers never learn a new tool. They ask the assistant they already use, and it runs the analysis for them. This is agentic-first: there is no dashboard and no app, the interface is the agent you already have.
Under the hood, P2Predict is two layers. A math layer trains on your spend, predicts the price, attributes it spec by spec, and puts a calibrated range on every estimate — the defensible statistics. A judgment layer sits on top and steers the conversation: which analysis to run, when to stop and ask you for more data, and whether a result is solid enough to quote or needs a real RFQ first. It's what keeps a weaker agent from overstating a number it shouldn't.
Everything runs on your own machine. Your purchasing data is your most sensitive commercial asset, so P2Predict never uploads it and no third party trains on it. Pair it with a local model to keep the whole loop offline, or with a cloud agent if you prefer; either way your raw spend stays put, with no data-residency conversation to have with legal.
It complements should-cost tools, it does not replace them. Bottom-up should-costing builds a part up from material, labor, and machine time to estimate what it should cost to make. P2Predict does not do that and is not trying to. It answers the other question every category manager actually asks: what has the market charged us for parts like this, and what should we expect to pay for the next one?
Proof on public data
Three worked case studies, each reproducible end to end on data anyone can download:
Battery Management ICs: the closest thing to a real procurement job. A small, realistic parts slice, a supplier-premium lever you can quote, and an honest read on where the model is strong and where it needs a real quote.
Used vehicles: the easy-to-follow walkthrough on prices that span orders of magnitude.
Aerospace fasteners: the honesty story. How P2Predict shows you when the data itself sets the limit, so you stop chasing accuracy the data can't give.
Each one leads with results, shows where to trust them, and points to exactly where every number comes from.
FAQ
Is it really free? Yes — including for for-profit companies. Any organization can use P2Predict internally for its own operations, procurement, and benchmarking at no cost. No seats, no trial period, no sales call. A commercial license is only needed if you use it to serve third-party clients (consulting or advisory work) or embed it in a paid product or service.
How reliable are the numbers? Every estimate comes with a range, not just a number, and the range is honest about how thin your data is: wide where history is sparse, tight where it's deep. On the public case studies, predictions land within roughly 16% of the actual price half the time, and within 73% nine times in ten. That's the model showing its work, not asking you to trust it blind.
Does my data ever leave my machine? No. P2Predict trains and predicts entirely on your own machine. The only thing that leaves is whatever your AI agent normally sends to its own model. Pair it with a local model to keep that offline too.
Which AI agents does it work with? Any agent that speaks MCP: Claude, GPT-based agents, or a model running locally. That's the core idea behind P2Predict: it's a capability for your agent to use, not a new tool for you to learn. No dashboard, no separate login. It just shows up wherever you already do your work with your agent.
Is this a should-cost tool? No. Should-costing builds a part up from material, labor, and machine time. P2Predict benchmarks against what the market has actually charged for comparable parts.
Why is this open source? Because a closed tool only ever does what its vendor decided to ship. Source-available means you, or your agent, can read every line, tune the code or the interfaces to your company's ecosystem, or wire it into your own stack and other tools without waiting on anyone's roadmap. That kind of ownership doesn't exist behind a closed API.
Who built P2Predict? It started in 2023 as a project in my free time, weekends and sporadic evenings. See below for the background it draws on.
Who do I talk to about rolling this out for my team? Nobody, really: it installs with one pip command and runs on your own data in a few minutes. Start with INSTALL.md and TECHNICAL.md. If you'd still like a hand, want a commercial license, or want to share a dataset for a future case study, reach out at ahmedhafsi.com/contact. Happy to help.
Who built it
P2Predict is built and maintained by Ahmed K. Hafsi.
Experience
Senior Manager, Negotiation Excellence — Infineon Technologies (2023–present). High-stakes deals across Asia, including foundry, OSAT, and strategic suppliers.
Senior Manager, Negotiation Excellence — Dyson (2020–2023). Architected and led the global Negotiation Excellence capability: governance, KPIs, tooling, and renegotiation campaigns.
Senior Management Consultant, Negotiation & Game Theory — TWS Partners, Munich & London (2014–2019). Advised industrial clients on negotiation strategy and applied game theory.
Education & training
Karlsruhe Institute of Technology (KIT), Electrical Engineering & Information Technology (2008–2014).
Harvard Law School, Program on Negotiation (2021). Negotiation Master Class: Advanced Strategies for Experienced Negotiators.
P2Predict comes out of that work: the tools a procurement team actually needs to walk into a negotiation knowing its number and its leverage. More at ahmedhafsi.com.
Try it / set it up
Set it up with your agent: see INSTALL.md to install, connect your AI assistant, and point it at your data.
How it works under the hood: the models, the math, the full reference live in TECHNICAL.md.
Related
P2CLPFD — the companion tool for the decision that comes next. P2Predict tells you what a part should cost; P2CLPFD decides who gets the volume: the lowest-TCO award across your suppliers under capacity, MOQ, share caps, dual-sourcing, and volume-discount rules, with the reasoning you can show in the room. Open source under the GPLv3. More at ahmedhafsi.com/p2clpfd.
Licensing
Source-available under the PolyForm Noncommercial License 1.0.0, with an additional grant of permission for internal company use.
Free for internal use, including for-profit companies. Any organization may use P2Predict internally for its own operations, procurement, and benchmarking at no cost. You do not need to buy anything, ask permission, or talk to anyone.
A commercial license is only required if you (1) provide consulting, advisory, or procurement services to third-party clients using P2Predict, or (2) embed or integrate it into a paid product or service.
Despite the "Noncommercial" in the underlying license name, commercial companies are explicitly covered by the additional grant — the restriction is on selling P2Predict's use to others, not on being a business.
For a commercial license, a partnership, or to share a procurement dataset, reach out: ahmedhafsi.com/contact.
© Ahmed K. Hafsi. P2Predict is a copyrighted work; all rights reserved except as granted under the license above.
Available Tools
12 toolsexplainA
Explain what is driving a part's predicted price, spec by spec.
Returns a business-ready view to quote directly — `starting_point` (the
baseline price every part starts from) and `price_drivers` (each spec /
supplier's effect in BOTH dollars and percent, biggest mover first) — plus
the underlying technical attribution for your own reasoning. top_n controls
how many top drivers are highlighted (default 3).
Reading it for the user: the explanation carries an axiom check
(baseline +/x contributions = prediction); if it fails the explanation is
unsound. SIGN-CHECK the drivers against intuition — a counterintuitive
sign (e.g. "more cells -> cheaper") on a LOW-importance driver means that
spec is under-sampled, not that the world is upside-down. Quote the
high-importance, correctly-signed drivers; flag the rest as noise. State
effects in the user's terms ("this supplier adds $0.72" / "+18%") — never
say "SHAP", "contribution", or "baseline" to a category manager.
| Name | Required | Description | Default |
|---|---|---|---|
| top_n | No | ||
| features | Yes | ||
| model_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the axiom check (baseline contributions = prediction) and warns that a counterintuitive sign on a low-importance driver indicates under-sampling. It also reveals that the output includes starting_point and price_drivers with both dollar and percent effects. This is substantial behavioral disclosure, though it omits error handling or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average (about 10 sentences) but well-structured: purpose first, then output structure, then usage guidance. Every sentence adds value, and the reading instructions are actionable. It is not bloated, though it could be trimmed without losing meaning. Front-loaded and logically organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (nested objects, output schema present, but schema descriptions absent), the description covers the output structure and interpretation thoroughly, but it does not specify constraints on features (e.g., must match training features), behavior if top_n exceeds driver count, or what happens if the axiom check fails. It also lacks any note on model_id validity. The output schema covers return values, but parameter details and edge cases are incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains top_n ('controls how many top drivers are highlighted, default 3') but does not describe features or model_id at all. Features is implied to be the specs/suppliers, but the format and required fields are undefined. model_id is never mentioned. Partial compensation, but key parameters remain underdocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Explain what is driving a part's predicted price, spec by spec.' It clearly distinguishes from siblings like predict by focusing on the explanation of drivers rather than the prediction itself. The purpose is unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: when you need a business-ready explanation to quote directly, and it gives detailed instructions on how to interpret and present the output (sign-check, quote high-importance drivers, avoid jargon). However, it does not explicitly contrast with siblings like predict or what_if, nor state when not to use it. The context is clear but exclusions are not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_reportA
Generate a procurement-style model-quality PDF report (3 pages).
Page 1: summary metrics + predicted vs actual scatter.
Page 2: error distribution + median % error by price band.
Page 3: top-N feature importance.
The PDF is the human deliverable; the return value also echoes the same
numbers as a structured `quality` block (identical to get_model_quality)
so you can both hand the user the file AND reason over the metrics.
Works best with models trained via the MCP train tool (which stores
holdout data). For older models, the report may be unavailable.
| Name | Required | Description | Default |
|---|---|---|---|
| model_id | Yes | ||
| output_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It explains that the PDF is the primary deliverable, that the return value additionally includes the same metrics as a structured quality block, and that availability depends on training via the MCP train tool. This is meaningful behavioral context, though it does not specify what happens when the report is unavailable or how output_path behaves when omitted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-organized and front-loaded: the deliverable is stated first, then the page breakdown, then the return-value behavior, then the constraint about older models. Each sentence adds necessary information and there is no filler or repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the report's content, the structured return value, and the dependency on holdout data from the train tool, which is the core non-obvious context. It is missing a few operational details such as output_path default behavior and how the tool signals unavailability, but it is largely sufficient for an agent to decide whether and how to invoke this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the input schema. It does not explain model_id requirements or the behavior of output_path (e.g., what happens when null or how the default path is chosen). The usage context implies model_id refers to a trained model, but explicit parameter-level guidance is missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific deliverable ('procurement-style model-quality PDF report') and details the exact page-level contents, so an agent knows precisely what this tool produces. It also differentiates the tool from siblings by noting the returned quality block is 'identical to get_model_quality', which clarifies the relationship to that nearby tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: use this when you need a human-readable PDF report, and it notes that the structured metrics are available for reasoning. It also warns that the report may be unavailable for older models and explains why (reliance on holdout data from the MCP train tool). It does not explicitly say 'use get_model_quality instead when you only need metrics', so it stops short of a complete routing guide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_model_infoA
Get detailed information about a specific model.
Returns a plain `say_to_user` line, the target, and the model's specs with
their types (Numerical/Categorical) and allowed categories — everything you
need to build a predict/explain call. Use it to understand what inputs a
model expects. Lead with `say_to_user`; do NOT surface algorithm / R² /
log-target (or the words 'SHAP', 'log-target', 'R²') to the user. Pass
include_internal=true for those raw fields when you need them to reason.
| Name | Required | Description | Default |
|---|---|---|---|
| model_id | Yes | ||
| include_internal | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden and delivers: return shape (plain say_to_user line, target, specs with Numerical/Categorical types and allowed categories) and the critical presentation constraint — 'do NOT surface algorithm / R² / log-target (or the words SHAP, log-target, R²) to the user.' It also explains that include_internal=true exposes those hidden raw fields for agent-side reasoning, exactly the non-obvious behavioral nuance agents need.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and output shape are front-loaded in the first two sentences, and the presentation rule and include_internal flag each earn their place. The only redundancy is 'Use it to understand what inputs a model expects' partly repeating the earlier 'everything you need to build a predict/explain call' framing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail is offloaded; the description covers what the schema cannot express — when to call, what to lead with, and words never to surface to the user. For a two-parameter informational tool with no annotations, nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds real semantics for include_internal — default false hides internal fields, true returns raw fields for reasoning. model_id is only implied as 'a specific model' rather than explicitly tied to the parameter, but the added detail about specs, types, and allowed categories gives parameter-level meaning the bare schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence 'Get detailed information about a specific model' uses a specific verb-resource pair and scopes it to a single model, distinguishing it from sibling list_models. It further orients the agent by naming predict/explain as downstream consumers, telling the agent this tool is the metadata prerequisite rather than the prediction itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs 'Use it to understand what inputs a model expects' and frames the result as 'everything you need to build a predict/explain call,' giving a concrete trigger condition. It does not name when-not-to-use alternatives (e.g., 'for a list of models, use list_models'), which keeps it from a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_model_qualityA
Structured, agent-readable model-quality report — the JSON form of the PDF.
Use this (not just generate_report, which only writes a PDF) when you need
to *reason about or relay* model quality. Every judgment is computed so you
don't eyeball thresholds:
- `assessment.verdict` — LEAD WITH THIS. One of: 'trustworthy' | 'usable'
| 'unreliable' | 'unknown' | 'insufficient_data'. It folds bias and
sample size into a plain `headline` you can quote verbatim — e.g. a
modest model that is even-handed reads 'usable', not just 'Needs
Improvement'. `assessment.confidence` is 'high' | 'limited' |
'insufficient'.
- `assessment.typical_bias_pct` — how far the model runs high or low on a
typical part, in percent. `assessment.bias_resolution_pct` is the
smallest offset this holdout could have detected; when the verdict is
'unknown', quote it as the bound you CAN rule out.
- `calibration_by_price_band[].reliability` — 'trust' | 'caution' |
'quote' per price range (with `low_confidence` when a band is thin).
Each band carries a `say_to_user` sentence in plain words — quote it to
tell the user which prices to benchmark vs. get a quote on.
- `feature_importance[].signal` — 'strong' | 'moderate' | 'weak', each
with its own `say_to_user` sentence. Only quote findings resting on
'strong' drivers to a stakeholder.
The default payload is deliberately business-only — every string is safe to
read to a category manager. NEVER say 'SHAP', 'R²', 'p-value', 'log-target',
'residual' to the user. The raw statistics (R², p-value, algorithm,
log-target) are NOT in the default response; pass include_metrics=true to
add them under `metrics`/`provenance` for your own developer-level reasoning.
Set include_holdout=true to also get the raw actual/predicted arrays, so an
agent with a code/plotting tool can draw its own charts (predicted-vs-actual,
residuals, error-by-band).
Requires a model trained via the MCP train tool (which stores holdout data).
| Name | Required | Description | Default |
|---|---|---|---|
| model_id | Yes | ||
| include_holdout | No | ||
| include_metrics | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so thoroughly. It discloses the safe business-only default, warns against exposing technical jargon, explains that raw statistics are excluded unless include_metrics is set, and describes what include_holdout adds. It also reveals the computation semantics, such as how verdicts fold bias and sample size, which is far beyond a minimal summary.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but appropriately structured with front-loaded purpose, usage guidance, bullets, and flags. Every sentence contributes meaningful information; there is no filler or repetition, and the formatting makes the key decision points easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity and the presence of an output schema, the description is remarkably complete. It covers tool selection, invocation parameters, output semantics, user-facing safety constraints, and prerequisites. There is no apparent gap an agent would face when deciding to call or invoking this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains include_metrics and include_holdout in meaningful detail, including what payloads they add and for what purpose. model_id's meaning is clear from context, and the description adds the important prerequisite that the model must have holdout data stored by the train tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it returns a structured, agent-readable model-quality report, positioned as the JSON form of the PDF. It explicitly distinguishes itself from the sibling generate_report by saying generate_report only writes a PDF, so an agent can immediately tell this tool apart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit guidance: use this instead of generate_report when you need to reason about or relay model quality. It also explains when to pass include_metrics and include_holdout, and notes the prerequisite that the model must be trained via the MCP train tool, leaving no ambiguity about selection and invocation context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_modelsA
List all trained P2Predict models in the configured models directory.
Call this first to discover which models are available. Each model carries
a plain `say_to_user` line, its target, and its specs. Lead with
`say_to_user`; do NOT read out raw fields like algorithm or R² (and never
the words 'SHAP', 'log-target', 'R²') to a category manager. Pass
include_internal=true only when you need the algorithm name / R² / log-target
flag for your own reasoning.
The response also carries a `server` block (version + git short-SHA + source
path) identifying the build this server process loaded — useful to confirm a
code change actually took effect (MCP servers are long-lived; a stale process
serves old code until restarted).
| Name | Required | Description | Default |
|---|---|---|---|
| include_internal | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It goes beyond the schema by describing the response contents: each model carries a say_to_user line, target, and specs, and a server block with version and git SHA. It also warns that MCP servers are long-lived and may serve stale code, which is genuinely useful operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the purpose, then gives usage guidance, parameter behavior, and an important operational note about stale MCP processes. Every sentence carries practical value and the length is appropriate for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with one optional parameter and an output schema, the description is complete. It tells the agent when to call it, what the response contains, how to present results, and how to handle the internal flag. Nothing needed for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only a boolean include_internal with a default of false and no description. The tool description fully compensates by explaining exactly when to pass include_internal=true (only when the algorithm name, R², or log-target flag is needed for the agent's own reasoning) and what fields are withheld by default.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'List all trained P2Predict models in the configured models directory.' This clearly differentiates it from sibling tools like get_model_info (which targets a single model) and predict (which uses a model). The purpose is unambiguous and immediately actionable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Call this first to discover which models are available,' giving a clear when-to-use directive. It also explains the condition for passing include_internal=true. However, it does not explicitly mention when to prefer sibling tools like get_model_info over this one, so the exclusion guidance is incomplete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
predictA
Predict the target value (e.g. price) for a single part.
Pass the model_id from list_models and a dictionary of feature values
matching the model's expected features. Example:
{"Weight": 15, "Region": "EU", "Supplier": "A"}
Returns an `in_domain` block alongside the prediction — read it before you
quote the number. 'out_of_domain' means the part sits outside the data the
model was built on (an unseen supplier, a spec far past anything observed),
so the estimate isn't reliable no matter how confident it looks; quote
`in_domain.say_to_user` and point the user at a real quote.
| Name | Required | Description | Default |
|---|---|---|---|
| features | Yes | ||
| model_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden; it explains the in_domain/out_of_domain result, warns that out-of-domain estimates are unreliable, and directs the agent to surface say_to_user. This is meaningful behavioral context beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence states purpose; the second gives usage and an example; the final block covers the critical reliability caveat. The prose is a bit detailed on out_of_domain, but every sentence earns its place and the important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter prediction tool with an output schema, the description covers invocation, expected feature shape, and the reliability caveat. It is complete enough for an agent to call it correctly and decide how to present the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, but the description compensates by explaining that model_id comes from list_models, that features must match the model's expected features, and by giving a concrete JSON example. It clarifies both required parameters beyond their names and types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence identifies a specific verb ('Predict'), a concrete resource ('target value ... for a single part'), and an example target ('price'). By explicitly limiting to a single part, it separates itself from siblings like predict_batch and predict_from_csv.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description instructs the agent to pass the model_id from list_models and a features dictionary matching the model's expected features, giving a concrete invocation context. It does not name when-not conditions or alternatives (e.g., when to prefer predict_batch), so it stops short of a full exclusionary guide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
predict_batchA
Predict the target value for multiple parts at once.
More efficient than calling predict repeatedly. Pass a list of
feature dictionaries, one per part. Returns one prediction per row.
Optionally enriches every row with the same views the single-part tools
give, so you don't have to fan out to explain / predict_interval:
- coverage (1-99): adds a likely-price range per row (conformal interval)
when the model carries calibration data; read `interval.reliability`
and `interval.say_to_user` per row exactly as predict_interval does.
Left None (default) for plain point predictions.
- with_explanation: adds the per-row price drivers (same `explanation`
shape as explain). Surface the high-importance, correctly-signed
drivers in dollars/percent; never say 'SHAP' to a category manager.
| Name | Required | Description | Default |
|---|---|---|---|
| rows | Yes | ||
| coverage | No | ||
| model_id | Yes | ||
| with_explanation | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden and meets it well. It explains the conditional behavior of coverage (only when the model carries calibration data), names the interval fields to read per row, and describes the exact shape of with_explanation output as matching explain. It even adds user-facing guidance about avoiding the term 'SHAP'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then moves into a clean bulleted list for the optional enrichment behaviors. Every sentence carries operational value, including the final output-guidance note about never saying 'SHAP' to a category manager.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema, so return-format details are covered structurally; the description focuses on what an agent needs to decide and invoke it correctly. It covers batching semantics, row-to-prediction correspondence, optional behavior, and edge conditions such as missing calibration data. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so this dimension depends entirely on the description. It defines rows as 'a list of feature dictionaries, one per part,' clarifies coverage as an optional integer 1-99 that toggles conformal intervals, and links with_explanation to the explain output shape. Model_id is not elaborated, but its purpose is evident and the other three parameters gain substantial semantic meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb and resource: 'Predict the target value for multiple parts at once.' It immediately distinguishes itself from the single-part tools by framing the batching and efficiency gain, and explicitly references predict, explain, and predict_interval as the alternatives it replaces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly states when to use the tool ('more efficient than calling predict repeatedly') and how the optional flags avoid fanning out to explain / predict_interval. The description tells the agent to pass a list of feature dictionaries and that each row returns one prediction, so the invocation pattern is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
predict_from_csvA
Batch-predict from a CSV file on the local filesystem.
The file-based sibling of predict_batch — use it when the user drops a
spreadsheet of parts. Reads csv_path, predicts every row, and returns one
prediction per row (point estimates by default).
The same opt-in enrichments as predict_batch apply per row:
- coverage (1-99): adds a likely-price range per row (conformal interval)
with its `interval.reliability` / `interval.say_to_user` read. Requires
a model with calibration data; an explicit coverage on an uncalibrated
model returns a no_calibration error. Left None (default) for plain
point predictions.
- with_explanation: adds the per-row price drivers (same `explanation`
shape as explain). State drivers in dollars/percent; never say 'SHAP'
to a category manager.
| Name | Required | Description | Default |
|---|---|---|---|
| coverage | No | ||
| csv_path | Yes | ||
| model_id | Yes | ||
| with_explanation | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the behavioral burden and performs well. It discloses that it reads csv_path, predicts every row, returns one prediction per row with point estimates by default, and explains both opt-in enrichments. It also reveals error behavior (no_calibration when coverage is requested on an uncalibrated model) and the internal reliability/say_to_user fields, providing richer transparency than typical definitions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently organized: a purpose statement and usage cue in the first two sentences, then the default behavior, followed by the two optional enrichments in a scannable list. Every sentence provides substantive information, though the with_explanation guidance ("never say 'SHAP'") is a contextual nicety rather than core mechanics. It is longer than strictly necessary but remains tight and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 4 parameters and an output schema, the description covers the essential calling context: input file, per-row predictions, default point estimates, both optional enrichments with their constraints and errors. The output schema accounts for return-value details, so the description does not need to repeat them. Minor omissions such as CSV format requirements or file-not-found behavior are acceptable at this level of complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does for the optional parameters: coverage is specified as 1-99, its default is None, its interval fields are named, and the no_calibration error condition is given; with_explanation is tied to the explanation shape from explain. However, the required parameters, model_id and csv_path, only appear in the flow of the text without additional format or constraint details, which is a modest gap given their importance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Batch-predict from a CSV file on the local filesystem.' It then explicitly differentiates itself from siblings by calling itself 'the file-based sibling of predict_batch' and giving the exact trigger condition ('when the user drops a spreadsheet of parts'). This leaves no ambiguity about what the tool does or how it differs from predict_batch and predict.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: 'use it when the user drops a spreadsheet of parts' and names the sibling alternative (predict_batch). It does not enumerate exclusions (e.g., when not to use it versus predict or predict_interval), but the file-based context and reference to predict_batch provide a clear decision point. This is strong usage guidance, though not exhaustive about all alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
predict_intervalA
Predict with a likely range (conformal prediction interval).
For a 90% interval, about 9 in 10 similar parts fall within the range.
coverage is an integer 1-99 (default 90). Requires a model trained with
P2Predict v0.5+ (which stores calibration data).
Reading it for the user: the band WIDTH is the per-part trust signal, and
the payload computes it for you — `interval.reliability`
('trust' | 'caution' | 'quote') and a plain `interval.say_to_user` sentence
you can quote directly. A tight band = predict with confidence; a very wide
band — or a lower bound at/below $0 on an additive (non-log) model — means
"get a quote, don't benchmark." Always surface the range, not just the point
estimate, when the user will act on the number.
| Name | Required | Description | Default |
|---|---|---|---|
| coverage | No | ||
| features | Yes | ||
| model_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explains the interval reliability output ('trust' | 'caution' | 'quote'), the meaning of tight versus wide bands, the lower-bound-at-$0 caveat on non-log models, and that the payload provides a ready-to-quote sentence. This is unusually transparent about expected behavior and interpretation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and then organized into calibration explanation and user-facing reading instructions. Every section adds practical value, though the second half is fairly long and could be tightened without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the tool's purpose, a prerequisite, the output interpretation, and actionable guidance for when to trust the prediction. The main completeness gap is the underspecified 'features' parameter, but the output semantics are thoroughly explained even though an output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does add meaningful semantics for the 'coverage' parameter (integer 1-99, default 90, and what it means probabilistically) and implies the model_id prerequisite. However, the 'features' object—the other required parameter—remains completely opaque; the description does not say what keys or shape are expected.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Predict with a likely range (conformal prediction interval).' It clearly distinguishes this from a point prediction (e.g., the sibling 'predict') by defining the interval semantics and even explains what a 90% interval means.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives strong contextual guidance: it requires a model trained with P2Predict v0.5+ and explicitly tells the agent when to surface the range rather than just the point estimate ('when the user will act on the number'). It does not explicitly name alternatives like 'predict' or state when not to use this tool, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
propose_training_planA
Inspect a training CSV and return a plain-language plan BEFORE training.
Call this first whenever a user wants to build a should-cost / pricing
model. It reads the CSV, decides what it would predict, which columns it
would use as specs, which it would leave out (and why — target leakage,
ID-like columns), and whether the target should use a log-target. It
trains nothing and writes nothing.
Relay `plain_summary` and `questions_for_the_user` to the user in their
own language, get confirmation, then call `train` (passing the agreed
`features` and `log_target`).
| Name | Required | Description | Default |
|---|---|---|---|
| target | Yes | ||
| csv_path | Yes | ||
| max_features | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses key behavioral traits: it reads the CSV, decides predictions, columns to use/leave out (with reasons like target leakage, ID-like columns), and whether to use a log-target. It explicitly states it trains nothing and writes nothing. While annotations are absent, the description carries the burden well. It could add more about side effects or error conditions, but it covers the main behavioral aspects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and front-loaded. The first sentence states the core purpose. The second paragraph provides workflow guidance. The third paragraph gives clear next steps. Every sentence earns its place, and the description is appropriately sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is quite complete for a planning tool. It explains the tool's role in the workflow, what it returns (plain_summary and questions_for_the_user), and how to proceed. It doesn't describe the output schema in detail, but the output schema exists and the description doesn't need to explain return values. It could mention potential errors or edge cases, but overall it's sufficient for an agent to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It mentions `features` and `log_target` as outputs to pass to `train`, but does not explain the input parameters `csv_path`, `target`, or `max_features` in detail. The description implies `csv_path` is the CSV to inspect and `target` is the target column, but it doesn't explicitly map them. This is a gap, but the description does provide some context about what the tool does with the CSV and target.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: inspect a training CSV and return a plain-language plan before training. It specifies the resource (training CSV), the action (inspect and propose), and distinguishes it from the sibling `train` tool by explicitly saying it trains nothing and writes nothing. This is a specific verb+resource combination that an agent can easily understand.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Call this first whenever a user wants to build a should-cost / pricing model.' It also provides a clear workflow: relay the summary and questions to the user, get confirmation, then call `train` with the agreed features and log_target. This is excellent usage guidance, including when to use it and what to do next.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trainA
Train a new P2Predict model from a local CSV file.
Prefer calling `propose_training_plan` first and confirming with the user
— this tool is the execution step. The CSV must have spec columns and a
price/cost target column. Training runs locally; no data leaves the
machine. The trained model is saved and immediately available.
Safe defaults (always surfaced in the returned `warnings` list):
- When features are auto-selected (features=None), columns that look
like target leakage — a near-duplicate of the price being predicted —
are excluded automatically.
- For a strictly-positive (price/cost) target where the automatic skew
test leaves the log-target off, the result recommends log_target="on".
algorithm: "auto" (default), "ridge", "random_forest", or "xgboost".
budget: "fast" (default) or "thorough".
log_target: "auto" (default), "on", or "off". Use "on" for prices.
allow_leaky_features: set True only to override the leakage guard and
train on an explicitly-requested feature that looks like leakage.
| Name | Required | Description | Default |
|---|---|---|---|
| budget | No | fast | |
| target | Yes | ||
| csv_path | Yes | ||
| features | No | ||
| algorithm | No | auto | |
| log_target | No | auto | |
| max_features | No | ||
| outlier_policy | No | warn | |
| allow_leaky_features | No | ||
| feature_outlier_policy | No | warn |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It discloses that training runs locally, that no data leaves the machine, that the trained model is saved and immediately available, and that safety defaults like leakage exclusion and log-target recommendation are surfaced in warnings.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action, then moves through workflow, constraints, behavioral defaults, and parameter values in a logically organized way. The bulleted safe-defaults section makes the leakage and log-target behavior easy to parse without excessive length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be repeated in the description. The description covers workflow, prerequisites, privacy, persistence, safety defaults, and several key parameters, but a few policy parameters such as outlier_policy and feature_outlier_policy still lack explicit semantics, leaving a minor gap for a 10-parameter training tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does document algorithm, budget, log_target, and allow_leaky_features with concrete values and guidance, but it leaves max_features, outlier_policy, feature_outlier_policy, csv_path, and target mostly to be inferred from their names rather than explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Train a new P2Predict model from a local CSV file.' It also differentiates this tool from its sibling by positioning it as the execution step after propose_training_plan, so an agent can clearly tell what this tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says to prefer calling propose_training_plan first and to confirm with the user before using this execution tool. It also states the required CSV shape and the privacy property that training runs locally, giving the agent concrete guidance on when this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
what_ifA
Compare a base scenario with a counterfactual where features change.
Returns a plain `summary` to quote directly (does the change add or save,
how many dollars, what percent, old vs. new price), plus both predictions,
the delta, and per-driver attribution of each change for your reasoning.
Answers "what if we switch from supplier A to B?" Set coverage to null to
skip intervals.
`summary.reliability` grades how far to trust the DIRECTION of the move —
'trust' | 'caution' | 'quote' — with a plain `summary.say_to_user` line to
quote. It fires when the changed spec is thinly sampled, or when the swing
is driven by knock-on effects rather than the change itself (a `quote` here
means the number is unstable — relay the caution, don't present the
direction as fact). It flags instability, not a wrong domain sign, so still
sign-check a 'trust' result against intuition.
| Name | Required | Description | Default |
|---|---|---|---|
| changes | Yes | ||
| coverage | No | ||
| features | Yes | ||
| model_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explains the return structure (summary, both predictions, delta, per-driver attribution), the meaning of summary.reliability ('trust' | 'caution' | 'quote'), when reliability fires (thin sampling, knock-on effects), and how to interpret a 'quote' (unstable number, relay caution). It even warns to sign-check 'trust' results. This goes far beyond basic transparency and is exceptionally detailed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is moderately long but every sentence adds meaningful information. It opens with the core purpose, then details the return format and reliability interpretation. The structure is logical and front-loaded. It could be slightly more concise, but it avoids fluff and each clause serves a purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (4 params, nested objects, output schema present), the description covers the key operational details: what the summary contains, how to quote it, how reliability works, and the coverage null behavior. It does not cover error cases or edge scenarios, but the output schema presumably handles return values. The description is sufficient for an agent to call the tool correctly in most situations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does explain the coverage parameter ('Set coverage to null to skip intervals') and hints at the shape of features/changes via the supplier example, but it does not detail the structure of the 'changes' object or how features are specified. This leaves room for ambiguity, but the example provides some semantic guidance. The description adds value but is not fully compensatory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Compare') and resource ('base scenario with counterfactual'), and immediately gives a concrete example ('switch from supplier A to B?'). It clearly distinguishes this tool from siblings like predict (which predicts without counterfactual comparison) and explain (which explains without scenario changes). The purpose is unambiguous and action-oriented.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear use case with a concrete example ('what if we switch from supplier A to B?') and explains when to use it (comparing scenarios with feature changes). It does not explicitly state when NOT to use it or mention alternative tools, but the example and the tool's name make the usage context clear. Minor gap: no explicit exclusion of cases better suited for predict or explain.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
12 tool updates
v1.1.1- First observed
explain - First observed
generate_report - First observed
get_model_info - First observed
get_model_quality - First observed
list_models - First observed
predict - First observed
predict_batch - First observed
predict_from_csv - First observed
predict_interval - First observed
propose_training_plan - First observed
train - First observed
what_if
TDQS
Scored across 12 tools
The prediction tools (predict, predict_batch, predict_from_csv, predict_interval, explain, what_if) have overlapping surface areas but each has a clearly distinct input mode or output focus. The one genuinely close pair is get_model_quality and generate_report; their descriptions do distinguish structured JSON from PDF generation, but an agent could still hesitate.
Most tools follow a clear snake_case verb or verb_noun pattern, and the predict_* family is consistently named. Minor deviations like 'explain' and 'what_if' break the pattern slightly but remain readable and predictable.
Twelve tools is well within the ideal range and each tool covers a distinct stage of the model workflow: discovery, training, prediction variants, explanation, quality assessment, and reporting. No tool feels redundant enough to remove.
The set covers model discovery, training, single/batch prediction, intervals, what-if analysis, explanation, quality review, and PDF reporting—very complete for a prediction-focused server. The main gaps are missing model deletion/update operations and a direct model-comparison tool, but these are not core to the stated purpose.
Maintenance
Related MCP Connectors
Parametric should-cost: P50/P80/P90 estimates, 801 materials, 25 countries, quote review.
Instant parametric cost estimates for custom manufacturing: CNC, molding, sheet metal, 3DP, PCB.
Turn agent intent into physical parts: engineering review, measured geometry, calibrated pricing.
Agent-native supply network for components, fabrication, industrial RFQs, offers, and fulfillment.
Related MCP Servers
- AlicenseAqualityAmaintenanceEnterprise-grade (40m+ lines) codebase intelligence in a zero-setup, private and local MCP: managed indexing, hybrid semantic search, polyglot code dependency graphs, and DB/API/infra knowledge. Benchmark: 61% less tokens, 84% fewer calls, 37x faster than standard AI grep.264,715 npm3,307AGPL 3.0
- AlicenseAqualityDmaintenanceEnables AI agents to discover, compare, and select the best AI models across multiple providers based on pricing, performance, and capabilities, with real-time cost estimation and benchmarking.939 npm1MIT
- AlicenseBqualityDmaintenancePredict the cost of an LLM call before you make it, and pick the cheapest model that still does the job, offline, from your editor.727 npmApache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables agent-assisted CAD engineering, allowing users to create, validate, and export CAD designs through natural language, with a deterministic engine that has zero LLM runtime dependency.Academic Free v1.1