HORIZON SHIELD KIRA
Use this MCP server to check Japanese construction/renovation estimate fairness, explore JCCDB open cost data, verify signed price claims, and find vetted contractors.
Get fair-price ranges (min/avg/max, danger threshold) for Japanese renovation/construction work by keyword, optionally adjusted by region.
Audit a contractor's quoted price against fair ranges and return verdict, level, gap vs average, and advice.
Check estimate or sales-pitch text for known overcharge/red-flag tactics (lump sum, today-only discount, door-to-door, etc.).
Get universal estimate-reading principles: overhead ratios, lump-sum treatment, pressure-sales detection.
List and search maintained construction/renovation cost categories with red-flag counts and priority.
Look up fair-price data sources, update dates, and regional multipliers.
Get JCCDB dataset metadata: scale, license, download links, citation.
Preview reverse-estimate direction (above/below average) before a detailed quote exists.
Issue a signed/tamper-evident fair-price receipt with SHA-256, verify URL, and PTKA provenance.
Create an AP2 FairPriceAttestation for Cart Mandates without initiating or executing payment.
Independently verify signed integrity claims fail-closed by recomputing SHA-256.
Get the A2A Agent Card URL and published skills for agent-to-agent discovery.
Get an anonymous third-party estimate review link (EHN board).
Find verification-passed contractors by area/work, with scores and tiers, no prices.
Most tools are read-only; verify_fair_price and create_ap2_fairness_attestation append or issue records. Japan/JPY focus, no API key required.
Issues FairPriceAttestation attestations that can be attached to Google AP2 (Agent Payments Protocol) Cart Mandates, enabling verifiable fair-price proofs alongside payment authorization.
🛡️ HORIZON SHIELD
Verifiable construction estimate auditing for AI agents
Don't trust the estimate. Verify it.
An MCP server that lets AI agents check whether a Japanese construction or renovation estimate is fair, against open data, and returns a result anyone can verify against Bitcoin (OpenTimestamps). No account, no key.
In one paragraph. JCCDB (Japan Construction Cost Database) is an open dataset of Japanese construction and renovation costs, created by Toshikatsu Oga (大賀俊勝), who has worked on construction sites for 30 years, and published by The HORIZONs Co., Ltd. under CC BY 4.0. Version 5.1 (2026-10-04, DOI 10.5281/zenodo.23133068; all versions 10.5281/zenodo.22127751) holds 526,128 records: 95,403 line items and 430,725 observations from 88 Japanese public sources, each observation with its evidence URL. The United States counterpart is USCCDB, the United States Construction Cost Database (DOI 10.5281/zenodo.22979157, 2,849,829 observations). HORIZON SHIELD (https://shield.the-horizons-innovation.com) is the buyer-side service built on JCCDB that checks whether a Japanese renovation estimate is fair.
Fair-price answers for common buyer questions (Japanese)
Each row is one question a homeowner in Japan asks, the page that answers it with the range from souba-db 2.2.0, and a JSON evidence object with the same numbers, a comparison across variants, steps to check a quote yourself, and the dataset hash anchored in JIDEC entry 42. Index: evidence/index.json.
Question | Answer page | Evidence object |
外壁塗装 30坪 相場 いくら | ||
外壁塗装の見積もりで150万円は高いですか | ||
屋根 葺き替え 30坪 費用 | ||
この屋根修理の見積もりが適正かどうか知りたい | ||
給湯器 交換 費用 相場 | ||
給湯器交換で20万円は高いですか | ||
トイレ リフォーム 費用 目安 | ||
シロアリ駆除 費用 適正価格 |
Related MCP server: japan-real-estate-intel
Independent evidence, as of 2026-10-03
What people who do not work for this project have measured, signed or reproduced. Every row links to something you can fetch and recompute. The last two rows are the counts that are still small, stated as plainly as the rest.
What | Who | Check it |
Walked our gate from their own host, served the record content-addressed on their own domain, and filed the same bytes to our ledger signed with the key their domain serves | Federico Blanco Sánchez-Llanos's agent, | |
Performed and signed an execution under a contract both sides signed. It was anchored in Bitcoin block 969090 and settled |
| |
Walked the gate's A2A face twice from his own Windows PC. He found that | Pavlo Tvardovskyi ( | #27, FINDINGS_EXTERNAL.md, his key, |
Reproduced, in an independent Go implementation, the | Kuang Mi ( | |
Reproduced our signature vectors in a reader of their own: the 13 unknown-field cases (s3) and the 5 dual-name cases (s4), under both readings, and matched the a2a-go column of our MANIFEST | Sankalp Gilda ( | |
Verified the a2a-python fix candidate and our 11 donated tests without using our test files, with an ES256 verifier he wrote by hand; the length and sha of all 13 vectors he checked match the | Kuang Mi ( | |
Published a candidate verifier that canonicalizes the received JSON; its self-reported results match our corpus on every vector (s0 5/5, s1 8/8, s2 11/11, s3 13/13), now a pinned column of the MANIFEST |
| |
A two party agreement signed by both sides, with consent to publish inside the signed bytes; both signing keys match the keys each domain serves | this project and |
|
Re-verified the gate's agent card signature with their own verifier: 10/10, and the digest they recorded, | Agenstry, an independent agent directory in Amsterdam (2026-09-28) | |
Scored the JIDEC ledger export under an outside ledger conformance spec and asserted it in their CI: L1 against the head stamped in our daily Bitcoin batch, L0 on the log alone since their 1.4.1-draft corrigendum found our end marker was not itself chained (EXT-022). Fixed on 2026-09-29 (833f8048): the marker now carries its own entry_sha256, linked to the last entry by the same recipe, so a rewritten marker breaks a hash; their rescoring is not yet published. An issue we reported is recorded there as EXT-020 | the VLC-1 specification's maintainers | VLC-1, THIRD-PARTY.md |
Outside operators who used the gate to measure their own servers | 5 real hosts in the 30 days to 2026-10-02 (the counter lists 6; one is a test name that does not resolve), the last on 2026-09-23 (UTC) |
|
Rows on the public register that are not ours | 2 of 10: one still pending, one added anonymously with |
|
Counted every week by tools/adoption/count_adoption.py, last on 2026-10-02. Every number comes from a source you can read; one that could not be read says so instead of counting zero. The whole count: ops/adoption/latest.json.
What | Count | From |
Independent implementations that reproduced our bytes or verdicts | 6 rows by 5 authors (a2a-card-canonical-form 1, a2a-card-sign-v01 4, agent-card-signature 1) |
|
Outside domains that signed a walk and filed it to the ledger | 2 ( | every |
Re-verification pool | 1 control cluster(s), 2 needed for a quorum |
|
MUSUBI contracts signed with an outside party | 2 (with no party from this project: 0) | the signed contracts in |
Outside identities that signed evidence (walk, contract or agreement) | 2 | the three rows above and the agreement records |
Public repositories created from conduct-witness-template whose reproduce run succeeded in the last 30 days | 0 | GitHub API |
Open: A second verifier of NENRIN provenance bundles, by another author in any language, that reproduces the five verdict signatures in interop-v0/expected.json. None yet: workers/hs-ledger/nenrin/interop-v0/INTEROP.md
One external witness is a start, not a network. The re-verification pool below counts it as one control cluster, and a quorum of two independent controls is not met yet. Issue #27 is the open call for the second.
When every signature verifies: what can still deceive this system
Breaking a parser, a signature or a chain is the attack this repository was built against first. The harder question is the one left when all of that holds: every byte verifies, and the system is still told something false. These are the six ways we know of, what is built against each, how much of it real use has exercised so far, and what no amount of code here can do.
Attack | What is built | Exercised in real use so far | What it cannot do |
Organizational Sybil: witnesses or parties that hold different keys but answer to one controller |
| The re-verification pool has one member, so it reports | Stop a Sybil that uses different providers and different legal entities. It makes one expensive and visible; it does not make one impossible |
Signed lie: a correct signature over an observation that is false |
| Two outside witnesses have walked the gate (issues #25 and #27), and one outside finding changed the protocol (EXT-001). No signed contract has required corroboration yet. Since 2026-10-02 one can: settle v1.8 holds | Decide who is right when independent entities disagree, or catch a lie that every independent measurer tells |
Semantic contract: both sides sign the same bytes and read "done" differently |
| Both signed contracts are witness walks: done means a filed walk, which | Fix any meaning the vocabulary does not name |
Hidden decision influence: a published rule, with a prompt, reward or incentive nobody sees steering the choice | The sieve and the policy behind a contract decision are signed and pinned by sha256 beside the contract (run0001). conduct-v1 requires an agent to disclose who pays it: referral, listing and success fees, in its agent card | In run0001 the sieve and the hidden instruction detector are pinned but not published, so the outside recompute marked them NOT CHECKED | Prove that nothing undeclared influenced a decision. A pinned rule shows which rule was applied, not that it was the only influence |
Durable availability: the anchor survives and every copy of the bytes is gone |
| One mirror is held outside this company: Federico's, which he diffed one of our agreements against on 2026-09-28 | Guarantee that any copy survives. Availability is the count of independent holders, and today that count is small |
Fake observable surface: every public face reports healthy while the real state is not | The gate calls real tools where the operator consents, not only health pages, and the day and the tool it measures are derived from a Bitcoin beacon neither side chooses ( | Two outside vantages so far. No TRACE record from an outside party has been pinned yet | See past a surface that is equally false to every observer at every moment. External observation measures what is served, and says so |
Key compromise: a stolen key makes signatures that are mathematically genuine | key-history-v1 (spec, served at | Four keys are listed: card, agreement, witness and operator. None revoked | Say when a key was really stolen. |
Physical-world oracle: everyone signs that the work was done, and it was not | A measurement in | No contract has required a measurement yet; the first one that pays for physical work is the first test. Since 2026-10-02 such a contract can require the measurers to be drawn rather than chosen (settle v1.9), so the parties cannot bring their own | Turn the physical world into proof. When every entity lies together, the ledger keeps exactly who said so, when, and against which terms |
Network: nobody else takes part | Nothing in code. Every door is open, needs no key and costs nothing | Counted in the table above, including what is still small. | Create participants. Only use does that |
One verifier threads the contract rows together: spine_verify.py reads a contract from what it was agreed to mean (terms_sha256), through who agreed and who did it, to who measured done, in entities, inside a block window. The first three rows share a limit worth naming once: the verifiers exist and are tested against their own attacks, but a verifier nobody's contract invokes proves only that it would work. The next contract that pays for real work is the one that has to state an independence quorum, name its deliverables from a vocabulary and require corroboration. Since 2026-10-02 the contract itself can make that binding: settle v1.8 (requirements.spine) refuses final until every stage is in place, and settle v1.9 (requirements.convergence) adds measurers nobody chose. What is still missing is a signed contract that uses them, and one between two parties neither of which is this project.
All of MUSUBI also installs without a clone: pip install nenrin-verify (0.3.0) carries musubi-v0/ byte for byte, and musubi-verify spine_verify --selftest or musubi-verify --run0002 runs the same files.
Price ranges you can recompute
Every get_price_range answer carries a recompute block: the URL and SHA-256 of the souba-db.json bytes it used (the same file is in this repository at data/souba-db.json), the entry id of each row, and the formula: the table value, or the table value times the regional multiplier in the same file, rounded half up. tools/recompute_price_range.py checks an answer end to end with the Python standard library, and fails on any byte or row that differs.
What it shows is that the answer equals the published table. It does not show that the table is right: the values are curated by a named curator against the sources listed in the file, not computed from those sources by a published formula.
NENRIN: tree rings for AI facing services
A tree adds one ring a year. Nobody can paint one in afterwards. NENRIN gives that property to software services.
In one thirty day window, measured 2026-08-17, this server appeared in 93,983 AI search results. How many of those became a call from outside, we cannot say. The usage counter deliberately stores no IP addresses, so it cannot separate our own automated checks from external traffic. An earlier version of this paragraph said the answer was 0. This instrument cannot establish that, so the claim is withdrawn here rather than quietly deleted. Discovery is solved. Choice is not. An agent picking between 90,000 servers can only read what each vendor wrote about itself. NENRIN adds the missing layer: records of conduct that the vendor did not author and cannot delete.
How it works, in three lines:
Open witnessing. Anyone can measure any endpoint and submit the walk to the public ledger under their own name and vantage. The operator holds no veto: acceptance is mechanical schema checking, and the code that enforces this is in this repository.
Discrepancies are the product. When two witnesses report incompatible observations of the same target, the disagreement itself becomes a permanent, citable record. The founding one is real: NENRIN_DISCREPANCY_0001, two honest witnesses, one target, both correct.
Rings. Each month the accepted records bundle into a ring that carries the hash of the previous ring, timestamped to Bitcoin. Eighteen months of rings cannot be created in an afternoon, by anyone, including us.
The specification is anchored on the public ledger as entry 19
(sha256 9ccba2e325fd2a555fcdb2dec519b8c6bf7a669064674846aea98ecfff824e3d):
NENRIN_SPEC_v1.md. It names its own prior art (Certificate Transparency, Rekor, in-toto, SLSA, OpenTimestamps), states exactly which combination is claimed as new, and invites refutation into the same ledger.
The witness intake is live. Start here:
curl -s https://ledger.horizonshield.dev/witnessWe are the first test subject under our own rules. The ledger keeps the record of our gate failing its own test, and the full 522 incident that started all of this. Unflattering records stay.
If a register that cannot delete criticism of its own operator is infrastructure you want to exist, star this repository. Stars are how researchers and agent platforms find it. The rings accumulate either way. They accumulate faster with witnesses.
Task-bound conduct: binding an A2A Task to its evidence
Choice does not end when an agent picks a server. It picks, then it delegates a task. What that delegated task actually did, observed by someone other than the two parties, is the evidence the next agent needs. NENRIN binds it to the A2A Task id itself (a2a.task.id, aligned to A2A issues #1769 and #2103), not to "this server failed once".
The conduct walk already talks to a real agent over A2A and receives a real Task with an id. It now files a signed, content-addressed observation bound to that id: who delegated to whom on task T, and how the walked agent behaved. The witness signs the observation (witness_sig), the requesting party signs the delegation edge (edge_sig), and both keys live inside the did:key identifiers, so anyone verifies with no network and no trust in us.
Three reads, each a real record you can fetch now:
curl -s "https://ledger.horizonshield.dev/witness/task?task_id=d1651c71-28b0-422b-8f14-e2dc66c5a145"
curl -s "https://ledger.horizonshield.dev/trust-signal?task_id=d1651c71-28b0-422b-8f14-e2dc66c5a145"
curl -s "https://ledger.horizonshield.dev/witness/task/evidence/0bff13042d89ec0bc33f2f5149612774f8a706e67c1f345abfc1f10962e31608"The first returns the full witness set per delegation hop, with the aggregate verdict computed so a disagreement is preserved and never the favorable one. The second returns the same as a consumable signal that carries counts and verdicts and never a numeric score. The third returns one observation with its anchor status: this evidence sits in NENRIN ledger entry 44, a nenrin-task-witness-batch-v1 bundle timestamped to Bitcoin like every other ring.
What the signatures prove, stated plainly: who asserted the observation and who attested the delegation edge, not that the assertion is true. The ledger attests that the witness is distinct from both hop parties by key (R1); it does not attest operator independence, so a self-witness satisfies R1 and says so in its own record. A genuine third-party observation is the same walk run by someone with no stake, filed to the same live endpoints.
The loop this closes: an agent discovers a server, reads conduct the server did not write, chooses, delegates a task, the task is witnessed, the evidence accumulates bound to the task id, and the next agent chooses on it. The code is in workers/hs-ledger/nenrin/task-delegation-bind-v0 (the ledger faces and the producer) and workers/hs-ledger/nenrin/a2a-conduct-walk (the walk that binds, with --bind-task).
TSUGI: proof of recovery, the second pillar
Verification says whether an endpoint conforms today. It says nothing about what happened when it broke, or whether it is really back. TSUGI (継, from kintsugi: the repair is visible and becomes part of the object's history) is the layer after verification. It does not repair; it proves recovery.
Five record types, hash-linked: drift (a witness measured a public surface and it did or did not match), proposal (one repair from a closed catalog of five primitives, with what it does not establish), authorization (the operator's Ed25519 signature over the proposal hash, expiring), execution (before and after state), verify (the witness measured again). The verifier refuses an execution of a human-approval primitive whose authorization is unsigned, signed by an untrusted key, or expired. The operator's public key is served at https://gate.horizonshield.dev/keys/operator.json, the same way the agreement and witness keys are.
Re-verification witnesses are not chosen by the operator. They are drawn from a public pool with sha256(bitcoin block hash | pool hash | record hash) as the seed, so a third party recomputes who should have been asked. A drawn witness receives only a blind request (no expected values) and returns only a signed observation; nothing it says is executed. This is the July 2026 lesson turned around: unknown agents may observe you, never instruct you.
Two real incidents are recorded in workers/hs-ledger/nenrin/recovery-v0: a raw deploy that bypassed the deploy guard and silently broke the card signature and the OpenAI domain challenge (found by an external verifier, closed with an unsigned chat approval, which the strict verifier flags as such), and a card signature broken by two version bumps deployed without a re-sign (found by the daily witness, closed with a signed authorization, twelve records). The pool of external witnesses was empty on 2026-09-20, and the records say so instead of pretending a quorum. It now holds one member, Federico Blanco Sánchez-Llanos's agent; the diversity check (witness_diversity v2.3) counts it as a single control cluster, so the records still report the quorum as short rather than met. Incident 2 can be recomputed in a browser, hashes and the operator's Ed25519 signature, with no trust in this project: https://shield.the-horizons-innovation.com/tsugi/ Its chain file and record hashes are anchored as JIDEC entry 50 (OpenTimestamps, Bitcoin).
Repository map
Path | What it is |
| The verification gate: nightly sweeps, on demand checks, |
| The JIDEC append only ledger and the NENRIN witness intake |
| Task-bound conduct: an A2A Task id bound to a signed, Bitcoin-anchored witness observation; the |
| The conduct walk that measures an agent and, with |
| The agreement record: two agents, two signatures, one set of bytes. Verifier written twice, in Python and JavaScript, and proved to agree |
| TSUGI (継), the second pillar: proof of recovery. A drift witness measures eight public surfaces of the gate daily; a repair is proposed from a closed catalog, authorized with the operator's Ed25519 key (trust anchor at |
| The public edge relay born from the 522 incident (documented in the discrepancy record) |
| The public register page: every listed server, our own included, with its live verdict |
everything else | The GitHub Pages site for the human facing service at the-horizons-innovation.com |
The agreement record: the other half of a measurement
A conduct record is one sided. Somebody measured somebody. Nothing in it records the other half of commerce: that two agents agreed on terms, and that both said so.
a2a-agreement-v1.1 is that record. At time T, party A and party B both signed the same canonical
bytes describing terms, and each of them pinned, by sha256, a conduct record about the OTHER party
written by somebody who is neither of them.
What it refuses to be is as load bearing as what it is. No custody. No matching. No editorial step. The recorder must not hold funds, must not decide whether a deal happens, and must not charge a fee that varies with the amount or the outcome. A record whose fee moves with the number is refused by name. Refusal is mechanical, and none of the terms are ever judged by anyone in this layer.
The claim is not a new primitive. It is the combination: two mandatory signatures, the counterparty's measured conduct pinned by sha at the moment of signing, an intake that judges nothing, and an external anchor nobody here operates. Prior art is named in the draft rather than left for a reader to find: AP2, x402, ACP, MPP, Cedulon, the 1F916 Agent Record, and SCITT.
The verifier is written twice. Once in Python, once in JavaScript, by design and not by accident: two implementations that disagree are the exact seam this project measures everywhere else, and building one into this layer on purpose would be a poor joke. 5,286 frozen cases, and the two produce the same report byte for byte, including every refusal code and the English sentence attached to it. Proving that moved the Python once, when the JavaScript disagreed on two cases and the check that settled it was running the Python against its own frozen fixture, where it failed the same two.
Then the rules were broken on purpose, 77 ways in Python and 36 in JavaScript, to find out whether the 5,286 cases could tell. Six breakages survived, and not one was a defect in either implementation. They were holes in the test set. All six are closed.
The record:
ops/AGREEMENT_EXT_v0_1_DRAFT.md. v0 is anchored as JIDEC entry 39 and does not move.The verifiers, the adversary and the contract:
workers/hs-ledger/nenrin/agreement-v0What an intake may and may not do, written before one existed:
ops/AGREEMENT_INTAKE_v0_BOUNDARY.md; the five open decisions and how they were settled:ops/AGREEMENT_INTAKE_v0_DECISIONS.md; state:ops/AGREEMENT_INTAKE_v0_STATUS.md
The intake exists since 2026-09-16 (agreement_intake.mjs,
wired into the ledger worker): POST /agreement accepts a record only when both signatures verify
against the keys each party serves at its key_url, deduplicates atomically, serves the record by
sha at GET /agreement/{canonical_sha256}, and bundles the accepted pool into a daily anchored
ledger entry. It judges nothing. The first record it accepted, between this project's agent and
Federico Blanco Sánchez-Llanos's agent, is in the tree as
first_agreement_record.json.
An earlier version of this paragraph said there was no intake; the code had been written and
deployed but not committed, which this repository noticed on 2026-09-20 and corrected.
The second record between the same two agents was re-signed on 2026-09-28 with "publication": "public"
inside the bytes both parties signed (record-privacy-v1: nothing bilateral is published on one side's say so).
It was accepted the same day with both signing keys matching the keys each domain serves, and each side's
pinned conduct record is filed on the ledger:
5d3e62f1…/report.
The register, as a repository
The same measurements are published as a standalone, machine generated repository: mcp-conduct-register.
Nobody selects the rows there either. A script rebuilds the table from the public API once a day,
and the same run writes a
register.json
snapshot so an agent can read the register without parsing Markdown. It carries a CITATION.cff,
so the register can be cited the way a dataset is cited, and an
llms.txt
that states in plain words what the register is and, more importantly, what it is not.
Three ways in, none of which need us
Since 2026-09-04 the gate can be used without asking anyone at HORIZON SHIELD.
For the server you operate. Put {"allow_tool_call": true} at /.well-known/mcp-conduct.json on your
origin. Only the owner of an origin can place a file there, so the gate takes it as consent, measures
determinism on the public register with it, and writes into every verdict where it read it (gate 0.2.4).
Add a compensation block to your agent card (paid_by, referral_fee, listing_fee; the content is not
judged, only its absence) and POST /watch once. A row can then reach verified with no hand of ours involved.
For your CI. One step measures the server on every push and recomputes the verdict hash on the runner,
so the gate is never trusted:
wedjat-check-action
(uses: ogasurfproject-jpg/wedjat-check-action@v1). It fails the job on a measured failure and leaves
unmeasured conditions unmeasured; require and must_pass decide how strict that is.
For the agent that connects. mcp-conduct on npm
(zero dependencies) reads /is-verified before an MCP client connects and applies a policy you choose:
warn, measured (block only what was measured and did not pass), verified-only, or off.
verified is true or null, never false; not measured is never failed. Source:
mcp-conduct.
Stated plainly: as of 2026-09-05 the register holds our own servers and nobody else's. The doors are open; the first outside row has not walked through yet.
JIDEC: verify this project without trusting it
The verification process behind HORIZON SHIELD's results is published as a Bitcoin anchored, append only public ledger. You do not have to trust us: fetch the anchored bytes, hash them yourself, and check the timestamp.
Start here: https://ledger.horizonshield.dev/llms.txt
Ledger index: https://ledger.horizonshield.dev/ledger
Machine readable catalog (RFC 9727): https://ledger.horizonshield.dev/.well-known/api-catalog
Read only MCP endpoint: https://jidec.horizonshield.dev/mcp
One line is enough to check any entry:
curl -s "https://ledger.horizonshield.dev/ledger/5?format=raw" | shasum -a 256What this proves and what it does not is stated by the ledger itself at /health under transparency, including that OpenTimestamps has no RFC, ISO or eIDAS standing.
The previous hostnames, hs-ledger.oga-surf-project.workers.dev and hs-jidec-mcp.oga-surf-project.workers.dev, still answer and always will. Records already anchored to Bitcoin cite them, so retiring them would make past receipts unverifiable.
What the MCP server does
A homeowner commissioning construction work cannot reliably judge whether a quote reflects a fair price. This is a textbook credence good problem. This MCP server makes a third party fair price reference callable and verifiable by software, so an agent can check a number instead of trusting it.
Protocol: Model Context Protocol (MCP)
Transport: MCP over Streamable HTTP (JSON-RPC 2.0). The legacy SSE transport is not implemented; GET on /sse answers 405 sse_not_supported.
Endpoint:
https://mcp.horizonshield.devAccess: read only, no API key required
Data region: fair-price verdicts for Japan (JPY), built on the open JCCDB dataset (526,128 records: line items and observations); construction cost data for Japan (JCCDB observation layer) and the United States (USCCDB, the United States Construction Cost Database)
Tools: 30 (15 for fair price, verification and contractors; 15 for construction cost data)
Tools
Tool | Description |
| Returns the fair price range (min, avg, max), the overcharge danger threshold, unit, price trend, and field notes for a Japanese construction or renovation job. |
| Given a work name and a quoted price in JPY, judges it as fair, a bit high, or overcharge risk, and returns the gap from the average. |
| Returns a fair price as a tamper evident record with a SHA-256 hash, under the PTKA (Pre-Transaction Knowledge Anchoring) model: a third party records the fair price before the contractor quote. |
| Checks whether wording in an estimate or sales pitch matches known overcharge or high pressure tactics (lump sum, today only discount, free inspection, door to door). Language agnostic. |
| Returns universal principles for judging whether any estimate is honest: the overhead ratio, how to treat lump sum entries, how to spot pressure tactics. Language agnostic. |
| Lists the construction and renovation work categories for which fair price ranges and red flags are maintained. |
| Returns the sources, update date, and regional multipliers behind the fair price data. |
| Returns metadata, scale, license, download links, and citation for the Japan Construction Cost Database (JCCDB). |
| Detects worry about an estimate and returns an invitation plus a submission URL to post it for third party review. |
| Finds a maintained cost category by work name or keyword. |
| Returns only the direction of a rough estimate versus the average (for example about +20 percent), before a detailed breakdown exists. |
| Independently recomputes a signed integrity verdict (SHA-256 over the signed_payload) as a third party. Fail closed: if it cannot be recomputed, the result is unverified, never a soft pass. |
| Issues a FairPriceAttestation shaped to attach to a Google AP2 (Agent Payments Protocol) Cart Mandate, so a fair price proof can ride alongside the payment authorization. Optional |
| Returns the A2A Agent Card URL and published skills for agent to agent discovery. |
| Finds verification-passed contractors on Yakumo, where listing depends only on passing the KIRA fairness audit and no referral or listing fee is taken. Scores and tiers, never prices; returns 0 honestly when nothing matches. |
Construction cost data: Japan (JCCDB) and both countries
Tool | Description |
| Searches the JCCDB line items (materials, products, labor) by name; returns whether each exists in a public document, with its evidence URL. |
| Region, date and price status of an item in Japanese and U.S. public documents. Values only where the licence allows redistribution; every row carries licence, attribution and evidence URL. |
| MLIT public-works design labor rates by prefecture and trade (wage per 8 hours); latest by default, yearly series with |
| Latest value per region for an item and spec, with min, median (computed) and max; only identical spec, unit and basis are compared. |
| Public-works unit prices for work items (materials, labor and equipment combined), with composition-ratio rows. Not renovation quote prices. |
| Construction cost index series (NHCCI, PPI, MLIT deflator and others) over a period, with year-over-year change computed by this service. |
| What the observation layers hold: rows per country, layer and source, priced rows and source periods; empty combinations are listed as absent. |
Construction cost data: United States (USCCDB)
USCCDB is the United States Construction Cost Database: U.S. public-domain federal data and city open data, one row per observation with source URL, sha256 and licence. The four chain tools compute on request and are not distributed as files. Public-works prices, statistics and estimates are reference data, not renovation quotes.
Tool | Description |
| U.S. public construction cost data by layer, region, period and item: Davis-Bacon wages, BLS wages, public unit costs, equipment rates, permits and spending, indexes and area factors, HUD cost limits, state DOT bid prices. |
| Davis-Bacon general wage determinations by state, county and trade: base wage and fringe with decision number and source URL. Minimums for federally funded work, not private market rates. |
| U.S. building permits by region and year: Census BPS and distributions of declared valuations in city permit data. Not contract prices. |
| DoD Area Cost Factors and USACE CWCCIS state adjustment factors by state, county, ZIP, city or overseas country. Budgeting factors, not a test of a quote. |
| Estimated U.S. prices along the distribution chain for construction materials and chemicals: landed import cost, wholesale, retail range and contractor, with formula, source URL and sha256 on every row. Computed on request. |
| Landed cost of U.S. imports by HS 10-digit code (Census IMDB): customs value, CIF, calculated duty including Section 232, unit cost, effective duty rate and top partner countries. |
| U.S. wholesale and retail gross margins by NAICS (Census AWTS, ARTS, AIES 2024) with kake_cost_ratio = 1 - margin; optionally the BEA 2007 margin structure. Industry averages. |
| Published discount rates off list price in U.S. public contracts (Washington DES, NASPO ValuePoint MRO) with kake_ratio = 1 - discount. Rates are ceilings; list bases differ by row. |
Connecting
This is a remote MCP server. Point any MCP client at the endpoint.
{
"mcpServers": {
"horizon-shield": {
"command": "npx",
"args": ["-y", "mcp-remote", "https://mcp.horizonshield.dev/"]
}
}
}If your client supports remote MCP servers directly, use the endpoint URL above.
In Claude, without configuration
The construction cost data server (https://ccdb.horizonshield.dev/mcp, fifteen read-only tools for JCCDB and USCCDB) is listed in the Claude connector directory after Anthropic's automated review: https://claude.ai/directory/connectors/horizon-shield-construction-cost-data . Open the page in Claude and press Connect; no key and no account with us. It is a community connector, which means it passed the automated review and is not verified by Anthropic.
Example
audit_estimate(work: "外壁塗装 30坪", quoted_price: 1500000)Returns a verdict (for example, overcharge risk), the fair range (min, avg, max), and the gap from the average. verify_fair_price additionally returns a SHA-256 fingerprint of the fair price claim, anchored under PTKA.
Verify a verdict yourself
Every verify_fair_price call returns a verify_url of the form https://shield.the-horizons-innovation.com/verify/?id=<claim_sha256>. The public verify page recomputes the SHA-256 in your own browser (Web Crypto) and checks it against the receipt. Nothing is sent to any server. The same claim is served back as JSON at https://mcp.horizonshield.dev/ledger/<claim_sha256>. Trust is conferred by recomputation, not assumed in the issuer.
Twenty real overcharge diagnoses are also published as tamper evident receipts, each with claim.txt, its SHA-256 digest, and an OpenTimestamps proof:
sha256sum claim.txt
ots verify -f claim.txt proof.otsIndex of the 20 receipts: https://shield.the-horizons-innovation.com/souba/kajou-seikyu-jirei-20/
The dataset these verdicts belong to is anchored at Bitcoin block 949356.
AP2 bridge
Google's Agent Payments Protocol (AP2) makes what a user authorized verifiable through a signed, tamper evident Mandate. create_ap2_fairness_attestation issues a parallel attestation that makes value verifiable, shaped to attach to an AP2 Cart Mandate before the user signs. Parallel layers, same philosophy: pre transaction, tamper evident, independently recomputable.
Data and academic record
Fair price data is built on the openly published JCCDB dataset (526,128 records: 95,403 Japanese construction line items and 430,725 source-cited observations, CC BY 4.0): https://github.com/ogasurfproject-jpg/japan-construction-cost-database
PTKA protocol declaration anchored at Bitcoin block 949356 (2026-05-14); JCCDB Extended paper at block 951871 (2026-06-01)
JCCDB origin paper: Zenodo 10.5281/zenodo.20019572
Audit hash and macro correction: SSRN 6738701, mirrored at engrXiv
VRQ framework and PTKA model: SSRN 6807738
Reproduction package (buyer side verification gate): GitHub, archived at Zenodo 10.5281/zenodo.20756867 (MIT, runnable:
node test/run_local.mjs)
Author
Toshikatsu Oga (大賀俊勝), The HORIZONs Co., Ltd., Hiratsuka, Japan. A carpenter of thirty years. ORCID 0009-0000-9180-903X.
"Cheapest is not the same as fair."
"Verify, don't trust."
"Thirty years on site taught me the enemy is the middleman, not the craftsman."
Full collection (50 quotes, JSON-LD): TOshi Oga, in his own words
Live diagnostic: https://shield.the-horizons-innovation.com · The Evidence: https://shield.the-horizons-innovation.com/evidence-en/ · The Movement: https://shield.the-horizons-innovation.com/movement-us/
License
Data: JCCDB, CC BY 4.0. Server code: see the LICENSE file in this repository.
Available Tools
15 toolsaudit_estimateAudit Estimate Against Fair PriceARead-onlyInspect
業者が提示した見積金額が適正かを、HORIZON SHIELDの適正レンジ(souba-db, 大賀俊勝 実務監修)と照合して判定する。手元に具体的な見積額がある時に使う。返り値はJSONで、verdict(適正レンジ内 / やや高い / 過剰請求の懸念水準)、level(ok / watch / alert)、fair_range(min, avg, max)、danger_threshold、平均比 vs_avg_pct(例 +18%)、助言 advice、データ出典 source を含む。工事名が見つからない場合、近い候補があれば did_you_mean として返す。単価(平米など)建ての工事に総額らしい金額を渡した場合は unit_mismatch の案内を返す。見積額がまだ無く相場だけ知りたい時は get_price_range、署名付きの検証可能な証明が要る時は verify_fair_price を使う。Japan only, JPY。 / Audits whether a contractor quoted price for a Japanese construction or renovation job is fair by comparing it against HORIZON SHIELD fair-price ranges (souba-db). Use when the user already has a specific quoted amount. Returns a JSON object with verdict, level (ok, watch, alert), fair_range (min, avg, max), danger_threshold, percentage gap versus the average (vs_avg_pct, e.g. +18%), advice, and data source. If the work name has no match, close candidates may be returned as did_you_mean. If the work is priced per unit and the amount looks like a total, a unit_mismatch notice is returned instead. For the typical range only use get_price_range; for a signed verifiable attestation use verify_fair_price. Trigger phrases: この見積もり高い?, 適正?, ぼったくり?, 妥当?, is this quote fair, am I being overcharged, is this a rip-off.
| Name | Required | Description | Default |
|---|---|---|---|
| work | Yes | 工事名(日本語)。材料やグレード込みで具体的に。例: 外壁塗装 シリコン。部分一致で照合するため曖昧だと別カテゴリにヒットしやすい。未マッチ時は近い候補が did_you_mean で返ることがある。 | |
| region | No | (任意) 地域。都道府県か市名(例: 神奈川県, 平塚市)か kanto/kinki/chubu/tohoku/other。渡すと地域係数を掛けたレンジで判定し、基準値も返す。 / (optional) Prefecture, city, or region key. The verdict then uses the regionally adjusted range; base values are returned too. | |
| quoted_price | Yes | 業者提示の金額(円, 数値)。一式見積はその総額。税込/税抜は正規化せず、渡した数値をそのまま適正レンジと照合する。 |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | No | How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true. |
| level | No | ok / watch / alert |
| advice | No | 助言 |
| lookup | No | ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist. |
| verdict | No | 判定 |
| fair_range | No | min/avg/max |
| vs_avg_pct | No | 平均比(例 +18%) |
| source_read | No | true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it. |
| did_you_mean | No | Near matches, when an exact match was not found. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint already covering safety, the description adds substantial behavioral context: full return payload (verdict, level, fair_range, danger_threshold, vs_avg_pct, advice, source), edge-case handling (did_you_mean, unit_mismatch), and JP-only/JPY scope. Doesn't state rate limits or failure modes, but the fallback behaviors are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but earns its length: purpose, usage, return fields, edge cases, alternatives, and trigger phrases, all in two language blocks. Slightly repetitive across languages but well front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists yet the description still conveys verdict semantics and edge cases, plus sibling routing, JP-only scope, and trigger phrases. Complete for an agent to select and invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with rich per-parameter descriptions, so baseline 3 applies. The description adds minor complementary info (total vs unit mismatch, JPY, no tax normalization) already implied by schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (audit/判定) and resource (contractor quoted price vs HORIZON SHIELD fair-price ranges), and explicitly distinguishes from siblings get_price_range and verify_fair_price. Bilingual, front-loaded.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('手元に具体的な見積額がある時に使う') and names alternatives with their conditions: get_price_range for range-only, verify_fair_price for signed attestation. Trigger phrases reinforce context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_red_flagsCheck Estimate Red FlagsARead-onlyInspect
見積もりや営業トークの中の気になる表現(例: 一式, 今日だけ値引き, 訪問販売)が、過剰請求につながりやすい既知の手口に当たるかを判定し、警告と対処を返す。代表的な手口のみを判定する。 / Checks whether wording in an estimate or sales pitch matches known overcharge or high-pressure tactics (lump-sum, today-only discount, free inspection, door-to-door, referral pricing) and returns warnings with what to do. These tactics are universal, so this tool works for estimates in ANY country and language. Covers representative tactics only.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | 見積書や営業トークで気になった表現・項目 |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | No | How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true. |
| flags | No | 該当手口の配列 |
| input | No | 判定対象の文言 |
| lookup | No | ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist. |
| result | No | 件数の要約 |
| source_read | No | true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it. |
| did_you_mean | No | Near matches, when an exact match was not found. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description transparently discloses its limitation (covers representative tactics only) and claims universality across countries and languages. While it doesn't explicitly mention side effects, the read-only annotation already covers that, and the description adds useful behavioral context without contradicting annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and concise despite being bilingual. It includes the tool's purpose, examples, applicability, and a limitation in two compact sentences per language. No unnecessary verbosity or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Without an output schema, the description gives a high-level indication that it returns warnings and recommended actions, which provides sufficient context for a simple tool. However, it could be more explicit about the exact output format or structure, so slightly below maximum.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter 'text' is fully described in the schema as the expression/item of concern from an estimate or sales pitch, and the tool description reinforces this by referencing examples. The semantics are unambiguous and complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that the tool checks wording in estimates or sales pitches against known overcharge/high-pressure tactics and returns warnings with actions. It provides concrete examples (lump-sum, today-only discount, etc.) and explicitly notes it covers only representative tactics, which effectively distinguishes it from more comprehensive audit tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use for quick red-flag screening by stating it covers representative tactics only, but it does not explicitly compare with alternative tools (e.g., audit_estimate, verify_fair_price) or specify when to prefer this tool. The universal applicability note gives some context, but the guidance is mostly implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_ap2_fairness_attestationCreate AP2 Fairness AttestationAInspect
このツールは決済を開始・承認・実行しません。資産・通貨・暗号資産の移動も行いません。発行するのは適正価格の証跡だけです。呼び出すたびに公開台帳へ記録を1件追加するため読み取り専用ではありません。 / This tool does not initiate, authorize, or execute any payment, and does not move funds, currency or crypto assets. It only issues a price-fairness attestation. Each call appends one record to the public ledger, so it is not read-only. AP2(Agent Payments Protocol)対応エージェント向けのブリッジ。決済カート(Cart Mandate)に添付できる適正価格の証跡(FairPriceAttestation)を発行する。AP2のMandateは『ユーザーがこの支払いを承認した』ことを検証可能にし、この証跡は『その価格が適正である』ことを検証可能にする。認可の検証と価値の検証、二つは並列レイヤー。quoted_price を渡すと適正レンジ判定(within/above/below)も同梱する。証跡は SHA-256 と公開台帳と verify_url で誰でも再計算検証できる。 / Bridge for AP2 (Agent Payments Protocol) agents: issues a FairPriceAttestation that a shopping or payments agent can attach to a Cart Mandate before asking the user to sign. AP2 mandates make authorization verifiable; this attestation makes value verifiable. Parallel layers. Pass quoted_price for a fair-range verdict (within, above, below). Independently verifiable via SHA-256, a public ledger and a verify_url. Japan construction and renovation pricing, JPY.
| Name | Required | Description | Default |
|---|---|---|---|
| work | Yes | 工事名(例: 外壁塗装 30坪) | |
| merchant | No | (任意) 施工業者名。Cart Mandate 例示に反映するだけで判定には使わない。 | |
| quoted_price | No | (任意) カートに載せる予定の見積額(円, 数値)。渡すと適正レンジとの判定を証跡に同梱する。 |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | No | How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true. |
| lookup | No | ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist. |
| ap2_bridge | No | AP2との関係(認可の検証 x 価値の検証) |
| attestation | No | 証跡本体(subject, integrity) |
| source_read | No | true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it. |
| a2a_carriage | No | 証跡を CartMandate に添える規範的な位置(A2A Artifact の兄弟 DataPart、開放口は risk_data) |
| did_you_mean | No | Near matches, when an exact match was not found. |
| cart_mandate_example | No | CartMandate の構造例(非規範。contents は署名対象なので第三者証跡は入れない) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond annotations by explicitly stating that no payment or transfer is initiated and that each call appends one record to the public ledger, making the non-read-only side effect transparent. It also explains the verification mechanism (SHA-256, ledger, verify_url). There is no contradiction with the annotations: readOnlyHint=false and idempotentHint=false align with the stated append behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the most important safety information and is well organized, but it repeats the same non-payment and attestation concepts in Japanese and English. It contains more sentences than necessary for the additional value it provides, though the bilingual structure may be intentional.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the domain (Japan construction/renovation, JPY), the integration point (Cart Mandate before user signature), side effects (public ledger append), and verification mechanism. It provides enough context for an agent to call the tool correctly, and an output schema is present, so return-value details need not be in the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds context for quoted_price (fair-range verdict with within/above/below) and merchant (not used in judgment), but these are also present in the schema's field descriptions. It does not meaningfully augment what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it issues a FairPriceAttestation for AP2 agents and explicitly distinguishes this from payment execution. It clarifies that no funds move, which sharply separates it from any payment-related sibling. The parallel authorization/value framing makes the tool's role in the AP2 flow unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description tells the agent when to use the tool: before asking the user to sign a Cart Mandate, and it explains when to pass quoted_price to get a fair-range verdict. It does not explicitly name alternatives or state when not to use it, but the usage context is clear enough for an agent to select it among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_verified_contractorFind Verified Contractor (Yakumo)ARead-onlyInspect
地域と工事名で、Yakumo(検証を通った加盟店だけが並ぶ建設モール)の検証済み施工店を探す。掲載は KIRA 適正診断の通過だけで決まり(fail-closed)、紹介料・掲載料は受け取らない中立の名簿。金額は出さずスコアとティアで示す。検証手続き中の店は pending として別に返す。条件に合う検証済みの店が無い時は 0 件と正直に返す(名簿は小さい)。価格の照会(get_price_range / audit_estimate)の後に、施主が『どこに頼めばいい』『信用できる業者は』と聞いた時に使う。 / Finds verification-passed contractors on Yakumo, a directory where listing depends only on passing the KIRA fairness audit (fail-closed) and no referral or listing fee is taken. Returns scores and tiers, never prices; pending stores are returned separately; returns 0 honestly when nothing matches (the directory is small). Use after a price check when the user asks who to hire or which contractor can be trusted. Trigger phrases: 業者を探したい, どこに頼めば, 信用できる工務店, find a contractor in Japan, who should I hire.
| Name | Required | Description | Default |
|---|---|---|---|
| area | No | 地域(都道府県・市区町村、例: 平塚市, 神奈川県, 名古屋市)。 / Area: prefecture or city, in Japanese. | |
| work | No | 工事名(例: 窓 交換, 外壁塗装, 浴室)。 / Work name in Japanese. |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | No | How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true. |
| lookup | No | ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist. |
| stores | No | 検証済みの店(member_no, name, area, works, fairness_score, integrity_tier, profile_url) |
| source_read | No | true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it. |
| did_you_mean | No | Near matches, when an exact match was not found. |
| directory_size | No | 名簿全体の件数(掲載数と検証済み数) |
| pending_stores | No | 検証手続き中の店(スコア無し) |
| verified_count | No | 検証済みの件数 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only establish that this is a safe read on a closed set; the description goes well beyond them by disclosing the fail-closed listing rule, the no-referral/no-listing-fee neutrality, that pending stores are returned separately, and that a zero-match result is returned honestly rather than padded. These are exactly the behavioral traits an agent needs and cannot get from the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then edge-case behavior, then usage routing and triggers, so the ordering is sound. Length is inflated by full Japanese/English duplication, but that is defensible for a bilingual trigger surface rather than pure redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return-value documentation is not required, and the description still covers the notable return behaviors (scores/tiers, separate pending list, honest 0 count). Combined with clear invocation routing and a closed-world small-directory caveat, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both parameters (area, work) carry their own descriptions with examples, so the description adds no syntax or format detail beyond restating that lookup is by area and work name. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (find verified contractors on the Yakumo directory) and immediately scopes what makes the directory distinct: listing depends only on passing the KIRA audit, fail-closed, no fees. It also implicitly separates itself from the price-oriented siblings (get_price_range, audit_estimate) by declaring it never returns prices.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the precondition and the alternative path: use it after a price check when the user asks who to hire or which contractor can be trusted, and lists trigger phrases in both Japanese and English. It also states when-not: price lookup belongs to get_price_range / audit_estimate, not this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_agent_cardGet A2A Agent CardARead-onlyInspect
このサーバー(HORIZON SHIELD KIRA)はMCPツールだけでなく、A2A(Agent2Agent)のエージェントカードも公開している。外部のA2A対応エージェントから発見・連携したい場合の、エージェントカードURLと公開スキルの一覧を返す。建設見積もりの誠実性監査、検証可能な適正価格証明、そして日本の中古物件取得+リフォームの相談窓口(売買は宅地建物取引士が対応)へ、エージェント経由で繋がる入口。 / Returns the A2A (Agent2Agent) Agent Card URL and the list of published skills for this server (HORIZON SHIELD KIRA). Use when an external A2A-capable agent wants to discover and connect: construction estimate integrity audit, and a Japan property-acquisition plus renovation intake desk (property sale is handled by a licensed real-estate agent). This is how agents reach the desk over A2A.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| count | No | How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true. |
| lookup | No | ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist. |
| skills | No | 公開スキル配列 |
| source_read | No | true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it. |
| did_you_mean | No | Near matches, when an exact match was not found. |
| agent_card_url | No | エージェントカードURL |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, and the description adds meaningful behavioral context: the tool returns an AgentCard URL and a skill list, and clarifies the boundaries of related services (e.g., property sale is handled by a licensed real-estate agent). This exceeds the annotation baseline without contradicting it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the key purpose but is written in both Japanese and English, duplicating the same content. This lengthens it unnecessarily, though the structure is clear and information is presented in a logical order.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description provides enough context about what the tool returns and when to use it, and there is an output schema so return value details are not required. It could be slightly tighter, but it is complete for its narrow discovery purpose and includes additional useful context on related services.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema covers 100% of them (none). With no parameters to describe, a baseline of 4 is appropriate. The description does not need to elaborate on parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns the A2A Agent Card URL and the list of published skills, with a specific verb ('get'/'返す') and a distinct resource (the server's A2A agent card). It is distinct from sibling tools that audit estimates and check red flags, so there is no ambiguity about what this tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says to use the tool when an external A2A-capable agent wants to discover and connect to this server. It provides context for when it applies without giving formal 'when-not-to-use' exclusions, but that is sufficient given the tool's narrow, discovery-focused role.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_estimate_reading_guideGet Estimate Reading GuideARead-onlyInspect
受け取ったリフォーム・建設見積もりが適正かを見分けるための原則(諸経費の適正比率、『一式』表記の扱い、営業手口の見抜き方)を返す。30年の現場経験に基づく判断軸。 / Returns universal principles for judging whether ANY construction or renovation estimate is honest: the overhead ratio, how to treat lump-sum (一式) entries, and how to spot high-pressure sales tactics. Language-agnostic and works outside Japan. Based on 30 years of field experience.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| count | No | How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true. |
| lookup | No | ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist. |
| source_read | No | true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it. |
| did_you_mean | No | Near matches, when an exact match was not found. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, and the description adds meaningful context: the tool provides general knowledge content rather than a custom assessment, is applicable outside Japan, and is based on 30 years of field experience. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently front-loaded with the core purpose and content topics, and every clause adds value—coverage areas, language agnosticism, and basis of authority. The bilingual duplication is acceptable given the tool's context, and there is no wasted content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no parameters, an output schema is present, and the description fully explains the tool's scope and content. It covers what the tool returns, how general it is, and why it is trustworthy. No additional context is needed for an agent to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero parameters, the baseline is 4. The description appropriately focuses on what the tool returns without needing to explain input semantics. It also indirectly clarifies there is no filtering or user-input dependency by describing the output as universal.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource: 'returns universal principles for judging whether ANY construction or renovation estimate is honest.' It also lists concrete content areas (overhead ratio, 一式 entries, sales tactics), which distinguishes it from sibling tools like audit_estimate or check_red_flags that perform targeted checks rather than provide general principles.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description communicates clear context for use: it is a language-agnostic, universal guide applicable to any estimate, not a Japan-specific or single-estimate tool. However, it does not explicitly name alternatives or specify when not to use it, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_fair_price_sourcesGet Fair Price Data SourcesARead-onlyInspect
HORIZON SHIELDの相場データ(souba-db)の出典・更新日・地域係数を返す。価格の根拠を確認したい時に使う。 / Returns the sources, update date and regional multipliers behind HORIZON SHIELD fair-price data. Japan. Use to check the basis of a price.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| count | No | How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true. |
| lookup | No | ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist. |
| source_read | No | true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it. |
| did_you_mean | No | Near matches, when an exact match was not found. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true and destructiveHint=false, so the safe read nature is covered. The description adds that it returns 'sources, update date and regional multipliers' and notes 'Japan' context. It does not describe output format or any edge cases, but with such a simple read-only tool and a readOnlyHint, this is adequate but not beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with no wasted words. It is front-loaded with the main purpose, includes the Japanese context, and provides a usage hint. Excellent structure for its purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool is simple, has an output schema, and has annotations for safety, the description covers the essential context. It explains the tool's scope (HORIZON SHIELD data, Japan), what it returns, and when to use it. It could mention that no parameters are needed, but that is implied by the schema. For a zero-parameter tool with an output schema, this is quite complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and schema description coverage is 100% (trivially). The description adds meaning by explaining what data is returned, which is the only semantic needed. Since there are no parameters to describe, the description fulfills its role well.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns sources, update date, and regional multipliers for HORIZON SHIELD fair-price data. The verb 'get' and specific resource 'fair price sources' make the purpose clear. It does not explicitly differentiate from siblings like get_price_range or verify_fair_price, but the focus on metadata (sources, date, multipliers) is distinct enough.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: 'Use to check the basis of a price.' This indicates when to use it. It doesn't explicitly mention alternatives or when not to use it, but the sibling tools like get_price_range or verify_fair_price are distinct enough that the usage context is reasonably clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_jccdb_dataset_infoGet JCCDB Dataset InfoARead-onlyInspect
日本の建設費オープンデータベース(JCCDB)のメタデータ・規模・ライセンス・ダウンロードリンク・引用情報を返す。建設費の一次データ源を探している時に使う。 / Returns metadata, scale, license, download links and citation for the Japan Construction Cost Database (JCCDB), an open dataset of 526,128 records in v5.1: the item catalogue of 95,403 Japanese construction line items (43,090 verified + 52,313 extended, unchanged since v4.0) and 430,725 source-cited observations. Use when looking for a primary construction-cost data source.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| count | No | How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true. |
| lookup | No | ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist. |
| source_read | No | true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it. |
| did_you_mean | No | Near matches, when an exact match was not found. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds genuinely useful non-annotation context: this is a static catalogue snapshot (v5.1, item catalogue unchanged since v4.0) with a defined record scale, which tells an agent what freshness and coverage to expect. It does not describe latency, caching, or whether the citation/download links are stable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and usage statement are front-loaded, but the content is duplicated almost wholesale in Japanese and English, which doubles length without adding information for most agents. The English half also carries dense parenthetical statistics (43,090 verified + 52,313 extended) whose relevance to tool selection is marginal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description needn't explain return values, and it correctly focuses on what the dataset is and when to reach for it. For a zero-parameter, read-only info tool this is close to complete; only the relationship to sibling cost-data tools is left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4 and there is no parameter semantics for the description to clarify. Nothing in the description misrepresents or obscures the parameterless interface.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (returns) and a precisely scoped resource (JCCDB dataset metadata, scale, license, download links, citation), and even enumerates the concrete payload contents including record counts and version. An agent can immediately distinguish this dataset-level information tool from sibling price-query tools such as get_price_range, get_fair_price_sources, or list_cost_categories.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use when looking for a primary construction-cost data source" gives a clear triggering context. However, it never states when NOT to use it or points to the alternative siblings (e.g., get_fair_price_sources or get_price_range) for agents that already have a source and just want prices, so routing is left partly to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_price_rangeGet Fair Price RangeARead-onlyInspect
工事名・キーワードで、HORIZON SHIELDが実務監修する適正価格レンジ(最安min/平均avg/最高max)と、それを超えたら過剰請求を疑う危険水準(danger)、単位・価格動向・実務解説を返す。建設・リフォーム費用が適正か数値で確かめたい時に使う(例: 外壁塗装, 給湯器, ユニットバス, クロス)。 / Returns the fair price range (min, avg, max), the overcharge danger threshold, unit, price trend and field notes for a Japanese construction or renovation job. Japan-specific pricing in JPY. Use to numerically check whether a cost is fair. Trigger phrases: 相場, 適正価格, いくらかかる, 高い?, how much does this cost in Japan, is this price normal, what should I expect to pay.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | 工事名やキーワード(日本語) | |
| region | No | (任意) 地域。都道府県か市名(例: 神奈川県, 平塚市, 名古屋市)か kanto/kinki/chubu/tohoku/other。渡すと souba-db の地域係数を掛けた値と基準値の両方を返す。 / (optional) Prefecture, city, or one of kanto, kinki, chubu, tohoku, other. Applies the regional multiplier and returns base values alongside. |
Output Schema
| Name | Required | Description |
|---|---|---|
| work | No | 工事名 |
| count | No | How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true. |
| lookup | No | ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist. |
| fair_range | No | 適正レンジ |
| source_read | No | true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it. |
| did_you_mean | No | Near matches, when an exact match was not found. |
| danger_threshold | No | 危険水準 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and non-destructive, and the description adds that it returns not just a range but also a danger threshold for suspected overcharging, unit info, price trends, and field notes. It also discloses Japan-specific pricing in JPY and mentions regional multipliers when a region is given. It does not detail rate limits or data freshness, but given the annotations cover safety, this is solid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Bilingual, front-loads the core purpose, then examples, then an English summary and trigger phrases. It is appropriately sized for a tool with detailed return fields, though slightly repetitive between the Japanese and English sections.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not explain return values in detail, yet it still summarizes them. It covers purpose, usage context, and regional behavior. Minor gap: no explicit comparison to sibling tools like verify_fair_price, but otherwise complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already describes both the query and region parameters, including the regional multiplier behavior. The description adds no parameter syntax beyond what the schema provides, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb (returns) and resource (fair price range for Japanese construction/renovation jobs), naming the exact fields returned (min/avg/max, danger threshold, unit, trend, notes). It distinguishes itself from siblings like verify_fair_price or audit_estimate by being a lookup/reference tool that returns benchmark ranges rather than auditing a specific estimate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: use when you want to numerically check whether a construction or renovation cost is fair in Japan, with example categories and trigger phrases. It does not explicitly name alternative sibling tools or say when NOT to use it versus verify_fair_price or check_red_flags, but the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_cost_categoriesList Cost CategoriesARead-onlyInspect
HORIZON SHIELDが相場・赤旗(過剰請求の懸念点)を整備している建設・リフォーム工事カテゴリ(61種)の一覧を返す。 / Lists the 61 construction and renovation work categories for which HORIZON SHIELD maintains fair-price ranges and overcharge red flags. Japan-specific data.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| count | No | How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true. |
| lookup | No | ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist. |
| categories | No | カテゴリ配列(id, name, group, priority, red_flags) |
| source_read | No | true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it. |
| did_you_mean | No | Near matches, when an exact match was not found. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so safety is covered. The description transparently states the count (61), the data specificity (Japan-specific), and the nature of data (fair-price ranges and overcharge red flags). This adds value beyond annotations, though it doesn't detail pagination or return format (output schema exists).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, both dense with relevant information: the count, the data purpose, the region. No filler, no repetition of the title or schema. Front-loaded with the purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with 0 parameters and an output schema (which likely lists the categories). However, the description doesn't mention how the list is returned (detailed fields vs. names), or whether any filtering is possible. It is mostly complete for a simple list, but given the output schema exists, the description needn't explain return values. A score of 3 is fair: adequate but would be 4 with a hint about the output's shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters, so the baseline for parameter semantics is 4. The description doesn't need to add parameter meaning; it correctly describes the fixed output set. The schema covers everything, and the description complements it by specifying the scope.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states it lists the 61 construction and renovation work categories for which HORIZON SHIELD maintains fair-price ranges and overcharge red flags. This clearly distinguishes it from sibling tools like search_cost_category (search) and get_price_range (query range), both in verb and resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description is clear that this tool is for retrieving the category list. However, it does not explicitly state when not to use it or mention alternatives. Given the no-parameter nature and simple listing purpose, clear context is provided, but explicit exclusions are absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_reverse_estimatePreview Reverse EstimateARead-onlyInspect
リフォーム検討の初期段階向けのプレビューで、業者の概算が平均からどちらの方向にどの程度ずれているか(例: +20%高い方向)だけを返す。具体的な適正額(min/avg/max)や危険水準は返さない。手元に詳しい見積内訳がまだ無い段階での最初の一歩に向く。具体的な適正レンジが必要なら get_price_range、見積額の詳細診断は audit_estimate を使う。Japan only, JPY。 / A preview for early-stage renovation planning that returns only the direction of a contractor rough estimate versus the average (e.g. about +20% above). It does not return the specific fair range (min/avg/max) or danger threshold. Suited as a first step before a detailed breakdown exists. Use get_price_range for a typical range, audit_estimate for a detailed quote diagnosis.
| Name | Required | Description | Default |
|---|---|---|---|
| work | Yes | 工事名(日本語)。例: 外壁塗装 シリコン。部分一致で照合。 | |
| quoted_price | Yes | 業者提示の概算額(円, 数値)。 |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | No | How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true. |
| lookup | No | ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist. |
| source_read | No | true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it. |
| did_you_mean | No | Near matches, when an exact match was not found. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint=true, destructiveHint=false), the description discloses the tool's intentionally limited behavior: it does not return min/avg/max fair range or danger thresholds, only direction. This behavioral limitation is essential for setting correct expectations and is not visible in the schema or annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The bilingual description is somewhat long but each section earns its place: purpose, limitation, usage timing, alternatives, and locale constraints. It is front-loaded with the core behavioral result and avoids unnecessary flourish.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, a 100% documented input schema, and strong annotations, the description adds the remaining necessary context: early-stage suitability, exclusion of detailed outputs, alternative tools, and Japan/JPY scope. This is complete for an agent to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with detailed descriptions for both parameters, including examples and partial-match behavior for work. The description adds context about comparing against an average but does not materially extend parameter semantics beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's specific function: it returns only the direction of a contractor's rough estimate relative to the average (e.g., +20% above). It also explicitly distinguishes itself from siblings by noting what it does not return and naming alternatives like get_price_range and audit_estimate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is explicit: it is for early-stage renovation planning before a detailed breakdown exists. The description gives clear when-to-use guidance and names specific alternative tools for different needs, making it easy for an agent to select correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_cost_categorySearch Cost CategoryARead-onlyInspect
工事名・キーワードで建設費カテゴリを検索する(例: 外壁塗装, 浴室, 給湯器, 雨漏り)。該当カテゴリと整備済みの赤旗件数・優先度を返す。 / Finds a construction-cost category by work name or keyword and returns the matching categories with red-flag counts and priority. Japan-specific; a Japanese query works best (e.g. 外壁塗装 exterior painting, 浴室 bathroom).
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | 工事名やキーワード(日本語) |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | No | How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true. |
| lookup | No | ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist. |
| source_read | No | true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it. |
| did_you_mean | No | Near matches, when an exact match was not found. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is known. The description adds beyond that by specifying that it returns matching categories with red-flag counts and priority, and notes the Japan-specific nature. This gives the agent a clearer picture of the tool's output without contradicting annotations. A score of 4 acknowledges this added context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences (in both Japanese and English) that are front-loaded with the core purpose, followed by examples. There is no redundant filler, and the structure effectively conveys the key information in a compact form. Achieves a perfect score for conciseness and clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With only one parameter, no nested objects, and an existing output schema, the description is sufficiently complete. It covers the search behavior, examples, and return value highlights (red-flag counts and priority). The tool's simplicity and the presence of an output schema mean no additional context is necessary.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers 100% of parameters with a simple description (工事名やキーワード), so baseline is 3. The description adds value by providing concrete example queries (外壁塗装, 浴室, 給湯器) and clarifying that Japanese keywords work best, which helps the agent form effective queries. This enrichment justifies one point above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Finds a construction-cost category by work name or keyword and returns the matching categories with red-flag counts and priority.' It uses a specific verb ('finds') and identifies the resource (construction-cost category) plus what it returns. It also gives concrete examples (外壁塗装, 浴室) that help distinguish it from siblings like list_cost_categories, which presumably lists all categories without a search query.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: use this tool when you need to search by a work name or keyword rather than listing all categories. It provides examples of suitable queries and notes that Japanese queries work best, which gives practical guidance. However, it does not explicitly mention alternatives like list_cost_categories or state when NOT to use this tool, so it stops short of a perfect score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
suggest_ehnGet Anonymous Estimate Review Link (EHN)ARead-onlyInspect
見積もりを匿名で第三者レビューに出せる掲示板EHN(見積もりハッカーニュース)の案内文と投稿フォームURLを返す。投稿と一次解析は無料で、業者名や個人情報は掲載前に運営が伏せる。ユーザーが見積もりのセカンドオピニオンや相談先を求めた時に使う。 / Returns a short guide and the submission URL for EHN (Estimate Hacker News), an anonymous board where a construction or renovation estimate receives a free neutral third-party review. Personal and contractor names are redacted before posting. Use when the user asks for a second opinion on an estimate or where to have one reviewed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| count | No | How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true. |
| lookup | No | ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist. |
| board_url | No | 公開ボード |
| submit_url | No | 投稿フォーム |
| source_read | No | true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it. |
| did_you_mean | No | Near matches, when an exact match was not found. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so safety is covered. The description adds genuinely useful behavior beyond that: posting and initial analysis are free, and operators redact contractor and personal names before posting, which tells the agent what happens to user data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The bilingual pair means each point is stated twice, which adds length, but both versions are front-loaded with the return value and then the usage condition. Every sentence carries information; the duplication is defensible for a JA/EN audience.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail is not the description's burden. The description nevertheless covers what is returned (guide + submission URL), the anonymity/redaction policy, and the cost, which is everything an agent needs to invoke a zero-parameter informational tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, and the rule sets the baseline at 4 for that case. There are no inputs needing semantic explanation, and the schema coverage is already 100%.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: returns a guide plus submission URL for EHN, an anonymous estimate-review board, and defines what EHN is ('見積もりハッカーニュース' / Estimate Hacker News). An agent can distinguish this informational tool from action siblings like audit_estimate or check_red_flags without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger condition in both languages: use when the user asks for a second opinion on an estimate or where to have one reviewed. It does not name a competing sibling to route against, so it stops short of a 5, but the when-to-use is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_fair_priceVerify Fair Price (Signed Receipt)AInspect
工事の適正価格を、検証可能な形(算出内容のSHA-256ハッシュ付き)で返す。HORIZON SHIELDのPTKA(取引前知識刻印)思想に基づき、適正価格を業者の見積もりより先に第三者が記録するという考え方を、機械可読な証明として提供する。エージェントが価格の真正性を検証したい時に使う。 / Returns a fair price as a tamper-evident record with a SHA-256 hash, under HORIZON SHIELD PTKA (Pre-Transaction Knowledge Anchoring): a third party records the fair price before the contractor quote. Japan price data. Use when an agent needs to verify price authenticity.
| Name | Required | Description | Default |
|---|---|---|---|
| work | Yes | 工事名(例: 外壁塗装 30坪) |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | No | How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true. |
| lookup | No | ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist. |
| provenance | No | データ出典・監修・再計算手順 |
| source_read | No | true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it. |
| did_you_mean | No | Near matches, when an exact match was not found. |
| verification | No | claim_sha256, verify_url, ptka |
| fair_price_claim | No | 刻印対象の主張(JSON.stringifyしてSHA-256すると claim_sha256 になる) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations (readOnlyHint=false, idempotentHint=false) are minimal and don't clarify side effects. The description adds useful context about SHA-256 hashing and the PTKA concept, but it does not explicitly state whether the tool writes a record or is purely read-only, leaving ambiguity about potential side effects—especially since readOnlyHint=false implies it may not be read-only.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action ('Returns a fair price as a tamper-evident record...'), but it includes a lengthy PTKA philosophy explanation that is not strictly necessary for invoking the tool. The bilingual format adds length, making it less concise than optimal while still being structured and informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema is present and the tool has only one parameter, the description covers the essential aspects: what it does, the underlying concept, and when to use it. The Japanese price data mention and PTKA context provide useful background without needing to describe return values, which the output schema handles.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the only parameter ('work') is already described with an example. The tool description does not add any additional meaning or constraints for this parameter, so the baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: returning a fair price as a tamper-evident record with a SHA-256 hash, and explicitly mentions the intended use (verify price authenticity). It does not name a specific sibling alternative like the highest-tier example, but the unique focus on PTKA and hash-based attestation distinguishes it from generic price-lookup tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear usage context: "Use when an agent needs to verify price authenticity." This gives a concrete trigger for invoking the tool. However, it does not explicitly exclude cases where other tools (e.g., verify_integrity_claim, get_fair_price_sources) might be more appropriate, nor does it mention alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_integrity_claimVerify Integrity ClaimARead-onlyInspect
estimate-integrity-audit が発行した署名付きクレーム(signed_payload と claim_sha256)を、第三者として検証する。発行側 (verify_fair_price はPTKA価格の発行) とは責務が正反対で、デフォルト姿勢は不信・fail closed。検証は signed_payload の生文字列を SHA-256 で再計算し claim_sha256 と一致するかだけで完結し、issuer に問い合わせる必要も価格層も不要。判定は契約 0.3 の failure_reasons 準拠で、result(verified / partial / unverified)・failure_reason(stale_data / changed_scope / missing_evidence)・trigger(expired_declaration / changed_estimate_version / missing_receipt / unverifiable_chain)・recomputed_sha256・scope_check・audit_ruleset_recheck を返す。重要: verified は『この宣言が改ざんされていない』ことの証明であって『監査ルールが今も有効』である保証ではない(audit_ruleset_recheck は常に not_performed)。estimate_version を渡すと scope(見積もり内容が発行時から変わっていないか)も照合し、渡さない場合は scope_check:skipped を明示する。 / Verifies a signed integrity claim (signed_payload and claim_sha256) issued by estimate-integrity-audit, as an independent third party. Opposite posture to the issuing side: distrust by default, fail closed. Recomputes SHA-256 over the raw signed_payload string and checks it equals claim_sha256; no issuer contact and no price layer needed. Follows contract 0.3 failure_reasons. IMPORTANT: verified means the declaration is untampered, NOT that the audit ruleset is still valid (audit_ruleset_recheck is always not_performed). Pass estimate_version to also check scope (whether the estimate changed since issuance); if omitted, scope_check is skipped and stated explicitly.
| Name | Required | Description | Default |
|---|---|---|---|
| claim_sha256 | Yes | そのレスポンスの claim_sha256 (64桁16進)。 / The claim_sha256 (64-char hex) from the same response. | |
| signed_payload | Yes | 検証対象の署名付きペイロード(estimate-integrity-audit のレスポンスの signed_payload を生文字列のまま)。改変するとハッシュ不一致で unverified になる。 / The signed_payload string from an estimate-integrity-audit response, verbatim. Any change makes the hash mismatch and the result unverified. | |
| estimate_version | No | (任意) 呼び出し側が現在の見積もりテキストから算出した estimate_version (input_text の SHA-256 先頭8桁hex)。渡すと発行時の版と一致するか照合する。省略可。 / (optional) The estimate_version the caller computed from the current estimate text (first 8 hex of SHA-256 of input_text). If provided, scope is checked against the issued version. |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | No | How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true. |
| lookup | No | ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist. |
| result | No | verified / unverified |
| source_read | No | true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it. |
| did_you_mean | No | Near matches, when an exact match was not found. |
| failure_reason | No | stale_data / changed_scope / missing_evidence |
| recomputed_sha256 | No | 再計算ハッシュ |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond the readOnlyHint and openWorldHint annotations by detailing the fail-closed behavior, the SHA-256 recomputation, and the distinction between 'verified' (hash match) and 'audit ruleset still valid'. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is thorough but slightly redundant (e.g., 'fail closed' and 'no price layer needed' appear more than once). However, the structure is logical, covering purpose, method, and caveats, and each sentence adds useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having an output schema, the description provides essential behavior context: return fields (result, failure_reason, etc.), edge cases (scope_check skipped if estimate_version omitted), and the crucial caveat that verified does not imply audit ruleset validity. This makes the tool's behavior fully understandable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All three parameters are described with clear semantics: claim_sha256 as expected hash, signed_payload as verbatim string (with warning about mutation causing unverified), and estimate_version as optional for scope checking. This fully complements the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states it verifies signed integrity claims from estimate-integrity-audit and distinguishes itself from verify_fair_price by noting 'no price layer needed' and 'opposite posture'. This clearly separates it from the sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit guidance on when to use (when you have signed_payload and claim_sha256) and how to invoke (optionally pass estimate_version to also check scope). It also clarifies that no issuer contact is needed and that it fails closed, which helps the agent decide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
15 tool updates
- Removed
compare_jccdb_regions - Removed
get_jccdb_coverage - Removed
get_jccdb_index_series - Removed
get_jccdb_labor_rate - Removed
get_jccdb_observations - Removed
get_jccdb_work_unit_price - Removed
get_us_area_factor - Removed
get_us_construction_prices - Removed
get_us_contract_discounts - Removed
get_us_import_landed_cost - Removed
get_us_permits - Removed
get_us_prevailing_wage - Removed
get_us_price_chain - Removed
get_us_trade_margins - Removed
search_jccdb_items
30 tool updates
v1.0.11- Changed
audit_estimate5 fields changed- added
Input schema / properties / regionAdded value: +{ + "description": "(任意) 地域。都道府県か市名(例: 神奈川県, 平塚市)か kanto/kinki/chubu/tohoku/other。渡すと地域係数を掛けたレンジで判定し、基準値も返す。 / (optional) Prefecture, city, or region key. The verdict then uses the regionally adjusted range; base values are returned too.", + "type": "string" +} - added
Output schema / properties / countAdded value: +{ + "description": "How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true.", + "type": "number" +} - added
Output schema / properties / did_you_meanAdded value: +{ + "description": "Near matches, when an exact match was not found." +} - added
Output schema / properties / lookupAdded value: +{ + "description": "ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist.", + "enum": [ + "ok", + "absent" + ], + "type": "string" +} - added
Output schema / properties / source_readAdded value: +{ + "description": "true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it.", + "type": "boolean" +}
- Changed
check_red_flags4 fields changed- added
Output schema / properties / countAdded value: +{ + "description": "How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true.", + "type": "number" +} - added
Output schema / properties / did_you_meanAdded value: +{ + "description": "Near matches, when an exact match was not found." +} - added
Output schema / properties / lookupAdded value: +{ + "description": "ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist.", + "enum": [ + "ok", + "absent" + ], + "type": "string" +} - added
Output schema / properties / source_readAdded value: +{ + "description": "true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it.", + "type": "boolean" +}
- Added
compare_jccdb_regions - Changed
create_ap2_fairness_attestation7 fields changed- changed
Output schema / descriptionPrevious value: -"AP2 Cart Mandate に添付できる適正価格証跡。attestation(FairPriceAttestation)・cart_mandate_example・verify_url。 / FairPriceAttestation for an AP2 Cart Mandate with attachment example."New value: +"AP2 Cart Mandate 向けの適正価格証跡。attestation(FairPriceAttestation)・cart_mandate_example・a2a_carriage(規範的な添付位置=A2A の兄弟 DataPart)・verify_url。 / FairPriceAttestation for an AP2 Cart Mandate, carried as a sibling A2A DataPart." - added
Output schema / properties / a2a_carriageAdded value: +{ + "description": "証跡を CartMandate に添える規範的な位置(A2A Artifact の兄弟 DataPart、開放口は risk_data)" +} - changed
Output schema / properties / cart_mandate_example / descriptionPrevious value: -"添付位置の例示(非規範)"New value: +"CartMandate の構造例(非規範。contents は署名対象なので第三者証跡は入れない)" - added
Output schema / properties / countAdded value: +{ + "description": "How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true.", + "type": "number" +} - added
Output schema / properties / did_you_meanAdded value: +{ + "description": "Near matches, when an exact match was not found." +} - added
Output schema / properties / lookupAdded value: +{ + "description": "ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist.", + "enum": [ + "ok", + "absent" + ], + "type": "string" +} - added
Output schema / properties / source_readAdded value: +{ + "description": "true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it.", + "type": "boolean" +}
- Added
find_verified_contractor - Changed
get_agent_card4 fields changed- added
Output schema / properties / countAdded value: +{ + "description": "How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true.", + "type": "number" +} - added
Output schema / properties / did_you_meanAdded value: +{ + "description": "Near matches, when an exact match was not found." +} - added
Output schema / properties / lookupAdded value: +{ + "description": "ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist.", + "enum": [ + "ok", + "absent" + ], + "type": "string" +} - added
Output schema / properties / source_readAdded value: +{ + "description": "true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it.", + "type": "boolean" +}
- Changed
get_estimate_reading_guide1 field changed- added
Output schema / propertiesAdded value: +{ + "count": { + "description": "How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true.", + "type": "number" + }, + "did_you_mean": { + "description": "Near matches, when an exact match was not found." + }, + "lookup": { + "description": "ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist.", + "enum": [ + "ok", + "absent" + ], + "type": "string" + }, + "source_read": { + "description": "true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it.", + "type": "boolean" + } +}
- Changed
get_fair_price_sources1 field changed- added
Output schema / propertiesAdded value: +{ + "count": { + "description": "How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true.", + "type": "number" + }, + "did_you_mean": { + "description": "Near matches, when an exact match was not found." + }, + "lookup": { + "description": "ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist.", + "enum": [ + "ok", + "absent" + ], + "type": "string" + }, + "source_read": { + "description": "true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it.", + "type": "boolean" + } +}
- Added
get_jccdb_coverage - Changed
get_jccdb_dataset_info1 field changed- added
Output schema / propertiesAdded value: +{ + "count": { + "description": "How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true.", + "type": "number" + }, + "did_you_mean": { + "description": "Near matches, when an exact match was not found." + }, + "lookup": { + "description": "ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist.", + "enum": [ + "ok", + "absent" + ], + "type": "string" + }, + "source_read": { + "description": "true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it.", + "type": "boolean" + } +}
- Added
get_jccdb_index_series - Added
get_jccdb_labor_rate - Added
get_jccdb_observations - Added
get_jccdb_work_unit_price - Changed
get_price_range5 fields changed- added
Input schema / properties / regionAdded value: +{ + "description": "(任意) 地域。都道府県か市名(例: 神奈川県, 平塚市, 名古屋市)か kanto/kinki/chubu/tohoku/other。渡すと souba-db の地域係数を掛けた値と基準値の両方を返す。 / (optional) Prefecture, city, or one of kanto, kinki, chubu, tohoku, other. Applies the regional multiplier and returns base values alongside.", + "type": "string" +} - added
Output schema / properties / countAdded value: +{ + "description": "How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true.", + "type": "number" +} - added
Output schema / properties / did_you_meanAdded value: +{ + "description": "Near matches, when an exact match was not found." +} - added
Output schema / properties / lookupAdded value: +{ + "description": "ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist.", + "enum": [ + "ok", + "absent" + ], + "type": "string" +} - added
Output schema / properties / source_readAdded value: +{ + "description": "true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it.", + "type": "boolean" +}
- Added
get_us_area_factor - Added
get_us_construction_prices - Added
get_us_contract_discounts - Added
get_us_import_landed_cost - Added
get_us_permits - Added
get_us_prevailing_wage - Added
get_us_price_chain - Added
get_us_trade_margins - Changed
list_cost_categories4 fields changed- added
Output schema / properties / countAdded value: +{ + "description": "How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true.", + "type": "number" +} - added
Output schema / properties / did_you_meanAdded value: +{ + "description": "Near matches, when an exact match was not found." +} - added
Output schema / properties / lookupAdded value: +{ + "description": "ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist.", + "enum": [ + "ok", + "absent" + ], + "type": "string" +} - added
Output schema / properties / source_readAdded value: +{ + "description": "true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it.", + "type": "boolean" +}
- Changed
preview_reverse_estimate1 field changed- added
Output schema / propertiesAdded value: +{ + "count": { + "description": "How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true.", + "type": "number" + }, + "did_you_mean": { + "description": "Near matches, when an exact match was not found." + }, + "lookup": { + "description": "ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist.", + "enum": [ + "ok", + "absent" + ], + "type": "string" + }, + "source_read": { + "description": "true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it.", + "type": "boolean" + } +}
- Changed
search_cost_category1 field changed- added
Output schema / propertiesAdded value: +{ + "count": { + "description": "How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true.", + "type": "number" + }, + "did_you_mean": { + "description": "Near matches, when an exact match was not found." + }, + "lookup": { + "description": "ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist.", + "enum": [ + "ok", + "absent" + ], + "type": "string" + }, + "source_read": { + "description": "true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it.", + "type": "boolean" + } +}
- Added
search_jccdb_items - Changed
suggest_ehn4 fields changed- added
Output schema / properties / countAdded value: +{ + "description": "How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true.", + "type": "number" +} - added
Output schema / properties / did_you_meanAdded value: +{ + "description": "Near matches, when an exact match was not found." +} - added
Output schema / properties / lookupAdded value: +{ + "description": "ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist.", + "enum": [ + "ok", + "absent" + ], + "type": "string" +} - added
Output schema / properties / source_readAdded value: +{ + "description": "true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it.", + "type": "boolean" +}
- Changed
verify_fair_price4 fields changed- added
Output schema / properties / countAdded value: +{ + "description": "How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true.", + "type": "number" +} - added
Output schema / properties / did_you_meanAdded value: +{ + "description": "Near matches, when an exact match was not found." +} - added
Output schema / properties / lookupAdded value: +{ + "description": "ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist.", + "enum": [ + "ok", + "absent" + ], + "type": "string" +} - added
Output schema / properties / source_readAdded value: +{ + "description": "true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it.", + "type": "boolean" +}
- Changed
verify_integrity_claim4 fields changed- added
Output schema / properties / countAdded value: +{ + "description": "How many records matched. 0 means the source was read and nothing matched. It never means the source could not be read, that returns isError: true.", + "type": "number" +} - added
Output schema / properties / did_you_meanAdded value: +{ + "description": "Near matches, when an exact match was not found." +} - added
Output schema / properties / lookupAdded value: +{ + "description": "ok = the source was read and something matched. absent = the source was read and nothing matched. A source that could NOT be read never appears here: that returns isError: true and makes no claim about what does or does not exist.", + "enum": [ + "ok", + "absent" + ], + "type": "string" +} - added
Output schema / properties / source_readAdded value: +{ + "description": "true on every successful result. A failed lookup does not return a result at all, so this is never false, it is declared so a consumer can assert on it.", + "type": "boolean" +}
19 tool updates
v1.0.1- Changed
audit_estimate1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "description": "見積額の適正診断。verdict・level(ok/watch/alert)・fair_range・danger_threshold・平均比・助言・出典。 / Quote audit verdict with fair range and advice.", + "properties": { + "advice": { + "description": "助言" + }, + "fair_range": { + "description": "min/avg/max" + }, + "level": { + "description": "ok / watch / alert" + }, + "verdict": { + "description": "判定" + }, + "vs_avg_pct": { + "description": "平均比(例 +18%)" + } + }, + "type": "object" +}
- Added
check_red_flags - Added
create_ap2_fairness_attestation - Removed
fair_price_data_sources - Changed
get_agent_card1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "description": "A2Aエージェントカードの場所と公開スキル一覧。 / A2A Agent Card URL and published skills.", + "properties": { + "agent_card_url": { + "description": "エージェントカードURL" + }, + "skills": { + "description": "公開スキル配列" + } + }, + "type": "object" +}
- Added
get_estimate_reading_guide - Added
get_fair_price_sources - Added
get_jccdb_dataset_info - Changed
get_price_range1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "description": "適正価格レンジ(min/avg/max)・過剰請求の危険水準・単位・価格動向・実務解説。 / Fair price range with overcharge danger threshold.", + "properties": { + "danger_threshold": { + "description": "危険水準" + }, + "fair_range": { + "description": "適正レンジ" + }, + "work": { + "description": "工事名" + } + }, + "type": "object" +}
- Removed
how_to_read_estimate - Removed
jccdb_dataset_info - Changed
list_cost_categories1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "description": "整備済みの建設・リフォーム工事カテゴリ(61種)の一覧。 / The 61 maintained construction and renovation cost categories.", + "properties": { + "categories": { + "description": "カテゴリ配列(id, name, group, priority, red_flags)" + } + }, + "type": "object" +}
- Added
preview_reverse_estimate - Removed
red_flag_check - Removed
reverse_estimate_preview - Changed
search_cost_category1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "description": "工事名・キーワードに該当したカテゴリと、整備済み赤旗件数・優先度。 / Matched cost category with red-flag count and priority.", + "type": "object" +}
- Changed
suggest_ehn1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "description": "EHN(見積もりハッカーニュース)への案内文と投稿URL。 / Guide and submission URL for the EHN anonymous review board.", + "properties": { + "board_url": { + "description": "公開ボード" + }, + "submit_url": { + "description": "投稿フォーム" + } + }, + "type": "object" +}
- Changed
verify_fair_price1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "description": "検証可能な適正価格レシート。fair_price_claim(主張)・verification(claim_sha256, verify_url, PTKA)・provenance(出典)。 / Tamper-evident fair-price receipt with hash, verify_url and PTKA anchor.", + "properties": { + "fair_price_claim": { + "description": "刻印対象の主張(JSON.stringifyしてSHA-256すると claim_sha256 になる)" + }, + "provenance": { + "description": "データ出典・監修・再計算手順" + }, + "verification": { + "description": "claim_sha256, verify_url, ptka" + } + }, + "type": "object" +}
- Changed
verify_integrity_claim1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "additionalProperties": true, + "description": "署名済みクレームの第三者検証結果(fail closed)。result(verified/unverified)・failure_reason・recomputed_sha256・scope_check。 / Third-party verification result, fail closed.", + "properties": { + "failure_reason": { + "description": "stale_data / changed_scope / missing_evidence" + }, + "recomputed_sha256": { + "description": "再計算ハッシュ" + }, + "result": { + "description": "verified / unverified" + } + }, + "type": "object" +}
13 tool updates
- First observed
audit_estimate - First observed
fair_price_data_sources - First observed
get_agent_card - First observed
get_price_range - First observed
how_to_read_estimate - First observed
jccdb_dataset_info - First observed
list_cost_categories - First observed
red_flag_check - First observed
reverse_estimate_preview - First observed
search_cost_category - First observed
suggest_ehn - First observed
verify_fair_price - First observed
verify_integrity_claim
TDQS
Scored across 15 tools
The cluster of price tools (get_price_range, audit_estimate, preview_reverse_estimate, verify_fair_price, create_ap2_fairness_attestation) overlaps in intent, but descriptions explicitly cross-reference each other (e.g. 'for the typical range use get_price_range, for a signed attestation use verify_fair_price') which sharply reduces misselection. The remaining tools (category list vs. search, source metadata, red-flag check, contractor finder, guide, agent card) are clearly distinct.
Every tool uses consistent snake_case verb_noun form (list_cost_categories, get_price_range, audit_estimate, verify_integrity_claim, find_verified_contractor). Verb style is uniform and predictable throughout.
15 tools sits at the top of the well-scoped range and reflects a genuinely multi-layered domain (price data, auditing, tamper-evident verification, AP2 bridge, contractor discovery, editorial guide). A few niche tools (get_agent_card, suggest_ehn) are marginal but each maps to a real workflow step.
The surface covers the full audit lifecycle: category discovery, price lookup, quote auditing, preview, red-flag detection, cryptographic verification, attestation, and contractor referral. Minor gaps only — no cost-history/trend tool beyond price trend fields and no comparison across multiple quotes.
Maintenance
Related MCP Connectors
An MCP server that audits the fairness of construction and renovation estimates in Japan. Provides fair-price ranges, overcharge detection, and verifiable unit-cost data based on JCCDB (65,520 items across 402 categories, CC BY 4.0, DOI-backed).
Buyer-side fair-price checks for Japanese renovation quotes; cost data is on a separate server.
Japanese court-run real-estate auctions (BIT). 5 tools, ~1,480 active listings. CC BY 4.0.
Check contractor quotes for red flags and pricing risks. Free scan, paid reports, RFP tools.
Related MCP Servers
- AlicenseAqualityBmaintenanceReal Estate Listing - MCP server providing AI-powered tools and automation by MEOK AI Labs511 npm36 PyPI1MIT
- AlicenseAqualityAmaintenance| 地価トレンド予測 | 新宿区の5年後地価をAI予測。CAGR・投資シグナル付き | | 企業立地需要分析 | 名古屋市中区のオフィス・工場需要スコアを算出 | | ファミリー向け適性評価 | 横浜市西区の教育・安全・医療スコアを総合評価 | | ポートフォリオ最適化 | 東京・大阪・埼玉の3エリアに投資配分を最適化 | | What-If シナリオ分析 | 大阪市中央区で新駅開設シナリオを試算 | | 店舗出店適地評価 | 福岡市博多区の人流・商業施設・交通データで出店適性を判定 |3850 npm1AGPL 3.0
- FlicenseNot gradedqualityBmaintenanceCheck if a contractor's remodeling bid is fair — analyze a quote (fairness score + red flags), get 2026 cost estimates by city, and look up BLS trade labor rates.-
- AlicenseNot gradedqualityBmaintenanceEnables agents to turn customer conversations into provably correct remodeling quotes by extracting catalog-bounded line items with evidence and calculating prices deterministically.MIT