Japan Company Info — MCP Server (Free Edition)
This server enables offline search of a Japanese corporate due diligence knowledge base covering 192 large-cap listed companies in the Free edition.
Search company information via
search_knowledgeusing corporate numbers, securities codes, Japanese/English names, or GAAP concepts.Choose search method:
auto,bm25(exact keyword),trigram(fuzzy/identifier),vector(semantic), orhybrid(RRF fusion).Control result count with
top_k(default 5).Get the local Web UI URL via
get_web_ui_urlwhen GUI mode is enabled.Leverage integrated official data: NTA corporate IDs, EDINET XBRL financials, e-Stat benchmarks, and gBizINFO certifications/subsidies (full edition adds 財務省 benchmarks, major shareholders, and 3,193+ companies).
Support use cases like cross-border equity analysis, KYB/due-diligence entity verification, M&A screening, and reading Japanese filings in English with correct J-GAAP/US GAAP/IFRS mapping.
Runs offline after initial embedding model download (no Docker required).
Provides verified corporate data for Toyota Motor Corporation, including EDINET statutory filings with 5-year XBRL financials, its 13-digit National Tax Agency corporate number, major shareholders, and gBizINFO certifications/subsidies. Toyota records can be retrieved by corporate number (1180301018771), securities code (72030), or Japanese/English company name via keyword, fuzzy, semantic, or hybrid search.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Japan Company Info — MCP Server (Free Edition)search 72030 and show Toyota's operating income vs ordinary income in English"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Japan Company Info — MCP Server (Free Edition)
The only 100% OFFLINE, all-in-one Japanese Corporate Due Diligence MCP. Integrates 4 official government sources: NTA 13-digit Corporate IDs, EDINET XBRL financials (US GAAP / IFRS mapped), e-Stat industry benchmarks, and gBizINFO certifications. Zero cloud leakage. Zero Docker.
Query EDINET statutory filings (5-year XBRL), National Tax Agency 13-digit corporate numbers, major shareholders, and gBizINFO certifications/subsidies — with a built-in J-GAAP / US GAAP / IFRS mapping dictionary that correctly distinguishes 営業利益 (Operating Income) from 経常利益 (Ordinary Income, a J-GAAP-only concept).
This package is the Free edition: 192 large-cap blue-chip listed companies (EDINET-listed, Nikkei 225 sample). The full dataset — 3,193+ listed companies with complete XBRL financials, major shareholders, gBizINFO details, and 財務省 法人企業統計調査 industry benchmarks — is available as a one-time purchase at mcporb.store.
Use cases: cross-border equity analysis, KYB / due-diligence entity verification, M&A target screening, and reading Japanese filings in English without mistranslating 営業利益 (Operating Income) vs 経常利益 (Ordinary Income, a J-GAAP-only concept).
How it works
The server bundles the mcporb-runtime binary and a pre-indexed Orb (BM25 + trigram +
optional dense-vector retrieval). Retrieval runs on your machine. On first launch the
runtime downloads its query-embedding model (~220MB) in the background to enable the
semantic vector method; until that finishes — and forever after, offline — the bm25,
trigram, and auto methods work without any network access. Your queries are not sent to
a third-party API by this server.
It exposes five domain tools that normalize your request into a precise query, plus the
generic search_knowledge tool as a fallback for open-ended questions:
edinet_financials_usgaap(company_name, metric?, fiscal_year?)— EDINET statutory financials (P/L, balance sheet, cash flow) and executive compensation, with accounting concepts mapped across J-GAAP / US GAAP / IFRS.japan_corporate_registry(company_name, info_type?)— 13-digit National Tax Agency corporate number, registered address, legal status, and gBizINFO certifications / subsidies.japan_shareholders(company_name, top_n?)— major shareholders and ownership structure from the 大株主 section of 有価証券報告書.japan_industry_benchmarks(industry, metric?)— sector benchmarks (operating margin, ordinary margin, equity ratio, ROE) from 財務省 法人企業統計調査 via e-Stat.japan_company_search(query, method?, top_k?)— keyword / fuzzy / semantic search to discover a company when the target is unknown or ambiguous.search_knowledge(query, method?, top_k?)— raw knowledge-base search (fallback).method∈auto(default) ·bm25(exact keyword) ·trigram(fuzzy / identifier) ·vector(semantic) ·hybrid(RRF fusion).
The five domain tools resolve into search_knowledge internally, so retrieval and the
.orb capsule stay completely generic.
Related MCP server: EDINET DB MCP Server
Quick start
Claude Desktop
Add to claude_desktop_config.json:
{
"mcpServers": {
"japan-company-info": {
"command": "npx",
"args": ["-y", "japan-company-info-mcp-bridge"]
}
}
}(Before the package is published to npm, use the GitHub form:
"args": ["-y", "github:dqj1998/japan-company-info-mcp-bridge"].)
Cursor
Add the same server under Settings → MCP → Add Server (command npx, args as above).
Local (from a clone)
npx .Platform support
Bundled runtime binaries in bin/ (no external system libraries required — run on minimal/headless images):
macOS Apple Silicon (arm64) —
mcporb-runtime-darwin-arm64Linux x86-64 —
mcporb-runtime-linux-x64Linux arm64 —
mcporb-runtime-linux-arm64Windows x64 —
mcporb-runtime-win32-x64.exe
Not bundled: macOS Intel (x86-64). index.js resolves mcporb-runtime-<platform>-<arch> and exits
with a clear message if no matching binary is found.
Example queries
Most reliable retrieval is by corporate number or securities code (exact identifiers), then Japanese company name; English company-name search covers companies with an official English name (best-effort otherwise).
Intent | Example |
By corporate number |
|
By securities code |
|
By Japanese name |
|
GAAP concept |
|
Coverage note: this Free edition indexes 192 blue-chip companies. Queries for companies outside that set return the closest available matches; unlock the full 3,193+ company dataset at mcporb.store.
Data sources & attribution
法人番号公表サイト (National Tax Agency) — corporate registration
EDINET (Financial Services Agency) — 有価証券報告書 XBRL financials
gBizINFO (METI) — certifications, subsidies, commendations
財務省 法人企業統計調査 — industry benchmarks (full edition)
Data is redistributed under each source's terms; see NOTICE.
License
Bridge code is MIT (see LICENSE). The bundled Orb data and mcporb-runtime
binary are not MIT — they are licensed separately; see NOTICE.
Keywords: MCP · EDINET · Japanese GAAP · US GAAP · IFRS · Operating Income · Ordinary Income · Balance Sheet · Statutory Audit · Corporate Number · National Tax Agency · Due Diligence · Entity Verification · AML · KYB · gBizINFO · JSIC · Operating Margin · Industry Benchmark · Credit Risk
Available Tools
7 toolsedinet_financials_usgaapA
Retrieve official EDINET statutory financials (income statement, balance sheet, cash flow) and executive compensation for a Japanese listed company, with accounting concepts normalized across J-GAAP, US GAAP, and IFRS. Runs fully offline on the local machine — no data leaves the host; the requested metric is internally mapped to its Japanese accounting term before searching, and results return as ranked records (company name, 13-digit corporate number, fiscal year, and the requested line items with values in ¥ millions) sourced from 有価証券報告書 XBRL filings. This Free edition indexes 192 blue-chip companies, so a company outside that set returns the closest available matches rather than an error. Use this for a specific financial figure or statement of a named company (e.g. 'Toyota operating income 2023', '経常利益 in US GAAP'); use japan_company_search to discover an entity first, or japan_corporate_registry for identity and certification data. Identify the company by 13-digit corporate number or 4-digit securities code for the most reliable match; Japanese names resolve better than English.
| Name | Required | Description | Default |
|---|---|---|---|
| metric | No | Financial metric, as an English GAAP term or a Japanese accounting term (e.g. 'operating_income', 'net_income', '経常利益'). Known English terms are mapped to Japanese automatically; unknown terms are searched as-is. | |
| fiscal_year | No | Fiscal year, e.g. 2023. Omit to retrieve the latest available year. | |
| company_name | Yes | Japanese or English company name, 13-digit corporate number, or 4-digit securities code. A 13-digit number or Japanese name resolves most reliably. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so thoroughly: it discloses the offline/no-data-leaves-host execution, internal metric-to-Japanese mapping, output shape as ranked records with company name and corporate number and ¥ millions, XBRL source, the 192-company limitation, and closest-match rather than error behavior. No annotation is present, but the description gives the agent an accurate behavioral model.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded, with every sentence carrying useful information. It is a long single paragraph, which makes quick scanning slightly harder, but there is no redundancy or fluff; many sentences earn their keep.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter query tool with no output schema and no annotations, the description covers the input formats, the output format (ranked records, specific fields, ¥ millions), the data source, the limitation to 192 blue-chip companies, and the fallback behavior. An agent has enough to invoke and interpret the result correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the input schema already explains the metric mapping, fiscal_year omission, and company_name identifiers, including the reliability note. The description repeats those facts but does not add new parameter-level meaning, so it stays at the high-coverage baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Retrieve official EDINET statutory financials ... and executive compensation') and adds accounting-standard normalization context. It also differentiates itself from sibling tools by noting the 192-company index and the closest-match fallback behavior, so an agent can tell it apart from japan_industry_benchmarks and japan_shareholders without inspecting those tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to use it: 'Use this for a specific financial figure or statement of a named company', and directs to alternatives: japan_company_search for entity discovery and japan_corporate_registry for identity/certification data. It also gives identifier guidance (13-digit corporate number or 4-digit securities code) and sets expectations for out-of-index queries.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_web_ui_urlA
Get the local Web UI URL for this Orb when GUI mode is enabled
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It reveals the GUI-mode precondition but does not say what happens when GUI mode is off (error, null, empty string), nor whether any auth or state is required. For a simple read-only getter this is a moderate but real gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. The resource ('local Web UI URL for this Orb') is stated before the conditional clause, which is the right ordering.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-argument getter with no output schema, the description covers the essential facts: what it returns and the condition under which it is meaningful. The remaining gap is the unspecified behavior when GUI mode is disabled, which is minor for a tool this simple.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema has nothing to document and there is no parameter semantics burden on the description. Baseline of 4 applies for a parameterless tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Get the local Web UI URL for this Orb'. It is unambiguous about what is returned. The only sibling, search_knowledge, is unrelated, so no explicit differentiation is needed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'when GUI mode is enabled' gives an implied precondition for calling the tool, which is useful context. However, it never states what to do if GUI mode is disabled or what alternatives exist, leaving usage guidance partial.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
japan_company_searchA
Search across Japanese corporate registries and listed-company filings by name, securities code, corporate number, or free-text intent — the entry point when the target company is unknown or ambiguous. Runs fully offline and returns ranked candidate companies (name, corporate number, securities code, and a snippet) ordered by relevance; this Free edition indexes 192 blue-chip companies. The 'vector' and 'hybrid' methods trigger a one-time ~220MB embedding-model download on first use, while 'bm25' and 'trigram' always work offline. Use this to resolve an entity first (e.g. '半導体メーカー 東京', 'Tokyo semiconductor maker'), then call edinet_financials_usgaap, japan_shareholders, or japan_corporate_registry for details; 'trigram' suits codes and identifiers, 'hybrid' suits natural-language queries.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | A company name, 4-digit securities code, 13-digit corporate number, or free-text intent describing the company you are looking for. | |
| top_k | No | Number of results to return. Omit to default to 5. | |
| method | No | Retrieval method. 'auto' (default) picks the best available; 'bm25' exact keyword; 'trigram' fuzzy / identifier; 'vector' semantic (first use downloads the model); 'hybrid' fuses all rankers. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden and does so well: it states that retrieval runs fully offline, that 'vector' and 'hybrid' require a ~220MB one-time model download, that the Free edition is limited to 192 blue-chip companies, and that results are ranked by relevance. This gives the agent accurate expectations about limitations, side effects, and output shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: purpose, scope, return values, limitations, download caveat, and routing to sibling tools are all present without fluff. It is front-loaded with the core purpose and then layers behavioral and usage details logically.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter search tool with no output schema and no annotations, the description is complete. It explains what the tool returns, when to use it, how to choose methods, what limits apply, what network/download behavior to expect, and which sibling tools to call next. No critical context for invoking the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for all three parameters, so the baseline is 3 (schema carries the load). The description adds meaningful method guidance beyond the schema by explaining offline behavior, the download triggered by certain methods, and suitability of 'trigram' vs 'hybrid' for different query types. It doesn't add much for top_k, but the schema already documents it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific verbs and a resource ('Search across Japanese corporate registries and listed-company filings') and explicitly frames the tool as 'the entry point when the target company is unknown or ambiguous.' It names the return payload (ranked companies with name, corporate number, securities code, snippet), clearly distinguishing it from the detail-oriented sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance ('Use this to resolve an entity first') and names the follow-up alternatives exactly: edinet_financials_usgaap, japan_shareholders, and japan_corporate_registry. It also provides method-level selection advice: 'trigram' suits codes and identifiers, while 'hybrid' suits natural-language queries.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
japan_corporate_registryA
Look up the official identity of a Japanese company: its 13-digit National Tax Agency corporate number, registered headquarters address, legal status, and METI gBizINFO certifications, subsidies, and commendations. Runs fully offline and returns the corporate number, registered name and address, status, and any gBizINFO records found, sourced from 法人番号公表サイト and gBizINFO; this Free edition covers 192 blue-chip companies, so out-of-set lookups return the nearest matches. Use this for KYB and entity-verification questions and for 'who or where is this company' lookups; use edinet_financials_usgaap for financial figures, or japan_company_search for fuzzy discovery. A 13-digit corporate number or an exact Japanese name gives the most reliable match.
| Name | Required | Description | Default |
|---|---|---|---|
| info_type | No | Which facet to focus on: 'identity', 'address', 'certifications', 'subsidies', or 'commendations'. Omit to retrieve all available registry information. | |
| company_name | Yes | Japanese or English company name, or 13-digit corporate number. A 13-digit number resolves most reliably. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It transparently states that the tool runs fully offline, returns a defined set of fields (corporate number, registered name/address, status, gBizINFO records), names its data sources (法人番号公表サイト and gBizINFO), and explicitly warns that out-of-set lookups return nearest matches rather than failing. This level of candor about limitations and output composition exceeds what is typical for a lookup tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is substantial but each sentence earns its place: it front-loads the core function, then covers coverage/limitations, usage routing, and input reliability. It is not padded, though it could be tightened slightly (e.g., the data sources sentence is somewhat long). Overall it is well-structured and information-dense without being wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter lookup tool with no output schema, the description is remarkably complete. It covers what the tool returns, its data provenance, its coverage boundary, the reliability of different input types, and how it relates to sibling tools. An agent has everything needed to decide whether to invoke it and how to construct a correct call; nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented. The description adds meaningful value beyond the schema by clarifying that a 13-digit corporate number resolves most reliably and by explaining the purpose of info_type ('Which facet to focus on') with concrete facet examples. This reinforces the schema's semantics without redundancy, earning a score above the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Look up') and a precise resource ('official identity of a Japanese company'), enumerating the exact data points (13-digit corporate number, address, legal status, gBizINFO certifications, subsidies, commendations). It clearly distinguishes itself from sibling tools by explicitly naming edinet_financials_usgaap for financials and japan_company_search for fuzzy discovery, leaving no ambiguity about its scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage guidance: 'Use this for KYB and entity-verification questions and for 'who or where is this company' lookups' and directly points to alternatives ('use edinet_financials_usgaap for financial figures, or japan_company_search for fuzzy discovery'). It also discloses the coverage limitation (192 blue-chip companies) and advises the most reliable input format (13-digit number or exact Japanese name), which is exactly the kind of operational guidance an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
japan_industry_benchmarksA
Retrieve sector-level financial benchmarks for Japanese industries from the 財務省 法人企業統計調査 (Ministry of Finance Corporate Enterprise Statistics Survey): operating margin (売上高営業利益率), ordinary margin (売上高経常利益率), equity ratio (自己資本比率), and ROE (自己資本ROE) by industry and data year. Runs fully offline and returns the benchmark row(s) for the requested industry with the available metrics and their data year, sourced from the Ministry of Finance survey via e-Stat; this Free edition includes a subset of industries. Use this to benchmark a company's profitability against its sector — pair it with edinet_financials_usgaap to pull the company's own figures, then compare. Provide the industry name, preferably the Japanese 業種 label (e.g. '輸送用機械器具製造業', '純粋持株会社').
| Name | Required | Description | Default |
|---|---|---|---|
| metric | No | Optional benchmark metric to focus on: 'operating_margin', 'ordinary_margin', 'equity_ratio', or 'roe' (English or Japanese). Omit to return all available benchmark metrics. | |
| industry | Yes | Industry / sector name, preferably the Japanese 業種 label (e.g. '輸送用機械器具製造業', '情報通信業'). English sector names are searched as-is. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and discloses important behavior: 'Runs fully offline', 'returns the benchmark row(s) ... with the available metrics and their data year', and 'this Free edition includes a subset of industries'. It also states the data source and the searched-as-is behavior for English names. It doesn't cover failure modes, but the core behavior is well disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but not wasteful; each clause contributes metrics, offline behavior, scope limitation, workflow, or input guidance. The main verb and resource are front-loaded, and the later sentences add rather than repeat. It is one long paragraph rather than tight bullets, but it remains efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter read tool with no output schema and no annotations, the description covers the source, the exact metrics, the return content, the edition limitation, the intended companion tool, and parameter conventions. It doesn't describe the precise response shape, but the 'benchmark row(s)' statement gives enough context for invoking the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds value by mapping metric codes to Japanese labels ('operating margin (売上高営業利益率)') and by providing preferred Japanese industry examples beyond those in the schema. This helps an agent supply correct parameter values, especially for metric and industry naming.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Retrieve sector-level financial benchmarks for Japanese industries'. It names the exact survey source and enumerates the available metrics, making the tool's purpose unmistakable. It also distinguishes itself from company-specific siblings by explicitly framing it for sector benchmarking.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit use case: 'benchmark a company's profitability against its sector' and tells the agent to pair it with edinet_financials_usgaap for company-side figures. It lacks explicit 'when not to use' exclusions, but the context is clear enough to route an agent correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_knowledgeA
Search the Japanese corporate knowledge base with raw keyword, fuzzy, semantic, or hybrid retrieval — the flexible fallback for open-ended questions the domain tools do not cover. Runs fully offline and returns ranked text records (company registry fields, EDINET financials, major shareholders, gBizINFO certifications, and industry benchmarks) ordered by relevance; this Free edition indexes 192 blue-chip companies. The 'vector' and 'hybrid' methods trigger a one-time ~220MB embedding-model download on first use, while 'bm25' and 'trigram' always work offline. Prefer the domain tools (edinet_financials_usgaap, japan_corporate_registry, japan_shareholders, japan_industry_benchmarks, japan_company_search) for structured questions; use this for cross-cutting or exploratory queries. 'bm25' suits exact keywords, 'trigram' suits codes and identifiers, 'vector' suits paraphrases, and 'hybrid' fuses all rankers.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search query | |
| top_k | No | Number of results (default: 5) | |
| method | No | Search method (default: auto). 'auto': automatically picks the best available method(s). 'bm25': exact keyword match, best for precise term lookup. 'trigram': fuzzy/typo-tolerant character-level match. 'vector': semantic similarity search, best for conceptual or paraphrase queries. 'hybrid': fuses all available rankers via RRF, recommended for mixed queries. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full burden. It discloses that 'vector' and 'hybrid' trigger a one-time ~220MB embedding model on first use while 'bm25' and 'trigram' always work offline, and that this Free edition indexes only 192 blue-chip companies. It does not mention caps on top_k or pagination behavior, which amounts to an otherwise minor omission.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is about 60 words in a single paragraph with a clear topic, method details, domain-alternative guidance, and edge cases. A significant quote here, 'bm25' for exact keywords, is scoped correctly. It is a bit dense and could be split with a sentence dedicated to method guidance, but every sentence contributes valuable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description covers many things: the result type ('ranked text records'), the content fields (registry, EDINET, major shareholders, gBiz pays, benchmarks), the index limit of 192 companies, and download-required methods. The only gap is an explicit result size/pagination limit, but this is not essential for invoking the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All three parameters have descriptions in the input schema, so the baseline is 3. The description adds guidance on method selection – 'bm25' for exact keywords, 'trigram' for codes/identifies, 'vector' for semantic, 'hybrid' for mixed – which gives an agent further reason for choosing a method.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb and resource: 'Search the Japanese corporate knowledge base' with explicit method mention ('raw keyword, fuzzy, semantic, or hybrid retrieval'). It also frames itself as the 'flexible fallback for open-ended questions the domain tools do not cover' and names the sibling tools as alternatives, making its distinct place in the family clear without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use for the domains: 'Prefer the domain tools (edinet_financials_usgaap, japan_corporate_registry, japan_shareholders, japan_industry_benchmarks, japan_company_search) for structured questions; use this for cross-cutting or exploratory queries.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v1.2.0- Added
edinet_financials_usgaap - Added
japan_company_search - Added
japan_corporate_registry - Added
japan_industry_benchmarks - Added
japan_shareholders
2 tool updates
- First observed
get_web_ui_url - First observed
search_knowledge
TDQS
Scored across 7 tools
Each data tool targets a distinct resource — discovery, registry identity, financials, shareholders, and benchmarks — and search_knowledge is explicitly positioned as a fallback. The only real ambiguity is between japan_company_search and search_knowledge, since both share retrieval modes and search the same 192-company index.
Naming is readable but mixed: get_web_ui_url and search_knowledge are verb-first, the japan_* tools are noun phrases with the action sometimes at the end (japan_company_search), and edinet_financials_usgaap breaks the japan_ prefix pattern. The japan_ prefix gives a recognizable family, but verb placement is inconsistent.
Seven tools is well-scoped for a read-only company information server: discovery, identity, financials, shareholders, benchmarks, a general knowledge fallback, and a UI utility. No tool feels redundant and the set is small enough for an agent to choose among quickly.
Core workflows are covered: resolve a company, verify identity/KYB, pull financials, inspect shareholders, and benchmark against industry data. Minor gaps exist — no dedicated officer/board tool and no multi-year financial history — but agents can work around them or use search_knowledge as a fallback.
Maintenance
Related MCP Connectors
Remote MCP for Japan's EDINET DB — 3,800 listed companies' financials & filings (OAuth)
Cross-market (US/JP/KR) structured financials, segments, ownership & metrics, traceable to filings.
Japanese EDINET XBRL facts: bilingual labels, financials, screening. No prices. Signup required.
Search Japanese subsidies and public company data using J-Grants, gBizINFO, and EDINET.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceMCP server for Japan's TDnet (Timely Disclosure network). Search and retrieve timely disclosure documents from listed companies on Japanese stock exchanges.5Apache 2.0
- AlicenseAqualityCmaintenanceStructured financial data for ~3,800 Japanese listed companies from EDINET regulatory filings — financials, major shareholders, segments, executive compensation, and corporate history. Remote MCP over HTTPS with OAuth 2.0, free tier.131MIT
- AlicenseAqualityCmaintenanceLets AI assistants query normalized financial statements (P/L, B/S, C/F) of 3,634 Japanese listed companies from official EDINET filings, unified across J-GAAP, IFRS, and US GAAP with English keys. Zero setup: npx -y edinet-mcp.46 npmMIT
- FlicenseNot gradedqualityDmaintenanceProvides AI agents with comprehensive Japanese market intelligence through 27 MCP tools, covering corporate data, macroeconomics, financials, and environmental data from 14 integrated sources.-