Skip to main content
Glama

MARKET_AI_HUB

金融研究、預測模型與驗證工具,封裝成 MCP tools 給本機 AI Client / Agent 使用。 把 MARKET_AI_HUB 接到任何支援 stdio MCP 的 AI Client / Agent(Cherry Studio / Claude Desktop / Codex / 其他 MCP Host),AI 不用自己跑模型,而是透過 MCP 呼叫本機 Python 做研究。

Cherry Studio 只是其中一個使用例子,不是唯一或必備軟體。

研究用途,不是自動交易系統。 這不是 AI 聊天模型本身,而是金融研究後端。

Release:v2.0.0-rc1(Phase 2 Research Release Candidate)

第一次使用?從這裡開始


Related MCP server: TdxQuant MCP Server

系統定位

MARKET_AI_HUB 讓 AI Agent 透過 MCP(Model Context Protocol)取得:

  • 官方與 proxy 市場資料(Data Lake)

  • 模型預測(Prediction Registry / Model Tournament)

  • 情境分析(Joint / Scenario Forecast + Dynamic Ensemble)

  • 歷史策略研究(Historical Edge Store)

  • 結構化分析封包(Analysis Packet)

真正大阪研究 target = OSE Nikkei 225 Micro FuturesOSE_NIKKEI225_MICRO_FUTURES,不是 ^N225)。

  • MCP tools:21(runtime introspection)

  • Skills:3(osaka-micro-analysis / taiwan-stock-v28 / model-validation-audit)

  • 完整測試:649 passed(DESKTOP_SAFE,無 GPU)


現在能做什麼

能力

狀態

大阪微型日經研究(OSE Micro TARGET / settlement / volume / OI)

✅ AVAILABLE

台股分析(2330 / 3706.TW 等)

✅ AVAILABLE

get_analysis_packet(compact/normal/audit)

✅ AVAILABLE

模型預測(Chronos-2 / TimesFM / XGBoost / LightGBM)

✅ AVAILABLE

NHITS / NBEATSx(training-only,runtime blocked)

⚠️ BLOCKED

Model Tournament + Best Baseline

✅ AVAILABLE

歷史策略研究(Edge Store / walk-forward / cost)

✅ AVAILABLE

自動研究循環 + 受控自動學習

✅ AVAILABLE

官方資料 live(JPX settlement/volume/OI + BLS/Cboe/EDGAR/BOJ)

✅ PARTIAL

現在不能做什麼

能力

狀態

Live Trading / 下單

❌ PROHIBITED

Yuanta realtime recorder

❌ DISABLED

TradingView 依賴

⚠️ OPTIONAL(未安裝)

自動 Champion promotion

❌ DISABLED(需 human approval)

Foundation model 微調

❌ DISABLED(僅 interface)

宣稱可獲利策略

❌ PROHIBITED(歷史 OOS + 因果 audit 已驗證:無可執行 edge)


Research State(誠實結論)

Historical statistical signal does NOT imply executable trading edge.

層級

結論

Proxy(^N225)OOS

NO_EVIDENCE

Direct Micro 歷史 OOS

VAR(1) = STATISTICAL_FORECAST_EVIDENCE(MASE 0.958, direction 62.2%)

Executability(因果 audit)

NON_EXECUTABLE_FORECAST_EDGE(edge 在 overnight gap,需當日 full close,pre-close 失效)

Strategy(含成本)

NO_ECONOMIC_EDGE

Forward Shadow

NONE_YET(activated 2026-09-20,尚無真實交易日 evidence)

Strategy candidate

NONE

Production candidate

NONE

詳見 docs/VALIDATION_EVIDENCE.mdPHASE2_RESEARCH_FREEZE.yaml


架構圖

Official Data Sources → Smart Data Lake → Data Quality → Feature Store
  → Regime/Events → Direct Models → Model Tournament → Joint/Scenario
  → Dynamic Ensemble → Historical Edge → Strategy Research
  → Analysis Packet → MCP → Skills → AI Client

詳見 docs/ARCHITECTURE.md


快速連結

安裝 / 驗證

# 1. 建立環境
scripts\setup_windows.ps1

# 2. 下載 required models
python scripts\download_models.py --required

# 3. 驗證
pytest tests/

# 4. 啟動 MCP
.venv\Scripts\market-ai-mcp.exe
# 或
python -m market_ai_hub.mcp.server

Research Gates

模型需通過 ENGINEERING_GATE / DATA_GATE / MODEL_PREDICTIVE_GATE / TRADING_EDGE_GATETRADING_EDGE_GATE 預設 UNPROVEN(無 forward-validated edge)。

Compute Resource Governor

預設 DESKTOP_SAFE(不是 TRAINING_MAX)。未來任何 training / fine-tuning / Optuna / GPU batch 不得吃滿整台電腦。GPU VRAM ≤ 65%(soft)/ 75%(hard,保留 ≥4GB);CPU 保留 25% 給系統; RAM ≤ 65%/75%;process priority BelowNormal;同時間最多一個 GPU_HEAVY job。

AUTO_TRAIN=false / AUTO_FINE_TUNE=false / AUTO_PROMOTE=false

config/resource_profiles.yamlPHASE2VF_RESOURCE_GOVERNOR_REPORT.md

Security

NO LIVE TRADING / NO ORDER MCP / NO BROKER CREDENTIAL。見 SECURITY.md

Known Limitations

  • No validated executable edge(VAR forecast 有統計訊號,但 non-causal / non-executable)。

  • Forward evidence NONE_YET(activated 2026-09-20,需持續合法更新 Micro data)。

  • OSE Micro settlement history 不足(JPX 公開源僅當日,歷史 404;需 J-Quants/Data Cloud)。

  • 225LABO 為 LOCAL_ONLY center-month continuous dataset(非 contract-level,不得發布)。

  • OSE Micro per-contract OHLC:CONTRACT_ONLY(官方 xlsx/csv 無)。

  • BEA / e-Stat / EIA / EDINET:NEEDS_CONFIG(需 API key)。

  • Yuanta Futures SPARK 0112:OBSERVED UNRESOLVED_EXTERNAL。

  • Yuanta Legacy Quote T 盤:REQUIRES_SESSION_AWARE_RETEST。

  • NHITS / NBEATSx:runtime blocked(training-only)。

  • TradingView:OPTIONAL(免費 15 分鐘延遲)。

  • 詳見 CURRENT_HANDOFF.mdPHASE2_RESEARCH_FREEZE.yamldocs/DATA_SOURCE_MATRIX.md

Yuanta API(四條 family,勿混用)

元大 API 分成四條:SPARK(證券+期貨)、Futures Legacy Quote、Futures Legacy Trading、 Leveraged Trading「槓桿全球贏家」Web API(CFD,獨立槓桿帳戶)

⚠️ 不要到「槓桿全球贏家」API 申請頁(ltm.yuantafutures.com.tw/member/api-apply)申請一般 Futures API。 OSE Micro / JNU 是 JPX Futures,應走一般 Futures / SPARK 路徑。

docs/YUANTA_SETUP_AND_LOGIN.mddocs/YUANTA_API_ARCHITECTURE.mddocs/YUANTA_LEVERAGED_TRADING_API.md


V1 Freeze 歷史

V1 為 Freeze 快照(build_id bbf3cb2f9a80d20e)。歷史報告見 V1_*_REPORT.mddocs/history/(HISTORICAL V1 SNAPSHOT)。

Available Tools

21 tools
analyze_osaka_nikkeiB

大阪日經 PROXY 分析(^N225 index + 跨市場;非 OSE micro 即時)。

requested_dates(可選):"YYYY-MM-DD..YYYY-MM-DD"。^N225 依 TSE 日曆交易;
若窗口含 TSE 休市日(OSE futures 可能 Holiday Trading)→ CALENDAR_TARGET_MISMATCH。
ParametersJSON Schema
NameRequiredDescriptionDefault
horizonNo1d
requested_datesNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It usefully discloses that the tool is proxy-based, not real-time, and that TSE holidays can trigger CALENDAR_TARGET_MISMATCH. But it does not describe return format, latency, or potential side effects, so meaningful behavioral information is still missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded, stating the core purpose first and then adding only the most relevant edge-case warning. It is slightly dense with parenthetical jargon, but each sentence contributes meaningful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the proxy nature and the holiday-mismatch error, but it omits any explanation of what the analysis returns or what output the agent should expect. Since there is no output schema, this is a major gap, and the horizon parameter also remains unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description compensates for part of the 0% schema coverage by explaining requested_dates format ('YYYY-MM-DD..YYYY-MM-DD') and its calendar behavior. However, it gives no semantic guidance for the horizon parameter, leaving one of the two parameters underspecified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as a proxy analysis of the Osaka Nikkei using the ^N225 index with a cross-market scope, and it explicitly states what it is not: OSE micro real-time. This gives the agent a concrete resource and scope, though the generic term 'analysis' leaves some room for ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when the tool is appropriate by framing it as ^N225 proxy analysis and excluding OSE micro real-time data, and it warns about TSE holiday mismatches. However, it does not name alternative tools or give explicit when-to-use versus when-not-to-use guidance beyond that one exclusion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_taiwan_stockA

台股分析(如 2330 / 3706.TW / 華邦電)。只回傳 structured evidence。

ParametersJSON Schema
NameRequiredDescriptionDefault
stockYes
horizonNo1d

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It usefully discloses that output is strictly structured evidence, but it does not state whether the operation is read-only, how the analysis is produced, or whether there are latency, data, or side-effect considerations. This is minimal but non-empty behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely terse and front-loaded: it states the domain, supplies examples, and adds the key output constraint in two short clauses. There is no filler or redundant restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A minimal call is feasible because the required 'stock' parameter has examples and 'horizon' has a default. Still, the description does not explain what the structured evidence contains, what horizon values are permitted, or any prerequisites, and there is no output schema to fill those gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description compensates for the 'stock' parameter with format examples such as '2330', '3706.TW', and '華邦電'. The optional 'horizon' parameter is not explained beyond its schema default of '1d', leaving time-range semantics ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies Taiwan-stock analysis as the resource and provides concrete examples of accepted ticker formats. It stops short of specifying what kind of analysis is performed, but it is not tautological and is distinguishable from the Osaka/Nikkei sibling at a market level.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the Taiwan-stock domain and the examples, and the sibling analyze_osaka_nikkei suggests a geographic split. However, the description does not explicitly state when to choose this tool over get_market_data or the prediction tools, and it provides no exclusion conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

backtestA

walk-forward backtest(yfinance 資料 + baseline 分類器)。

輸出含分類指標(accuracy/balanced_accuracy/macro_f1/mcc)與 baselines
(uniform_random=1/3、majority_class_baseline、naive direction baseline)。
ParametersJSON Schema
NameRequiredDescriptionDefault
modelNolgbm
periodNo1y
symbolNo^N225
n_splitsNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It discloses the computation style (walk-forward), data source, and the exact output metrics/baselines, but it does not mention side effects, network/live-data dependency, execution time, or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense, front-loaded sentences cover the core behavior and output without redundancy. Every phrase earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with no output schema and no annotations, the description is incomplete: it omits parameter semantics, usage boundaries, and operational realities. An agent could call it with defaults, but customizing the invocation requires guesses.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description needed to explain model, period, symbol, and n_splits. It does not define any of them; the defaults are visible but their meaning must be inferred from the parameter names and the phrase 'walk-forward backtest.'

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('walk-forward backtest'), a data source (yfinance), and the classifier-based methodology. It also enumerates the output metrics and baseline values, making it easily distinguishable from the sibling prediction and analysis tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied by the name and purpose: run a walk-forward backtest on yfinance data. However, it does not explicitly state when to prefer this tool over siblings like run_ts_validation or predict_ensemble, nor does it mention when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_analysis_archive_statusB

Analysis Archive 狀態(不可變分析 / append-only outcome / reanalysis)。

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It offers domain terms like 'immutable analysis' and 'append-only outcome' but does not state whether the call is read-only, what kind of status payload is returned, or whether there are any side effects or dependencies.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single compact line with no filler, and it front-loads the resource and operation. It is slightly cryptic due to mixed language and jargon, but it earns its place by conveying the key conceptual categories succinctly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of an output schema, the absence of annotations, and a large sibling set, the description is too sparse to fully orient an agent. It does not specify what status data is returned, whether the status is a snapshot or live, or how this tool differs from similar archive/status tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and an empty schema, so there are no parameter details for the description to clarify. The baseline of 4 applies because nothing is left undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource ('Analysis Archive') and the operation ('Status'), and adds specific aspects (immutable analysis, append-only outcome, reanalysis) that go beyond the bare tool name. However, it does not explicitly contrast with nearby status tools like get_forward_test_status or get_data_source_status, so differentiation is implied rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus any of the many sibling tools, nor any context about what conditions warrant calling it. The only usage signal is the tool name itself, which is implicit rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_analysis_packetC

正式分析封包(backend 先完成大部分工作)。

market: osaka | taiwan。target 例:OSE_NIKKEI225_MICRO_FUTURES / 3706.TW。
horizon: 1d/2d/5d/10d。detail_level: compact | normal | audit。
^N225 只能是 PROXY/REFERENCE,不得當 execution target。
ParametersJSON Schema
NameRequiredDescriptionDefault
marketNoosaka
targetNoOSE_NIKKEI225_MICRO_FUTURES
horizonNo1d
detail_levelNocompact
save_analysisNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It does add some useful context ('backend 先完成大部分工作', and the N225 execution-target restriction), but it does not disclose side effects such as whether save_analysis=true persists data, what the return value contains, or any latency/caching implications. For a tool with no annotations Magnitude this is a notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact, scannable, and free of filler. The parameter lines are front-loaded and informative. The opening phrase is slightly tautological, but the added parenthetical about backend precomputation earns its place. Overall, every line contributes useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations according to outline and no output schema, the description should explain what an analysis packet is, what it returns, and what side effects may occur. It does not mention the save_analysis behavior curves or the format of the response, and it offers no differentiation from sibling analysis tools. The description is parameter-focused but contextually incomplete for a tool with this little structured metadata.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage curves, but the description compensates by enumerating valid values for market (osaka | taiwan), horizon (1d/2d/5d/10d), detail_level (compact | normal | audit), and giving a concrete target example. The only missing parameter is save_analysis, whose meaning is not addressed at all. This is strong added value over the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the tool as '正式分析封包' (official analysis packet) with a note that the backend does most work first, which implies a precomputed analysis retrieval. This is enough to know the general resource, but it does not clearly distinguish itself from siblings like get_official_release_snapshot, get_analysis_archive_status, or analyze_osaka_nikkei/analyze_taiwan_stock. It reads more like a label than a complete functional definition.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides parameter usage constraints, such as valid market/horizon/detail_level values and the warning that ^N225 can only be PROXY/REFERENCE, but it gives no guidance on when to use this tool versus alternatives. There is no explicit mention of situations where another sibling should be chosen instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_data_coverageA

大阪微型日經各 factor 資料覆蓋摘要(LIVE_VERIFIED/CONTRACT_ONLY/NEEDS_CONFIG/...)。不得隱藏缺口。

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description is the only source of behavioral context. It usefully discloses the status taxonomy and the guarantee '不得隱藏缺口' (must not hide gaps), which is a meaningful behavioral trait. However, it does not describe output shape, possible side effects, or any other operational behavior, so it is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single compact sentence that includes the scope, the status categories, and a key behavioral guarantee. There is no redundant wording or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple no-argument, no-output-schema tool, the description identifies scope and key statuses, which is largely sufficient. It could be more complete by explaining how to interpret or act on the coverage categories, but that is not essential for selecting the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the input schema is an empty object, so there is nothing for the description to add. Per the baseline for 0-parameter tools, this is appropriately handled.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool returns a data coverage summary for Osaka Micro Nikkei factors, with specific status categories like LIVE_VERIFIED and CONTRACT_ONLY. This is a clear verb+resource pairing. It does not explicitly distinguish itself from sibling tools such as get_data_source_status, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool instead of related siblings like get_data_source_status, get_market_data, or health_check. There is no explicit context or exclusion, leaving the agent to infer the appropriate use case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_data_source_statusB

各資料來源狀態(TWSE/FinMind/FRED/yfinance/JQuants/TradingView/Broker)。

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the full burden. It only says 'status of each data source' without explaining what status means, whether it checks connectivity/latency/last update, or whether it is read-only or potentially slow. The description adds little beyond the tool name.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, short sentence that names the resource and lists the covered sources. There is no redundant wording or filler, and the core meaning appears immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (no parameters, no output schema), the description is mostly sufficient: it tells the agent what entity is being reported on. However, it lacks any detail on the nature or format of the status values Mkay and does not clarify how this differs from similarly named sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and schema description coverage is 100%, so no parameter documentation is needed. The description appropriately focuses on the output scope rather than parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb+resource: 'get data source status' and enumerates the specific sources (TWSE, FinMind, FRED, etc.). It is unambiguous about what the tool returns, though it does not explicitly differentiate itself from siblings like health_check or get_data_coverage.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used when you want to know the status of the listed data sources, but it gives no explicit guidance on when to use it versus alternatives such as health_check or get_data_coverage. No exclusions or conditions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_event_calendarA

近期重要官方事件日曆(BOJ/Fed/CPI/NFP/PCE/GDP/MOF),只回 Top-N。

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo
top_nNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations at all, the description carries the full burden of behavioral disclosure. It adds one useful behavioral trait: only the top N events are returned, and it scopes the event categories. It does not discuss read-only guarantees, data freshness, or response shape, but those are less critical for a simple read-oriented calendar lookup.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one compact sentence that front-loads the resource, scopes the event categories in parentheses, and ends with the output rule. Every element earns its place and there is no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter read-only tool, the description covers the core purpose and the Top-N limitation. However, it lacks parameter detail, does not describe the output format, and does not disambiguate from get_official_release_snapshot. It is adequate but leaves several practical details to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage and provides only names, types, and defaults. The description's 'Top-N' hints at the top_n parameter and 'recent' loosely maps to days, but neither parameter is explicitly defined and the defaults are not restated. Some meaning is added, but not enough to fully compensate for the missing schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies a tool that retrieves a calendar of recent important official events and enumerates the covered categories (BOJ/Fed/CPI/NFP/PCE/GDP/MOF). It also notes the Top-N return behavior. However, it does not contrast itself with the closely named sibling get_official_release_snapshot, so it lacks explicit sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended usage is implied by 'recent important official events' and 'only returns Top-N', which suggests when an agent should query this tool. However, there is no explicit when-to-use, when-not-to-use, or alternative guidance, particularly against the similar get_official_release_snapshot.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_forward_test_statusB

Forward test 註冊狀態(Prediction Registry 中 joint/scenario/ensemble 預測是否已結算)。

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure, but it only states what status is queried. It does not mention that the operation is read-only, what the response format looks like, or any potential side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single compact sentence with no filler. It could be slightly clearer by aligning 'registration status' with 'settlement status', but it is efficiently sized and front-loaded with the core concept.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter status tool, the description conveys the essential meaning, including the specific predictions covered. However, with no output schema and no annotations, it leaves the response shape and exact status values to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. There is no parameter semantics gap for the description to fill, and the schema already confirms no arguments are required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies a specific resource—forward test status in the Prediction Registry—and clarifies that it reports whether joint/scenario/ensemble predictions are settled. It is distinguishable from sibling status tools, though it lacks an explicit verb and uses the slightly inconsistent '註冊狀態' alongside '是否已結算'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus alternatives such as get_data_source_status, get_analysis_archive_status, or get_target_instrument_state. The usage context is only implied by the tool name and resource reference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_market_dataC

取市場資料(yfinance symbols,或數字代碼走台股流程)。僅回傳結構化資料。

ParametersJSON Schema
NameRequiredDescriptionDefault
periodNo6mo
symbolYes
intervalNo1d

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It does state that only structured data is returned, which is a useful constraint, but it does not disclose whether the operation is read-only, whether it has rate limits, what failure modes exist, or how the Taiwan-stock process differs behaviorally. This is too minimal for a tool with no annotation support.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short and front-loaded, with no filler. Every clause adds at least some information: the resource, the routing nuance, and the structured-only return constraint. It earns a high conciseness score, though it is arguably too terse to fully compensate for the missing schema context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has no annotations and no output schema, so the description must provide most of the context. It explains the symbol routing and the structured return type, but it does not describe valid periods or intervals, what the structured data contains, or any prerequisites or edge cases. For a tool an agent must invoke correctly, this is incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds meaning for the symbol parameter by explaining yfinance symbols and numeric codes, but it does not clarify period or interval values, allowed formats, or how defaults behave. With three parameters and no schema descriptions, the coverage is insufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool retrieves market data, with a specific distinction between yfinance symbols and numeric codes routed through the Taiwan stock process. This gives a concrete verb and resource, and the Taiwan-stock note helps separate it from more generic sibling tools, though it does not explicitly name any sibling it is not.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to prefer this tool over alternatives such as analyze_taiwan_stock or get_data_source_status. The only usage signal is the routing rule for numeric codes versus yfinance symbols, which is a helpful condition but does not provide when-to-use or when-not-to-use guidance relative to sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_model_leaderboardD

模型 leaderboard(tournament PerformanceStore,含 BEST_BASELINE)。

ParametersJSON Schema
NameRequiredDescriptionDefault
targetNo
horizonNo

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. The description does not state whether this is a read-only operation, whether it requires authentication, what data it returns, or any side effects. The mention of 'tournament PerformanceStore' and 'BEST_BASELINE' hints at internal data sources but does not explain behavior. For a tool with no annotations, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short (one sentence), which is concise, but it is under-specified rather than efficiently informative. The parenthetical '(tournament PerformanceStore,含 BEST_BASELINE)' adds jargon without explaining meaning. It is not front-loaded with the most important information because the core action is missing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 2 parameters, no output schema, no annotations, and a description that does not explain the operation, parameters, or return value. Sibling tools suggest a complex domain (model performance, backtesting, validation), but this description provides no context for an agent to select or invoke the tool correctly. The description is completely inadequate for a tool with undocumented parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain the 'target' and 'horizon' parameters at all. The description only mentions the leaderboard concept and BEST_BASELINE, which does not map to the parameters. With two parameters completely undocumented in both the schema and the description, the agent has no way to know what values to provide. This is a critical gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is a noun phrase in Chinese: '模型 leaderboard(tournament PerformanceStore,含 BEST_BASELINE)' which translates to 'model leaderboard (tournament PerformanceStore, including BEST_BASELINE)'. It identifies the resource (model leaderboard) but lacks a verb stating what the tool does (e.g., 'retrieve', 'list', 'get'). The parenthetical mentions a tournament PerformanceStore and BEST_BASELINE, which is cryptic and does not clarify the operation. It does not distinguish itself from sibling tools like get_model_performance or get_analysis_packet.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. The description does not mention any context, prerequisites, or exclusions. Sibling tools like get_model_performance and get_analysis_packet exist, but the description does not explain how this leaderboard tool differs from them. The only hint is the parenthetical about tournament PerformanceStore and BEST_BASELINE, which is too vague to serve as usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_model_performanceB

讀取歷史 backtest 績效(model 空白 = 全部)。

分類指標不再硬用 50% 門檻:baseline_threshold = max(majority_class_baseline, uniform_random_baseline)。
沒有資料的欄位回 null(不是 0)。
ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
modelNo

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description takes on the full burden of behavioral disclosure. It meaningfully discloses the classification threshold rule (baseline_threshold = max of two baselines) and the null-return behavior for missing fields, both of which go beyond what a generic read operation implies. It does not describe the return format, but the disclosed behaviors are substantial and non-obvious.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three short sentences with no filler. The core purpose is front-loaded, and each subsequent sentence adds distinct behavioral value: model blank handling, threshold logic, and null semantics. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers key filter semantics and two important behavioral quirks, but the absence of both annotations and an output schema means more burden falls on the description. It does not specify the response shape, how limit interacts with results, or how this tool relates to performance-related sibling tools. Enough for basic use, but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the 'model' parameter's empty-means-all behavior, which is valuable, but it gives no semantic guidance for the 'limit' parameter beyond its schema default. With only one of two parameters clarified, the coverage is incomplete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb-resource pair: '讀取歷史 backtest 績效' (read historical backtest performance). It clarifies an important scope condition ('model 空白 = 全部'), but it does not explicitly distinguish this tool from siblings like get_model_leaderboard or get_forward_test_status, so it falls slightly short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to choose this tool over alternatives such as get_model_leaderboard or backtest. It mentions the 'model blank = all' behavior, which is parameter usage rather than tool-selection guidance. No exclusions or alternative routing is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_official_release_snapshotB

官方 macro 來源狀態快照(LIVE_VERIFIED / NEEDS_CONFIG / CONTRACT_ONLY)。

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing behavior. It implies a read-only snapshot but does not state side effects, caching, authentication requirements, or the response format, leaving significant gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence that names the resource and the key status values with no wasted words. It loses a point because the extreme brevity omits clarifying context that would aid an agent unfamiliar with the domain.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, no-output-schema, no-annotation tool, the description is minimal but does state what the snapshot covers and the possible statuses. However, it omits what the returned object contains and why the statuses are meaningful, so it is only borderline adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the input schema is fully complete and there is nothing for the description to add. Per the rubric, the baseline for 0 parameters is 4, and the description correctly focuses elsewhere.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource (official macro source status) and action (snapshot) and lists the three status values, giving concrete scope. However, it does not differentiate from the similarly-named sibling get_data_source_status, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like get_data_source_status or health_check, nor any mention of prerequisites or exclusions. The single sentence offers no usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_research_gatesB

Research validation gates:ENGINEERING_GATE / DATA_GATE / MODEL_PREDICTIVE_GATE / TRADING_EDGE_GATE。

每次評估基於 current runtime(singleton 載入),帶 evaluated_at / build_id / evidence。
ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Since no annotations are provided, the description carries the full burden of behavioral disclosure. It usefully mentions the singleton-based current runtime and the returned fields evaluated_at, build_id, and evidence, but it does not clarify whether the operation is purely read-only, what side effects may occur, or how failures are presented.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately short and front-loads the core gate types before adding runtime details. Every sentence adds useful information, though the mixed punctuation and somewhat cryptic phrasing keep it from being maximally polished.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should clarify what the returned data represents. It names evaluated_at, build_id, and evidence but does not explain their types, possible values, or how gate results are encoded. This is sufficient for a simple zero-parameter read, but an agent still lacks a complete return contract.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and an empty input schema, so there is no parameter burden for the description to handle. The baseline of 4 applies; the description does not need to explain parameters it does not have.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource as research validation gates and enumerates the exact gate types (ENGINEERING_GATE, DATA_GATE, MODEL_PREDICTIVE_GATE, TRADING_EDGE_GATE). It is more specific than simply restating the tool name, though it does not explicitly distinguish itself from sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives some context by noting that each evaluation is based on the current runtime, but it does not explain when to prefer this tool over alternatives such as get_forward_test_status, get_model_leaderboard, or get_analysis_packet. No exclusions or alternative routing are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_system_infoB

Python / GPU / CUDA / models(role + task + dual status)/ ensemble / MCP tools / gates / build fingerprint。

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden for behavioral disclosure. It hints at the output surface by listing environment/components, but it does not say whether the tool performs live checks, whether it is read-only, how long it might take, or how the returned information is structured.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very compact and front-loaded with useful keywords, but it is a slash-separated fragment rather than a well-formed sentence. The mixed-language parenthetical '(role + task + dual status)' adds ambiguity and reduces readability even though each item likely earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters, invoking the tool is trivial, and the description gives a high-level list of what to expect: Python, GPU, CUDA, model status, ensemble, MCP tools, gates, and build fingerprint. However, there is no output schema and the description does not define key terms like 'dual status', 'gates', or 'build fingerprint', leaving the returned payload only partially specified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema is an empty object with zero parameters, so there are no parameter semantics for the description to add. The baseline of 4 applies because there is nothing to document beyond the schema, and the description appropriately focuses on result content instead.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The tool name 'get_system_info' clearly identifies the resource, and the description enumerates concrete content categories: Python, GPU, CUDA, models, ensemble, MCP tools, gates, and build fingerprint. This is more specific than a tautology and helps distinguish it from sibling health/status tools, though it lacks an explicit verb phrase stating what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus siblings like health_check, get_research_gates, or get_data_source_status. The description only lists content areas and gives no conditions, exclusions, or alternative routing, so an agent must infer the intended use from the name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_target_instrument_stateB

真正交易標的狀態(OSE_NIKKEI225_MICRO_FUTURES)+ Micro settlement + 角色標記。

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

無任何註解,描述承擔全部行為揭露責任。描述僅列出狀態內容,但未說明此操作是否為唯讀、是否需要權限、是否有副作用,或失敗情境。工具名含有 get 暗示唯讀,但描述本身未確認,對代理的風險評估幫助有限。

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

描述僅一句話,精簡無冗餘,核心資訊(標的、狀態項目)有呈現。但因過度精簡導致部分語意模糊(如「角色標記」的具體含義),結構上可接受,但前段「真正交易」語意略顯籠統。

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

對於 0 參數、無輸出 Schema 的工具,描述提供了基本標的資訊,但缺少使用時機、回傳格式和唯讀性質的說明。在兄弟工具眾多的環境中,描述未能充分幫助代理區分使用情境,屬於「最低可用但有明顯缺口」。

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

工具無任何參數,Schema 覆蓋率 100%,因此描述無需補充參數意義。基於 0 參數的基準,此維度不需扣分,描述中提到的 Micro settlement 和角色標記可視為回傳內容而非參數。

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

描述明確指出具體的資源(OSE_NIKKEI225_MICRO_FUTURES)和狀態的組成(Micro settlement、角色標記),能與其他 get_* 工具區分。雖然沒有明確的動詞,但工具名稱已含 get,整體目的仍屬清晰,且非單純重述名稱。

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

描述完全沒有說明何時應使用此工具,也未提及與任何兄弟工具的替代關係或排除條件。'真正交易標的'暗示某種區分,但未具體說明,代理無法從中判斷與 get_market_data 或 get_data_source_status 的分工。

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

health_checkA

整體健康檢查:models / providers / 環境 / build fingerprint。

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states what is checked but not whether it is read-only, what side effects or rate limits exist, or what the response format is. For a diagnostic tool, this is a notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that conveys the purpose and scope efficiently. Every word earns its place, and the enumeration of areas adds value without fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless health-check tool, the description adequately lists the components examined. It could mention the return format (e.g., a summary report), but given the simplicity and no output schema, the description is sufficient for an agent to understand the tool's purpose and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. The description adds no parameter details (correctly, since none exist), and the schema is empty with 100% coverage, so no compensation is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it performs an overall health check and enumerates the specific areas covered (models, providers, environment, build fingerprint). This distinguishes it from siblings like get_system_info which likely target a single subsystem.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives. The purpose implies it is the go-to for overall status, but it doesn't mention conditions for choosing a sibling like get_system_info for detailed system diagnostics.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

predict_chronosB

Chronos-2 多步預測。horizon "Nd" = N trading bars(1d/2d/5d/10d)。含 build fingerprint。

ParametersJSON Schema
NameRequiredDescriptionDefault
periodNo6mo
symbolYes
horizonNo1d

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden of behavioral disclosure. It adds some useful context: the tool produces multi-step forecasts, defines horizon semantics, and mentions that output includes a build fingerprint. However, it does not disclose the output format, whether it is a read-only operation, or any side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely compact and front-loaded with the core purpose. Every clause adds useful information—model identity, prediction type, horizon semantics, and build fingerprint—without any filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter tool with no output schema and no annotations, the description is incomplete. It omits usage guidance, parameter details for symbol and period, return structure, and any behavioral context beyond horizon and fingerprint. An agent would need to infer too much.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description explains only one of three parameters: horizon, including its format and allowed values ('1d/2d/5d/10d'). Symbol and period are not described, so the description only partially compensates for the missing schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states it is a Chronos-2 multi-step forecasting tool, which clearly identifies the resource (Chronos-2 model) and the action (multi-step prediction). It differentiates from siblings like predict_timesfm and predict_ensemble by naming the specific model, though it does not explicitly mention that a symbol must be provided.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to choose this tool over predict_timesfm, predict_ensemble, or other siblings. The mention of 'Chronos-2' implies model-specific usage, but there are no conditions, exclusions, or alternative recommendations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

predict_ensembleC

Chronos + TimesFM + XGBoost + LightGBM ensemble。

V1.1 語義:price_ensemble(PRICE_FORECAST 真 quantile)與
direction_ensemble(DIRECTION_CLASSIFICATION)分層;legacy 欄位標 legacy_research_only。
ParametersJSON Schema
NameRequiredDescriptionDefault
periodNo6mo
symbolYes
horizonNo1d

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden of behavioral disclosure. It does reveal that output is split into price_ensemble and direction_ensemble and that legacy fields are marked legacy_research_only, but it does not explain what a caller receives, how the ensemble combines inputs, or what 'true quantile' means operationally. The behavioral picture is partial and underspecified.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and avoids verbosity, but it is structurally fragmented: the first line states models, then a version marker introduces internal semantic labels. It is compact enough to read quickly, yet the organization feels like internal notes rather than a coherent tool spec.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description needs to explain parameters, return shape, and usage context; it addresses none of the parameters and only hints at output semantics. The mention of price_ensemble and direction_ensemble is useful, but it is far from sufficient for an agent to confidently construct arguments or interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain any of the three parameters: symbol, period, or horizon. Since the schema itself provides only names and defaults, the missing parameter semantics leaves the agent unable to know what values are valid or how they affect the forecast. The description does not compensate for the schema gap at all.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description implies a forecasting ensemble by naming Chronos, TimesFM, XGBoost, and LightGBM)Skip and by distinguishing PRICE_FORECAST from DIRECTION_CLASSIFICATION, but it never states a concrete verb like 'predicts' or 'returns'. It gives meaningful semantic hints about price vs. direction outputs, yet lacks the explicit resource-action framing needed to fully clarify what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to choose predict_ensemble over predict_chronos, predict_timesfm, or backtest, despite these siblings existing. The text mentions internal output categories but provides no conditions, exclusions, or alternative routing. This is essentially no usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

predict_timesfmC

TimesFM-3.0 多步預測(weights 非商業授權)。含 build fingerprint。

ParametersJSON Schema
NameRequiredDescriptionDefault
periodNo6mo
symbolYes
horizonNo1d

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

沒有提供任何註解,因此描述必須承擔所有行為揭露的責任。描述只提到權重非商業授權和包含構建指紋,但沒有說明預測行為的細節,如輸出格式、是否會修改數據、是否有速率限制或任何其他行為特徵。對於一個預測工具,這遠不足以讓代理理解其行為。

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

描述非常簡短,只有一句話,沒有冗餘,但其中「含 build fingerprint」這個資訊對代理選擇或呼叫工具幾乎沒有幫助,可能只是內部細節。簡短是好的,但缺乏實質內容,因此只是最低可接受。

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

這是一個預測工具,有3個參數(其中2個有預設值),沒有輸出模式,也沒有註解。描述應該解釋預測的性質、輸出是什麼、如何解讀結果,以及任何限制。描述僅提供模型名稱和授權資訊,對代理實際呼叫工具所需的資訊嚴重不足。

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

架構描述覆蓋率為0%,因此描述必須解釋每個參數的含義,但描述完全沒有提及任何參數(symbol、period、horizon)。代理無法從描述中得知這些參數的用途或格式,只能依賴架構中的預設值,這對正確使用工具是嚴重的障礙。

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

描述明確指出這是使用TimesFM-3.0模型進行的多步預測,提供了具體的動詞(預測)和資源(TimesFM-3.0),並提及了權限限制。然而,它沒有與兄弟工具(如predict_chronos或predict_ensemble)做出區分,因此雖然清晰但未能說明與其他預測工具的差異。

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

描述完全沒有提供任何使用指引,沒有說明何時該用此工具而不是其他預測工具,也沒有提及任何條件、前置要求或替代方案。在有多個預測工具的環境中,這是一個嚴重的缺失。

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_ts_validationB

時間序列模型 rolling-origin OOS 驗證(V1.1 pipeline)。

- 多 rolling origins、no look-ahead、每 fold 預測下一 bar
- 比較 last-price naive / drift / moving-average baselines(MAE/RMSE/MASE)
- interval coverage:p10-p90 nominal 80% 的實測覆蓋率 + calibration error
- deterministic promotion(services/validation.PROMOTION_RULES),結果持久化
ParametersJSON Schema
NameRequiredDescriptionDefault
modelNochronos-2
periodNo1y
symbolNo^N225
n_foldsNo

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral disclosure burden and does well: it reveals no look-ahead behavior, per-fold next-bar prediction, baseline comparisons, interval coverage computation, deterministic promotion, and result persistence. It does not mention side effects of promotion or whether runs overwrite previous results, but it is substantially transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the main purpose and uses a tight bulleted structure. Every sentence adds distinct technical value, covering methodology, baselines, metrics, and persistence without filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description thoroughly covers the validation methodology and evaluation metrics, but there is no output schema, no return format, no parameter guidance, and no usage context versus siblings. For a 4-parameter tool with no annotations and no output schema, this leaves notable gaps despite strong methodological detail.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description needed to explain model, period, symbol, and n_folds, but it does not map any parameter to behavior. Parameter names and defaults are somewhat self-explanatory, which prevents a 1, but the description adds no explicit parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific operation: rolling-origin out-of-sample validation of time-series models, with concrete details such as multiple rolling origins, no look-ahead, next-bar prediction, and baseline comparisons. It does not explicitly differentiate itself from sibling tools like backtest or predict_*, so it stops short of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool instead of alternatives such as backtest, predict_chronos, or get_model_performance. The purpose is implied but no usage context, exclusions, or decision rules are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 21 tool updatesv0.1.0
    • First observedanalyze_osaka_nikkei
    • First observedanalyze_taiwan_stock
    • First observedbacktest
    • First observedget_analysis_archive_status
    • First observedget_analysis_packet
    • First observedget_data_coverage
    • First observedget_data_source_status
    • First observedget_event_calendar
    • First observedget_forward_test_status
    • First observedget_market_data
    • First observedget_model_leaderboard
    • First observedget_model_performance
    • First observedget_official_release_snapshot
    • First observedget_research_gates
    • First observedget_system_info
    • First observedget_target_instrument_state
    • First observedhealth_check
    • First observedpredict_chronos
    • First observedpredict_ensemble
    • First observedpredict_timesfm
    • First observedrun_ts_validation

TDQS

C2.9/5.0

Scored across 21 tools

Disambiguation3/5

Most tools target distinct resource/action pairs, but several status/snapshot tools overlap (health_check vs get_system_info, get_data_source_status vs get_official_release_snapshot vs get_data_coverage), and analyze_taiwan_stock/analyze_osaka_nikkei vs get_analysis_packet have unclear boundaries. The detailed descriptions help, but the tool set is not immediately unambiguous.

Naming Consistency4/5

The dominant patterns get_<noun>, predict_<model>, and analyze_<market> are consistent and readable. Minor deviations like health_check, backtest, and run_ts_validation break the pattern slightly but do not create serious confusion.

Tool Count3/5

21 tools is on the heavy side for an MCP surface, with several status/coverage/snapshot tools that could plausibly be consolidated. The broad hub scope makes the count defensible, but it is borderline and requires agents to absorb a large tool list.

Completeness4/5

The surface covers health, data ingestion, prediction, backtesting, time-series validation, market analysis, and monitoring/archive status. Minor gaps exist around triggering forward tests and retrieving raw prediction outputs, but the core workflows are represented without major dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    C
    quality
    D
    maintenance
    Provides comprehensive Taiwan stock market data and analysis through MCP tools. Enables querying real-time stock prices, historical data, company information, technical analysis, and market overviews for TWSE and TPEx listed companies.
    8
    16
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    Provides AI clients access to TdxQuant/通达信 financial data and trading capabilities through MCP. Enables retrieval of market data, financial reports, sector information, and trading operations via stdio or HTTP/SSE connections.
    29
    -
  • A
    license
    A
    quality
    A
    maintenance
    Enables AI agents to operate a local financial terminal, including market data, backtesting, paper portfolio management, and news digest, through safe, gated tools over MCP.
    6
    MIT