mcp-server-gamenumerics
gamenumerics
Deterministic numeric engines for game design — exposed to your AI coding agent via MCP. Import xlsx config tables, audit growth curves, run battle/gacha simulations, and reverse-engineer the formulas behind game spreadsheets. Every numeric answer comes from zero-dependency pure-function engines — zero LLM guessing.
Who it's for: game designers and solo/small teams doing roguelike / deckbuilder balance work, spreadsheet-driven numerical design, or live-ops tuning audits — no engine integration required, just an xlsx export of your config tables.
Quick start (MCP server)
The fastest way to use these engines is through mcp-server-gamenumerics — a stdio MCP server exposing 17 read-only tools to Claude Code, Cursor, ZCode, or any MCP host:
npx -y mcp-server-gamenumerics{ "mcpServers": { "gamenumerics": { "command": "npx", "args": ["-y", "mcp-server-gamenumerics"] } } }Full per-host config (Claude Code / Cursor / ZCode), the 17-tool reference, and the performance baseline live in mcp-server/README.md.
Related MCP server: Excel MCP Server
What's inside
Path | What it is |
| 19 pure-function numeric engines: battle simulation, gacha probability + Monte Carlo, EHP×EDPS power model, growth projection, economy production/consumption, progression curves, level systems, difficulty mirroring |
| xlsx/spreadsheet normalization: dual-row header merge, column pattern detection (arithmetic/geometric), base×coefficient relation inference, Theil-Sen robust curve fitting |
| Tool definitions (JSON Schema + handlers) wrapping the engines — the read/audit surface |
| The npm-distributed stdio MCP server ( |
| xlsx → normalized JSON workspace import pipeline |
Tool surface (17 = 14 mapped + 3 meta)
import_xlsx · list_tables / read_table · battle_simulate / simulate_gacha / compute_power / power_curve / eval_formula / audit_column / infer_column_rule / infer_table_relation · grade_workspace / profile_table / infer_foreign_keys · list_workspaces / set_workspace / read_memory
Read-only by design — write tools stay behind the web workbench's human-in-the-loop confirm flow (separate product, not part of this repo).
Development
npm install
npm test # vitest: engine + table + mcp-server suites
npm run check # tsc --noEmit (lib + scripts)
npm run mcp:check # tsc for mcp-server (own tsconfig)
npm run mcp:build # esbuild bundle → mcp-server/dist/index.jsLicense
Available Tools
20 toolsaudit_columnARead-only
列审计(对账):用锚点公式重算指定列的每一行(变量 x=varColumn 该行值,i=行序),报告偏离公式超过阈值(默认 1%)的行——识别成长曲线的手调断点。先看 list_tables 的 columnPatterns 了解列的既有模式。
| Name | Required | Description | Default |
|---|---|---|---|
| table | Yes | 目标表名 | |
| column | Yes | 被审计的数值列 | |
| varColumn | Yes | 作为变量 x 的列(通常是等级/序号列) | |
| expression | Yes | 期望公式(x=varColumn 值,i=行序),如 100 + 5 * (x - 1) | |
| thresholdPct | No | 偏离阈值百分比(默认 1) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint=true annotation, the description fully explains the behavior: per-row recalculation using x=varColumn value and i=row sequence, a default threshold of 1%, and reporting only deviating rows. It also discloses a dependency on list_tables column patterns. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no wasted words: the first front-loads the purpose and algorithm, and the second appends a short actionable prerequisite. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter tool with no output schema, the description covers core semantics, default threshold, use case, and a prerequisite. It does not specify the exact shape of the returned rows (e.g., whether each row includes row index, original value, or deviation), but the combination of 'rows' and formula-variable explanation gives a reasonably complete picture.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaningful value by defining expression semantics (x and i, with an example), noting varColumn is typically a rank/sequence column, and specifying thresholdPct's default of 1%. This goes beyond the schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a column audit tool: it recomputes each row of a specified column against an anchor formula and reports rows that deviate beyond a threshold, with a specific goal of finding manual adjustment breakpoints. However, it does not explicitly contrast itself with similar siblings like recon_diff or eval_formula, so it lacks explicit sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a concrete prerequisite ('先看 list_tables 的 columnPatterns 了解列的既有模式') and an intended use case (identifying hand-tuned breakpoints in growth curves). It stops short of naming alternative tools or giving when-not-to-use guidance, so it is below the 5 level.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
battle_simulateARead-only
战斗模拟:玩家面板 × 敌人面板的对拼计算(确定性引擎,非 LLM 估算)。返回击杀刀数、被击刀数、咬合比 killRatio、评级(easy/balanced/tight/impossible)、期望 DPS;可选 Monte Carlo 胜率。用于改数值后评估战斗咬合变化。
| Name | Required | Description | Default |
|---|---|---|---|
| enemy | Yes | 敌人面板(来自 enemy/等级成长) | |
| player | Yes | 玩家面板(来自 hero/成长表 + 装备加成) | |
| simulations | No | Monte Carlo 局数(默认 0 不算胜率) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true; the description adds meaningful behavioral context: deterministic core engine, not LLM-estimated, and optional Monte Carlo stochasticity when simulations are provided. It also discloses the returned metrics, which is useful beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence that front-loads the core purpose, then lists outputs, optional behavior, and intended use case. No filler or redundancy; every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by listing all key return values and rating scale. It also covers the optional simulation parameter and the tool's use case, making it complete enough for an agent to select and invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters and defaults. The description adds the overall simulation semantics and maps simulations to 'optional Monte Carlo win rate,' but it does not add per-parameter meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific action and resource: deterministic player-vs-enemy battle simulation, explicitly distinguishing itself from LLM estimation and listing concrete outputs (kill slashes, received slashes, killRatio, rating, DPS, optional Monte Carlo win rate). This clearly separates it from siblings like simulate_gacha and compute_power.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context: 'used to evaluate combat balance after changing numbers.' It implies when to use it but does not explicitly name alternative tools or exclusion conditions, though the deterministic-vs-estimation distinction gives some routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compute_powerARead-only
战力计算(确定性引擎,EHP×EDPS 开方模型):输入属性面板(中文键如 攻击/体力/暴击率/暴击伤害),返回战力、EDPS、EHP 与各属性边际价值。用于对比改数值前后的战力变化。
| Name | Required | Description | Default |
|---|---|---|---|
| attrs | Yes | 属性面板,键可用中文(攻击/体力/暴击率/暴击伤害/攻速/防御)或英文 | |
| stdAtk | No | 满级标准攻击锚(默认取面板攻击值) | |
| stdSpeed | No | 满级标准速度锚(默认 1,攻速倍率口径) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already signals a safe read-only operation, and the description adds useful behavioral context by calling it a deterministic engine and explaining the calculation model. This goes beyond the annotation without contradicting it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence that front-loads the purpose and model, then covers input, output, and use case without wasted words. It is concise yet complete.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description helpfully lists the returned values (power, EDPS, EHP, marginal values). Combined with the schema and annotation, an agent has enough to call the tool correctly, though output structure and edge cases are not detailed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds example Chinese keys and the formula model, but the schema already documents attrs, stdAtk, and stdSpeed, including defaults, so no major semantic gap exists.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that the tool computes power using a deterministic EHP×EDPS square-root model, takes an attribute panel as input, and returns power, EDPS, EHP, and marginal values. This is specific enough to distinguish it from siblings like battle_simulate or power_curve.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says it is for comparing power before and after changing values, giving a clear intended context. However, it does not mention exclusions or alternatives such as battle_simulate for actual battle behavior.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
eval_formulaARead-only
确定性求值数值公式(四则/幂/括号/变量,支持中文变量名)。用于验证锚点公式、快速算数——数值永远由引擎算而非估算。
| Name | Required | Description | Default |
|---|---|---|---|
| variables | No | 变量取值,如 {等级: 60} | |
| expression | Yes | 公式,如 100 * 1.08 ^ (等级 - 1) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds meaningful behavioral context: evaluation is deterministic and values are always computed by the engine rather than estimated, which tells the agent the result is exact and reproducible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise clauses, front-loads the core function, and every sentence adds useful information. There is no filler or repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter formula evaluator with full schema coverage, the description is largely complete: it covers supported syntax, variable-name language, determinism, and intended use. It does not discuss error behavior or the exact return shape, but the numeric-result nature is strongly implied by '数值永远由引擎算而非估算'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds beyond the schema by specifying the supported formula grammar (four arithmetic operations, powers, parentheses) and explicitly confirming Chinese variable names are allowed. This helps the agent construct valid expressions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('求值') and resource ('数值公式'), and enumerates supported constructs (四则/幂/括号/变量) including Chinese variable names. It also names concrete uses—verifying anchor formulas and quick arithmetic—so an agent can distinguish it from sibling calculation/simulation tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use the tool: for verifying anchor formulas and quick arithmetic. It does not explicitly name alternatives or exclusion criteria, but the clear use-case framing gives enough context for an agent to select it among formula-related siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_tableARead-only
导出当前工作区的数值表为 lua record 数组模块(return { [1] = { Key = v, ... } })或 JSON 数组文本(format 缺省 lua)。中文列名自动转 ["列名"] 字符串键;columns 可选收窄导出列。产物超 1MB 拒绝并引导收窄。先 list_tables 拿表名。
| Name | Required | Description | Default |
|---|---|---|---|
| table | Yes | 表名(来自 list_tables) | |
| format | No | 导出格式(缺省 lua) | |
| columns | No | 导出列子集(可选;大表收窄产物) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses the precise output structure, the automatic conversion of Chinese column names to string keys, the optional column narrowing behavior, and the over-1MB rejection path. This gives the agent a detailed picture of what invoking the tool entails.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the primary action and output format, followed by key options, size-limit behavior, and prerequisite. There is no redundant or filler text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter tool with no output schema, the description covers output shape, default format, column-name handling, size-limit rejection, and how to obtain the required table parameter. An agent has enough information to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameter descriptions already document table, format, and columns. The tool description mostly restates these facts and adds no novel parameter-level semantics beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool exports the current workspace's table as a Lua record array module or JSON array text, naming the exact output formats. It implies distinction from read_table by focusing on serialized output, but it does not explicitly contrast itself with any sibling tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear usage context: get the table name from list_tables first, use the columns parameter to narrow large exports, and expect rejection with guidance when output exceeds 1MB. It does not state when to prefer this over alternatives or explicitly list exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
grade_workspaceARead-only
工作区分级(逆向接入第一步):按输入契约三级分档扫描全部表——A 规范(双行表头点分列名比例/列生成模式命中率达标)、B 半规范(单行中文表头)、C 裸表(英文驼峰/无中文语义表头),输出工作区判级、逐表分级、覆盖率与孤儿表初判(无外键候选列且无角色候选信号)。分级结果可用 save_structure 固化到 structure.json。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint: true, and the description is consistent, describing a scanning and output operation with no mutation. It adds valuable behavioral detail beyond the annotation by specifying exactly what it outputs (grade levels, coverage, orphan table judgment) and that it does not persist results itself (delegates to save_structure). This enriches the read-only context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense run-on sentence with multiple embedded clauses and technical criteria. While it front-loads the purpose, the excessive detail (e.g., specific header patterns) makes it harder to parse quickly. It could be split into clearer sentences without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters and no output schema, the description must carry the burden of explaining what the tool does and what it returns. It lists the outputs (workspace grade, per-table grade, coverage, orphan table detection) and the integration with save_structure. This is adequate for an agent to call it correctly, though it doesn't specify the exact data structure of the output.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and 100% schema coverage (vacuously), so the baseline is 4. The description doesn't need to explain parameters; it instead explains the logic and outputs, which is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: grading a workspace by scanning all tables and classifying them into A/B/C levels based on specific criteria. It also mentions the outputs (workspace grade, per-table grade, coverage, orphan table detection), making it highly specific and distinct from sibling tools like profile_table or infer_foreign_keys.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It positions itself as the 'first step of reverse integration', giving clear context on when to use it. It also references save_structure as a follow-up action, which helps an agent understand the workflow. However, it doesn't explicitly state when not to use it or mention alternative tools for other steps, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_xlsxA
从本地 xlsx 文件(绝对路径)导入并创建工作区:读取工作簿全部 sheet(或 sheets 参数指定子集),双行表头自动合并为「父.子」复合列名,数值列自动识别等差/等比/常数模式,装配为规范化 tables/*.json + workspace.json 并设为当前工作区。headerRows 缺省由启发式探测(结果随摘要回显,探测失误可显式指定 1/2 重试);同名工作区整体覆盖重建。mode=check 只做九规则格式检查(表头结构/类型合法性/外部工作簿引用/公式结构一致性——含公式矩阵 cell.f),不建工作区,返回 findings 清单。导入后即可用 list_tables / read_table / audit_column / export_table 等工具审计导出。
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | import(缺省)= 建工作区;check = 只跑九规则格式检查不建工作区(导表前体检) | |
| sheets | No | 导入 sheet 子集与覆盖(可选;缺省导入全部 sheet,module=import,headerRows 自动探测) | |
| xlsxPath | Yes | 本地 xlsx 文件绝对路径(Windows 正反斜杠均可,内部归一) | |
| workspaceName | No | 工作区名(可选;缺省取文件名去扩展名。仅允许字母/数字与 -._,首字符须为字母/数字) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries full weight for behavioral disclosure. It explicitly discloses the destructive behavior ('同名工作区整体覆盖重建'), the non-destructive check mode, the heuristic headerRows detection plus retry fallback, the auto-merging of two-row headers, and the side-effect of setting the current workspace. This is unusually transparent for a tool with no annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-organized, front-loading the primary purpose (import and create workspace) before detailing modes, parameter behavior, and post-import usage. Every clause carries useful information, though the long first sentence with many subordinate clauses makes it slightly harder to parse quickly than a bulleted structure would be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (two modes, heuristic behavior, destructive overwrite, normalized output artifacts), the description covers most of what an agent needs: input path, sheet selection, header handling, check-mode behavior, and follow-up tool suggestions. There is no output schema, and the description does not fully specify the return shape for the import mode beyond mentioning echo and findings; slightly more detail about the success response would make it complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds meaningful semantic context beyond the schema: it explains why headerRows matters (double-row headers become 'parent.child' composite names), what mode=check actually validates, and that same-name workspaces are overwritten. It doesn't add much about workspaceName or sheets beyond what the schema already states, but the added behavioral context raises it above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('import'), a concrete resource ('local xlsx file absolute path'), and the resulting artifact ('normalized tables/*.json + workspace.json'). It clearly distinguishes itself from the post-import sibling tools by noting those are usable after import, and it enumerates its two modes (import vs check) so an agent knows exactly what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: to build a workspace from an xlsx file, and mode=check as a pre-import validation step. It also names downstream tools (list_tables, read_table, audit_column, export_table) that should be used after import. It does not explicitly explain when not to use this tool or name direct alternatives, but the mode distinction provides solid practical guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
infer_column_ruleARead-only
列规则推断(确定性拟合,逆向分析的起点):对指定数值列拟合三大曲线族(等差/等比/幂律),返回最优规则——参数、吻合占比 fitPct、可直接使用的表达式——以及偏离规则的断点行(疑似手调)。分析成长/消耗曲线的构成规律先用它;得到规则后把 expression 交给 audit_column 复核,或用 write_table 的 apply_curve 按规则整列重算。
| Name | Required | Description | Default |
|---|---|---|---|
| table | Yes | 目标表名(来自 list_tables) | |
| column | Yes | 要推断的数值列 | |
| thresholdPct | No | 偏离阈值百分比(默认 1),超出即计为断点 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already covers safety; the description adds that fitting is deterministic and that the tool returns both the optimal rule and suspected manually-adjusted breakpoint rows. This gives an agent a clear behavioral model without contradicting the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but front-loaded with the tool's role and outputs, followed by usage workflow. Every sentence adds information, though the single long sentence with semicolons and the reference to an unlisted write_table slightly reduce readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description takes responsibility for explaining the return value (optimal rule with parameters, fitPct, expression, breakpoint rows), which it does. It also gives the follow-up workflow. It doesn't cover failure modes or non-numeric input, but for a read-only inference tool the essentials are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description restates that the column must be numeric and mentions breakpoint rows, but adds no meaningful detail about thresholdPct beyond the schema's '偏离阈值百分比(默认 1)'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific role ('列规则推断') and states the exact operation: fit three curve families (arithmetic/geometric/power) to a numeric column and return the optimal rule plus breakpoint rows. This clearly distinguishes it from related siblings like audit_column and power_curve by naming it the starting point for reverse analysis.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says to use this tool first when analyzing growth/consumption curves, then hand the expression to audit_column for review. This gives concrete workflow context. It does not explicitly state when not to use alternatives, and it references 'write_table' which is not present in the sibling list, slightly weakening the routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
infer_foreign_keysARead-only
表间外键推断(值域包含检测,确定性):from 列非空值域 ⊆ to 列值域、悬挂率 <5% 才成立,from 非空值 ≥3 防巧合;confidence = (1-悬挂率)×值域大小因子。输出候选关系边(kind=foreign_key,不落盘,供用户确认后经 save_structure 固化)。可指定 fromTable/toTable 定向检测;缺省全表两两扫描(输出按置信度排序的 topN 防爆炸)。
| Name | Required | Description | Default |
|---|---|---|---|
| topN | No | 全表扫描模式下输出候选数上限(默认 20,最大 50) | |
| toTable | No | 被引用方表名(与 fromTable 成对指定) | |
| fromTable | No | 引用方表名(与 toTable 成对指定) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark readOnlyHint=true, and the description adds substantial behavior: deterministic algorithm, exact acceptance conditions, confidence formula, no persistence, and output as unpersisted candidate edges for user confirmation. This goes well beyond the annotation baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence carries meaningful information: algorithm, thresholds, confidence formula, output behavior, and usage modes. No filler or redundancy; appropriately compact for a nontrivial tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only inference tool with no output schema, the description covers the algorithm, validity criteria, confidence calculation, output form, persistence behavior, and both invocation modes. An agent has enough context to call it correctly and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all three parameters. The description adds extra semantic value by explaining topN's purpose as preventing output explosion and clarifying fromTable/toTable as a paired directional-detection mode.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: inferring foreign keys between tables using value-domain containment detection, and specifies the output as candidate relation edges with kind=foreign_key. It is specific and informative, but it does not explicitly differentiate itself from overlapping siblings like infer_table_relation or suggest_refs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context: targeted detection via fromTable/toTable or default full-table pairwise scanning, with topN limiting output. It does not explicitly state when not to use this tool or name alternative tools for other scenarios, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
infer_table_relationARead-only
表内列间派生关系推断(确定性拟合):判断「B 列 ≈ A 列 × k」的系数关系并标出偏离行(取整容限内免误报),或对整表数值列两两配对、自动发现基准锚点列与各列系数。还原「各装备属性 = 基准 × 装备系数」这类生成结构时用;发现系数后可与系数表(如 read_table 装备划分)对照确认同源性。列内时序规律(随行序怎么走)用 infer_column_rule,本工具管列与列的结构。
| Name | Required | Description | Default |
|---|---|---|---|
| table | Yes | 目标表名(来自 list_tables) | |
| columns | No | 参与列。恰 2 列时输出该对的详细关系;多列为子集两两;缺省为全表数值列自动两两 | |
| thresholdPct | No | 相对容差百分比(默认 1,与取整容限取大) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint=true so the agent knows this is a read-only operation. The description adds behavior such as '确定性拟合', tolerance handling (取整容限内免误报), and default behavior (auto-pairwise on all numeric columns) which goes beyond the schema. However, it doesn't elaborate on output format or failure behavior, but annotations cover the safety profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and packed with domain-specific terminology in Chinese, but it front-loads the core function first. It includes example structures and sibling references but could be slightly more concise. It earns a 3 because while informative, it's somewhat dense and may bury the simplest summary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity, the description explains the two modes (pair-specific vs auto-pairwise), default behavior, tolerance semantics, and relationship to sibling tools. There's no output schema, so the description carries the burden of conveying behavior, and it does so adequately. Minor gaps: no explicit statement of what 'deviation rows' output looks like, but context is otherwise complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema documents all three parameters. The description adds context: columns default to all numeric columns, thresholdPct tolerance semantics (relative tolerance, max with rounding tolerance). This is useful but the baseline with full schema coverage is 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: infer linear column relationships (B ≈ A × k) and identify deviating rows, or automatically discover baseline anchor columns and coefficients across numeric columns. It distinguishes this from infer_column_rule explicitly, making its scope unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use it: for reconstructing generation structures like 'equipment attributes = baseline × equipment coefficient', and mentions comparing results with coefficient tables like read_table. It explicitly excludes column-internal temporal patterns, directing to infer_column_rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_tablesARead-only
列出当前工作区的全部数值表清单:表名、所属模块、行数、列名。回答数值问题前先用它定位要查哪张表、有哪些列可用。
| Name | Required | Description | Default |
|---|---|---|---|
| module | No | 按模块过滤(可选):settings/growth/equipment/level/hero/economy/rogue/talent/enemy/gamedata |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes a safe read operation. The description adds valuable behavioral context beyond that: the tool returns metadata for all tables in the current workspace, scoped to a 'numeric tables' inventory, and includes the specific fields returned. This is adequate for a simple listing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler: the first states the resource and output fields, the second provides the usage trigger. Every sentence earns its place and the key scoping information (workspace, table metadata) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a metadata-listing tool with one optional parameter and no output schema, the description is complete: it names what is listed, the fields returned, and when to invoke it. No further return-format or pagination details are necessary for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single 'module' parameter has 100% schema description coverage, listing valid filter values, so the schema does the heavy lifting. The description adds no additional parameter-level detail beyond implying the table list can be scoped by module, which keeps this at the baseline score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('列出' / list) and resource ('当前工作区的全部数值表') and enumerates the output fields (table name, module, row count, column names). It also positions the tool as a discovery step before querying, which separates it from read_table and other table-manipulation siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: use it before answering numeric questions to locate the relevant table and available columns. It does not name alternatives or state when not to use it, but the trigger context is clear enough for an agent to choose this tool over read_table or export_table.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_workspacesARead-only
列出全部可用工作区(位于本机工作区根,GND_WORKSPACES_DIR 环境变量控制,缺省 ~/.gamenumerics/workspaces)。切换工作区用 set_workspace。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description reinforces this by stating it lists (a read-only action). It adds context about the workspace root and the controlling environment variable, which is beyond what annotations provide. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that states the action and resource first, then the location, then a pointer to the sibling tool. No wasted words, efficient and complete.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless list tool with readOnlyHint annotation and no output schema, the description covers the key operational detail (where workspaces are found) and points to the related tool. An agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so per the rubric baseline is 4. The description adds no parameter details because none exist, but it does clarify the source of the list, which indirectly explains why no parameters are needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (list) and the resource (all available workspaces), and specifies the location via environment variable with a default. It distinguishes itself from set_workspace by explicitly naming the sibling for a different purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the alternative tool (set_workspace) for the related but distinct operation of switching, giving clear context on when to use which. However, it doesn't explicitly state when not to use this tool beyond that, but for a list operation the intent is obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
power_curveARead-only
战力曲线(批量):读成长表逐行(可采样)属性面板调 computePowerStats,一次产出「等级→战力」全曲线与形态摘要(首/中位/末档战力、成长倍率、形态判定:匀速/后期加速/台阶断点)。属性列自动按常用名识别(攻击/攻击力→attack、体力/生命→hp、防御→defense、攻速/速度→speed、暴击率→crit),识别不到时用 columns 显式指定。适合全曲线分析与成长×装备联合推导,替代多次单点 compute_power。
| Name | Required | Description | Default |
|---|---|---|---|
| table | Yes | 成长表名(如 hero/成长表) | |
| columns | No | 显式列映射 {attack:'属性.攻击', hp:'属性.体力', defense?, speed?, crit?}——自动识别失败或需换列时用 | |
| levelColumn | No | 等级/序号列名,缺省取首列 | |
| sampleEvery | No | 采样间隔(默认自动:行数>60 时取 ceil(行数/60)),首末行恒在采样内 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, it discloses sampling behavior, automatic column-name recognition with fallback to explicit columns, and the concrete output composition including first/median/last power, growth multiplier, and shape classification. This is substantial behavioral context that the annotation alone does not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every clause earns its place: input, sampling, output, auto-mapping, fallback behavior, and intended use case. It is front-loaded with the core batch-curve purpose and contains no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still enumerates the output fields and shape classifications. The schema covers all parameters and the annotation marks the tool read-only, so nothing essential is missing for correct invocation and interpretation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real parameter semantics: concrete auto-name mappings (攻击力→attack, 体力→hp, etc.), explicit column-mapping examples, and the guarantee that first/last rows are always included when sampling. This raises it above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific action plus resource: it reads a growth table and produces a level-to-power curve with shape summary metrics. It explicitly distinguishes itself from compute_power by presenting itself as the batch/full-curve version and stating it replaces repeated single-point calls.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states when to use the tool: full-curve analysis and growth × equipment joint derivation, and it names compute_power as the single-point alternative it replaces. It does not explicitly state a when-not-to-use condition, but the context is clear enough for an agent to choose correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
profile_tableARead-only
单表画像(确定性统计,无 LLM):行数、列类型分布、枚举度(distinct/行数)、单调性(递增/递减/无)、主键候选(唯一且非空列)、累计列校验(某列≈另一列逐行前缀和,取整容差)、角色候选(growth=有等级列且存在等差模式列;master_data=有索引/序号类列;cost=累计列校验成立)。逆向理解单表结构先用它,结论可经 save_structure 固化。
| Name | Required | Description | Default |
|---|---|---|---|
| table | Yes | 目标表名(来自 list_tables) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, it discloses that the tool is deterministic and LLM-free, and it fully specifies the behavioral criteria for each output category, including formulas like distinct/rows and prefix-sum with rounding tolerance. This gives an agent confidence about repeatability and side-effect-free execution.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The key purpose and nature are front-loaded ('单表画像,确定性统计,无 LLM'), and subsequent clauses compactly enumerate the computed metrics and heuristics. It is dense but every element carries information; only a slightly monolithic long sentence prevents a perfect structure score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with no output schema, the description clearly lists what the agent will get back and even gives role-candidate rules and the follow-up save_structure action. It does not specify exact return shapes or edge-case behavior, but the essential selection and invocation information is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents the only parameter ('目标表名(来自 list_tables)') with 100% coverage, so the baseline is 3. The description does not add parameter-specific guidance beyond saying the profile is for a single table, which is not additional semantic value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: profiling a single table with deterministic statistics, and it enumerates the exact metrics returned (row count, type distribution, cardinality, monotonicity, PK candidates, cumulative-column check, role candidates). It explicitly positions itself for '逆向理解单表结构先用它', which distinguishes it from relation/column-level siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear usage context: use this first when reverse-engineering a single table, and route results to save_structure for persisting conclusions. It does not explicitly state when not to use it or name alternatives for exclusion, so it stops short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_memoryARead-only
读取工作区全部项目记忆(项目画像 PROFILE.md + 事实流 facts.md)。回答与历史结论相关的问题前先查记忆。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description aligns with that by describing a read operation. The description adds useful context by naming the two memory artifacts it reads, but it does not disclose return format, size, or whether it reads from the current workspace specifically. This is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with no filler. The first sentence states exactly what the tool does, and the second sentence gives practical usage guidance. Information is front-loaded and every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple parameterless read-only tool with an annotation covering safety, the description is largely complete: it states the resource, the files involved, and when to use it. It does not explicitly describe the return format, but the nature of reading memory is clear enough for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so parameter semantics are inherently satisfied. The description does not need to explain parameter behavior. Baseline 4 is appropriate for a parameterless tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('read') and resource ('workspace all project memory'), and even names the concrete files involved (PROFILE.md and facts.md). This differentiates it from sibling tools like read_table and list_workspaces, which operate on tabular or workspace-level data rather than project memory.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides an explicit usage context: check memory before answering questions related to historical conclusions. It does not explicitly state when not to use it or name alternatives, but the guidance is clear enough for an agent to know when this tool is relevant.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_tableARead-only
读取工作区内指定数值表的数据行。支持列选择、条件过滤(eq/ne/gt/gte/lt/lte/contains)、分页(offset/limit,默认前 20 行)。先用 list_tables 拿到表名和列名。
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | 返回行数上限(默认 20,全表导出可设大) | |
| table | Yes | 表名,如 hero/成长表(来自 list_tables 的 table 字段) | |
| filter | No | 过滤条件,多条件 AND(可选) | |
| offset | No | 跳过前 N 行(默认 0) | |
| columns | No | 只返回指定列(可选) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the operation as read-only, and the description adds meaningful behavioral details: default limit of 20 rows, supported filters (eq/ne/gt/gte/lt/lte/contains), column selection, and pagination via offset/limit. This goes beyond the annotation without contradicting it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core action, then quickly covers capabilities and the prerequisite workflow. Every sentence earns its place, with no redundant filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only table query tool with five parameters, the description covers the prerequisite, supported filters, pagination, and column selection. It does not describe the return format, but the absence of an output schema is mitigated by the simplicity of the operation and the explicit read-only annotation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the input schema already documents every parameter. The description reinforces the filter operator list and pagination defaults but adds little new meaning beyond the schema; the schema itself already mentions the list_tables source and default limit. This meets the baseline but does not significantly extend it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('读取...数据行') and the resource ('工作区内指定数值表'), and specifies the main capabilities (column selection, filtering, pagination). It also distinguishes itself from list_tables by telling the agent to call list_tables first for table and column names, avoiding confusion with sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides concrete usage context: use this tool to read rows from a table, and explicitly instructs the agent to first call list_tables to obtain table and column names. It does not explicitly state when not to use this tool or mention alternative siblings, but the workflow guidance is sufficient for correct selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recon_diffARead-only
跨表对账(L3):当前工作区两张表按声明键列对齐、值列逐格比对。left 为明细/被检方,right 为基准/主数据方;left 键在 right 缺失 → error(外键断链同构);值比对容差 = 绝对半格(誊抄取整,缺省 0.5)+ 相对 1e-9。ratio 可声明倍率誊抄(如品质系数 1.4)。leftValues/rightValues 省略时为纯键域对账(FK 形态,多对一引用合法)。先 list_tables 拿表名与列名。
| Name | Required | Description | Default |
|---|---|---|---|
| ratio | No | left ≈ right × ratio(缺省 1) | |
| leftKeys | Yes | 左表对齐键列(可多列组合键) | |
| leftTable | Yes | 左表(明细/被检方,来自 list_tables) | |
| rightKeys | Yes | 右表对齐键列 | |
| tolerance | No | 绝对容差(缺省 0.5 誊抄取整半格) | |
| leftValues | No | 左表对账值列(可选;省略 = 纯键域对账) | |
| rightTable | Yes | 右表(基准/主数据方) | |
| rightValues | No | 右表对账值列(与 leftValues 按序配对) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, it discloses precise behavior: missing left keys in right produce an error (FK broken-chain isomorphic), the tolerance formula is absolute half-grid (default 0.5) plus relative 1e-9, ratio handles multiplier transcription, and omitted values select pure key-domain reconciliation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every clause adds operational value: purpose, roles, error semantics, tolerance formula, ratio behavior, optional mode, and prerequisite. No filler or repetition of schema names.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool with no output schema, it covers input roles, error behavior, tolerance, optional reconciliation mode, and table discovery. The only notable omission is the shape or semantics of the result/return value, which would round it out.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though schema coverage is 100%, the description significantly enriches parameters: it defines left as detail/checked and right as baseline/master, explains tolerance's default and formula, gives a ratio example (quality coefficient 1.4), and clarifies that omitting leftValues/rightValues yields FK-style reconciliation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with '跨表对账(L3)' and defines the exact operation: align two tables by declared key columns and compare value columns cell by cell. It specifies left vs right roles and even names the prerequisite list_tables, making the tool's unique role among siblings evident.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for use: current workspace, two-table reconciliation, left/right orientation, and the explicit prerequisite '先 list_tables 拿表名与列名'. It does not name alternatives or when-not-to-use conditions, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_workspaceA
切换当前工作区(后续全部数值工具调用均作用于它)。name 须来自 list_workspaces 的列表。
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | 工作区名(来自 list_workspaces) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of disclosing behavior. It does disclose the key behavioral trait: switching the workspace affects all subsequent numeric tool calls. It also reveals the constraint that name must be a valid workspace from list_workspaces. It doesn't describe invalid-name behavior or persistence, which would be nice but is not essential for such a simple setter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is only two sentences, with the core purpose and side effect stated first and the parameter constraint second. Every sentence earns its place, and there is no vague filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (one parameter, no output schema, no nested objects), the description is complete enough: it states what the tool does, the side effect on subsequent calls, and where the valid value comes from. An agent has sufficient information to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents the name parameter and its source. The description reiterates that name must come from list_workspaces, which matches the schema description and adds no new semantic information beyond the schema. Hence the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'switch current workspace', and clearly states the tool's role as a state setter for all subsequent numeric tool calls. The context distinguishes it from sibling tools like list_workspaces and read_table by emphasizing that it changes the active workspace rather than querying or transforming data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: call this to change the workspace that subsequent numeric tools will act on. It also states the prerequisite that name must come from list_workspaces, which is actionable guidance. It does not explicitly mention exclusions or alternatives, but there is no real alternative for this action among the siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
simulate_gachaARead-only
概率模拟(确定性引擎):mode=tiers 模拟无保底分层概率池(rogue 技能三选一、宝箱品质掉落)N 次抽取的各层次数分布与分位;mode=pity 模拟带保底的抽卡(基础概率+软/硬保底)抽数分布与期望。
| Name | Required | Description | Default |
|---|---|---|---|
| mode | Yes | tiers=分层池抽取;pity=保底抽卡 | |
| draws | No | mode=tiers:单轮抽取次数(如一局 15 次三选一 = 15 抽) | |
| tiers | No | mode=tiers 必填:各层名称与单次命中率(概率和须为 1) | |
| baseRate | No | mode=pity:基础概率(如 0.006) | |
| hardPity | No | mode=pity:硬保底抽数 | |
| iterations | No | 模拟轮数(默认 1000) | |
| softPityStart | No | mode=pity:软保底起始抽数 | |
| softPityIncrement | No | mode=pity:软保底每抽概率增量 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, so the description is not burdened with stating read-only behavior. The description adds the 'deterministic engine' trait, but does not elaborate on output format, performance, or error handling. Given the annotation coverage, a 3 is appropriate – the description adds some context but misses richer behavioral details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, dense with information, and front-loads the mode concept before detailing each variant. No wasted words; every clause adds meaning. Structure is exemplary for a multi-mode tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Without an output schema, the description must explain return values. It mentions 'distribution and quantiles' and 'distribution and expectation' but does not specify the exact structure (e.g., array of counts, percentiles). For a tool with 8 parameters and two complex modes, this is a noticeable gap, though the description covers the essential functional scope.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description adds value by explaining mode-dependent parameter requirements (mode=tiers needs tiers; mode=pity needs baseRate, hardPity, etc.), clarifying conditional semantics not obvious from the schema alone. This elevates it to 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb (simulate) and resource (gacha/probability draws), and explicitly distinguishes two modes (tiers for layered pools without pity, pity for gacha with guarantee). Concrete examples (rogue skill three-choice, chest quality drops) anchor its purpose and differentiate it from sibling tools like battle_simulate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly specifies when to use each mode (mode=tiers for layered pools, mode=pity for guaranteed draws), which effectively routes the agent to the correct parameter set. However, it does not mention when not to use this tool or alternative tools for similar probabilistic tasks, so it falls short of explicit exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
suggest_refsARead-only
引用完整性候选发现:扫描当前工作区全部表,找出「A 表某列的值疑似引用 B 表某键列」的候选映射(值域包含推断,词面相似度仅排序)。候选不是结论——须向用户确认映射后,用 recon_diff 对确认的映射执行零容忍断链扫描。典型流程:suggest_refs 列候选 → 用户确认 → recon_diff 逐对检查(悬挂行即配置断链:关卡表配置的装备等级在装备等级表无对应条目)。
| Name | Required | Description | Default |
|---|---|---|---|
| topN | No | 候选数上限(可选,1-100,缺省 20) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint already marks this as read-only, and the description adds valuable behavioral context: outputs are candidates, not conclusions; inference is based on domain containment; and lexical similarity is only used for ranking. It also warns that confirmed mappings still require recon_diff, which clarifies the tool's limitations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description packs purpose, workflow, caveats, and an illustrative example into one coherent paragraph. The key constraint that candidates are not conclusions is front-and-center, and every sentence contributes useful information, though the paragraph is somewhat dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description adequately explains what the tool scans, how candidate discovery works, and what the follow-up step should be. It stops short of describing the exact output shape, but the workflow and boundary against recon_diff are sufficient for an agent to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter topN is fully described in the input schema with a range and default, so schema coverage is 100%. The description adds no additional parameter-level semantics, which is acceptable given the schema already carries that burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening phrase '引用完整性候选发现' names an explicit task, and the description specifies scanning all tables in the current workspace to find candidate mappings where a column in table A may reference a key column in table B. It also differentiates itself from recon_diff by framing outputs as unverified candidates rather than confirmed broken-link results.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear workflow: use suggest_refs to generate candidates, get user confirmation, then use recon_diff for zero-tolerance broken-link checks. This gives concrete context for when to use the tool and how it relates to recon_diff, though it does not explicitly contrast it with other sibling discovery tools such as infer_foreign_keys.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
20 tool updates
v0.1.0- First observed
audit_column - First observed
battle_simulate - First observed
compute_power - First observed
eval_formula - First observed
export_table - First observed
grade_workspace - First observed
import_xlsx - First observed
infer_column_rule - First observed
infer_foreign_keys - First observed
infer_table_relation - First observed
list_tables - First observed
list_workspaces - First observed
power_curve - First observed
profile_table - First observed
read_memory - First observed
read_table - First observed
recon_diff - First observed
set_workspace - First observed
simulate_gacha - First observed
suggest_refs
TDQS
Scored across 20 tools
Most tools have distinct targets, and the descriptions are detailed enough to separate table I/O, column inference, simulation, and workspace management. However, infer_foreign_keys and suggest_refs are near-duplicate FK-candidate scanners, and read_table/export_table plus several column-analysis tools have partial boundary overlap.
The vast majority of names follow a clear snake_case verb_noun pattern (list_tables, read_table, compute_power, set_workspace). Minor deviations like battle_simulate, power_curve, and recon_diff break the pattern but remain readable and predictable.
At 20 tools the server is on the high side, but the scope is genuinely broad: import/workspace management, table I/O, structural inference, reconciliation, simulations, and calculation utilities. It is slightly heavy due to a couple of redundant FK tools, but not unreasonable for the domain.
The core read/import/analyze/simulate path is well covered, but key workflow tools referenced in the descriptions are missing: write_table (apply_curve) and save_structure cannot be called, so applying inferred rules or persisting structural results creates dead ends. There is also no update/delete or memory-write capability beyond workspace import.
Maintenance
Related MCP Connectors
Deterministic reasoning stack for AI agents: simulate, decide & compute, plus cross-domain tools.
Structured financial modeling for AI agents: build, version, audit models, export to Excel.
Precision math engine for AI agents. 203 exact methods. Zero hallucination.
500+ deterministic tools for AI agents: math, conversion, validation, hashing, encoding, date/time.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to create, read, and manipulate Excel files without requiring Microsoft Excel installation. Supports comprehensive spreadsheet operations including formulas, formatting, charts, pivot tables, and data validation.MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to create, read, and modify Excel workbooks without requiring Microsoft Excel, supporting operations like formulas, charts, pivot tables, formatting, and data validation.MIT
- AlicenseAqualityDmaintenanceEnables AI assistants to perform comprehensive Microsoft Excel operations including data analysis, cell editing, advanced formatting, and VBA execution on Windows systems. It provides a structured workflow for managing workbooks and worksheets through a dedicated Model Context Protocol interface.5101 npm4MIT
- AlicenseNot gradedqualityFmaintenanceEnables AI agents to analyze Excel files through atomic operations like filtering, aggregation, and grouping without loading raw data into context.46AGPL 3.0