Skip to main content
Glama
USHIKUNDESUYO

shiwake-mcp

shiwake-mcp

仕訳データの一次スクリーニングを行う MCP サーバー。依存ライブラリはゼロ。

総勘定元帳から仕訳を受け取り、15個のルールを当てて、先に人間が目を通すべき順番に並べ替えて返します。監査の現場でいう仕訳テスト(Journal Entry Testing)を、AIエージェントから呼べる形にしたものです。

English

依存ライブラリを持たない理由

npm install で入るものが1つもありません。package.json の dependencies は空です。

会計データを扱う道具に外部依存を足すと、導入のたびに「このパッケージは何をしているのか」を説明する必要が出ます。監査法人や会計事務所のネットワークで動かすとき、その説明コストは実装の手間より高くつきます。依存がゼロなら、読むべきコードはこのリポジトリの中だけで閉じます。

MCP の stdio トランスポートは、行区切りの JSON-RPC 2.0 です。SDK を使わなくても 200 行ほどで書けます。

Related MCP server: sigmodx-mcp

動かす

Node.js 20 以上が必要です。

MCP サーバーとして繋ぐ

npm に公開しているので、npx で起動できます。事前のインストールは要りません。ダウンロードされるのはこのパッケージ1つだけです。依存がないので、ほかには何も入りません。

Claude Code なら1行です。

claude mcp add shiwake -- npx -y shiwake-mcp

名前の前に --scope project を付けると、プロジェクト直下の .mcp.json に書き込まれ、チームで共有できます。

Claude Desktop は設定ファイル(claude_desktop_config.json)に追記します。

{
  "mcpServers": {
    "shiwake": {
      "command": "npx",
      "args": ["-y", "shiwake-mcp"]
    }
  }
}

バージョンを固定したい場合は shiwake-mcp@0.1.0 のように指定します。コードを読んでから動かしたい場合は、リポジトリを clone して "command": "node"、"args": ["/path/to/shiwake-mcp/src/server.js"] と直接指定してください。

大きな元帳はファイルで渡す

数百件を超える元帳は、仕訳を会話に並べずに、ファイルのパスを渡します。仕訳を会話で渡すと、AI がそれを全部書き出すことになり、数万件では収まりません。

サーバーが読んでよいフォルダを、起動時に --data-dir で指定します(何度でも指定できます。環境変数 SHIWAKE_DATA_DIR でも指定できます)。

claude mcp add shiwake -- npx -y shiwake-mcp --data-dir /path/to/ledgers

Claude Desktop なら "args": ["-y", "shiwake-mcp", "--data-dir", "C:\\audit\\ledgers"] です。

指定したフォルダの外にあるファイルは、AI からパスを渡されても開きません。シンボリックリンクやジャンクションで外を指している場合も、たどった先で判定して読みません。

あとは「2025年度_仕訳帳.csv を screen_journals で見て」のように頼めば、ツールに file が渡ります。相対パスは、最初に指定したフォルダを起点にします。読めるのは .json と .csv(UTF-8・Shift_JIS)で、上限は 256MB です。手元の計測では、38万件(CSV 76MB)で約9秒、応答は約8万字でした。

まず手元で試す

リポジトリを clone すると、同梱のサンプルデータ(合成データ383件、既知の異常を混ぜてあります)で挙動を確かめられます。npm install は要りません。

git clone https://github.com/USHIKUNDESUYO/shiwake-mcp.git
cd shiwake-mcp
npm run demo
検査対象   383 件
検出       54 件 / 対象仕訳 25 件
重要度     high 13 / medium 16 / low 25
ベンフォード  MAD 0.012241 → 許容の限界(n=383)

ルール別:
     営業時間外の入力                11 件
     キリのよい金額                   9 件
  !! 承認限度額の直下                 7 件
  !  重複仕訳                         5 件
  !! 起票者と承認者が同一             4 件
  !  期末直前の大口計上               4 件
  !  計上日と入力日の乖離             3 件
  !  稀な勘定科目の組み合わせ         3 件
     休日の計上                       3 件
  !! 貸借不一致                       2 件
     摘要が空                         2 件
  !  期末後の入力                     1 件

確認の優先順位(上位 20 件):
  [ 23] JV-0382    2025-09-06       3000000  self_approval, rare_account_pair, weekend_or_holiday, after_hours, round_amount, missing_description
  [ 19] JV-0365    2026-03-30       8900000  self_approval, period_end_large, after_hours, round_amount
        開発委託費
  [ 17] JV-0362    2026-01-22       1200000  unbalanced, rare_account_pair, round_amount
        業務委託費計上
  (以下省略)

383件が25件に絞られます。スコアは各ルールの重要度の合計で、複数のルールに同時に当たった仕訳ほど上に来ます。

ツール

ツール

何を返すか

screen_journals

全ルールを当て、リスクスコア順に並べ替えた一覧

check_balance

貸借が一致しない仕訳と、その差額

benford_analysis

金額の先頭桁の分布、MAD、χ²

detect_duplicates

完全に一致する仕訳のグループ

list_rules

実装されているルールの一覧と趣旨

どのツールも、MCP のツール注釈で「読み取り専用(readOnlyHint)・外部に触れない(openWorldHint: false)」と宣言しています。何かを書き換えたり、外のサービスに送ったりはしません。

仕訳は journals に並べて渡すか、file でファイルのパスを渡します(前節)。

screen_journals が返す個々の検出(findings)は、上位 top 件(既定 50)の仕訳に関わるものだけです。件数の集計は常に全件です。全件が必要なら allFindings: true を渡します。check_balance と detect_duplicates の一覧も、既定で 200 件までです。

入力の仕訳は、簡易形と明細形のどちらでも受けます。

{
  "id": "JV-0001",
  "date": "2026-03-31",
  "entered_at": "2026-04-02T23:41:00+09:00",
  "debit_account": "売掛金",
  "credit_account": "売上高",
  "amount": 12000000,
  "description": "3月度売上計上",
  "created_by": "acc01",
  "approved_by": "mgr01"
}

消費税や複合仕訳のように行数が増えるものは、明細形で渡します。

{
  "id": "JV-0002",
  "date": "2026-03-31",
  "lines": [
    { "account": "外注費", "debit": 1000000 },
    { "account": "仮払消費税", "debit": 100000 },
    { "account": "買掛金", "credit": 1100000 }
  ]
}

1件でも読めない仕訳があると、既定では全体を止めて、何件目のどこが読めないかを返します。読めない行を除外して続けたいときは skipInvalid: true を渡します。除外した行は、何件目か・伝票番号・理由を添えて invalidRows に返ります。

entered_at は日付だけでも受けます。その場合は「計上日と入力日の乖離」には使い、「営業時間外の入力」には使いません。

CSV の読み方

会計ソフトから書き出した CSV を、そのまま渡せます。列名は次の候補から自動で当てます。全角と半角、空白、「(税込)」のような括弧の注記の違いは無視します。

項目

列名の候補

伝票番号

伝票番号・伝票No・仕訳番号・取引番号・No など

日付

日付・取引日・計上日・伝票日付・仕訳日 など

入力日時

入力日時・登録日時・作成日時・入力日 など

借方科目・貸方科目

借方勘定科目・借方科目、貸方勘定科目・貸方科目

金額

借方金額と貸方金額、または 金額

摘要

摘要・内容

起票者・承認者

入力者・起票者・作成者、承認者

当たらない列は columns で指定します(例: { "date": "伝票日付", "amount": "金額(税込)" })。どの列を使ったかは応答の source.columnsUsed に返るので、確かめてから結果を読んでください。

1行に借方と貸方を持つ形を前提に、伝票番号と日付が同じ行を1つの仕訳にまとめます。複合仕訳の相手に使われる「諸口」は、同じ伝票の中で借方と貸方が同額なら取り除きます。合わないときは、貸借のずれを隠さないよう残します。

日付は 2026/3/31・20260331・2026年3月31日・R8.3.31・令和8年3月31日 などを読みます。金額の桁区切りと円記号は落とし、△ と括弧は負の数として扱います(負の金額は、既定では止まります)。見出しの前に表題の行があっても、見出しの行を探して読みます。エラーと除外の位置は、CSV の何行目かで返します。

弥生会計の「弥生インポート形式」(見出しの行が無い25項目または27項目の CSV)は、1項目めの識別フラグで見分けて、列の位置で読みます。2000・2111 は1行で1仕訳、2110 から 2101 までの行を1つの仕訳にまとめます。取引日付は 20260331・2026/3/31・R08/03/31 のどれでも読みます。列の並びは、弥生会計サポート情報「仕訳データの項目と記述形式」の表に合わせています。この形式には入力日時の列が無いので、入力日時を使う3つのルール(計上日と入力日の乖離・期末後の入力・営業時間外の入力)は動きません。

ルール

ID

内容

重要度

何を示すか

unbalanced

貸借不一致

high

手入力、取込不良、改変のいずれか

self_approval

起票者と承認者が同一

high

職務分掌が効いていない

threshold_avoidance

承認限度額の直下

high

分割計上による承認回避

duplicate

重複仕訳

medium

二重計上、または正当な定期計上

reversal

取消・訂正仕訳

medium

誤りの取消・訂正。期末をまたぐ組は期間帰属の確認へ

backdated

計上日と入力日の乖離

medium

期間帰属の誤り、遡及計上

post_period_entry

期末後の入力

medium

決算整理と締めたあとの修正。統制の無効化が現れやすい

period_end_large

期末直前の大口計上

medium

利益調整が現れるならこの窓

rare_account_pair

稀な勘定科目の組み合わせ

medium

通常の取引フローから外れた処理

weekend_or_holiday

休日の計上

low

業務サイクルの外での処理

after_hours

営業時間外の入力

low

単独では弱いが、重なると効く

round_amount

キリのよい金額

low

見積、概算、付け替え

missing_description

摘要が空

low

監査証跡としての品質

description_keyword

摘要のキーワード

low

事後の手直し、内容の定まっていない計上

voucher_gap

伝票番号の欠番

low

削除・取消された伝票、出力の漏れ

前提条件はオプションで渡します。

{
  "fiscalYearEnd": "03-31",
  "businessHours": [9, 18],
  "holidays": ["2026-01-01", "2026-01-12"],
  "approvalThresholds": [1000000, 5000000],
  "backdatedDaysThreshold": 30
}

approvalThresholds と fiscalYearEnd を渡さなければ、対応するルールは動きません。関係のないルールが空振りして偽陽性を増やすより、明示的に止まるほうがよいという判断です。

「休日の計上」は、月末日付の仕訳を既定で対象から外します。月次・期末の整理仕訳は、土日でも月末の日付で計上されることが多いためです。3月31日が日曜だった2024年3月期のような年は、外さないと期末の整理仕訳がすべて当たります。月末も含めて見る場合は exemptMonthEnd: false を渡します。

「休日の計上」は、土日に加えて日本の祝日(振替休日・国民の休日を含む、2000〜2099年)を自動で見ます。祝日は内閣府の一覧を取り込まず、祝日法の規定から計算しています。2000〜2027年の全日で、内閣府の一覧と一致することを確かめました。年末年始のような会社独自の休日は holidays で足します。日本以外の元帳に使う場合は japaneseHolidays: false を渡します。

「重複仕訳」と「稀な勘定科目の組み合わせ」は、借方・貸方それぞれの科目を並べ替えてから比べます。明細の行の順番は結果に影響しません。

「期末後の入力」は、期末日より後に入力された、期末日以前の日付の仕訳を拾います。決算整理と、締めたあとの修正がここに集まります。「計上日と入力日の乖離」は既定で30日を超えた遅れしか拾わないので、3月31日付を4月10日に入力したような短い遅れは、こちらで拾います。fiscalYearEnd と entered_at がそろっているときだけ動きます。

「取消・訂正仕訳」は、同じ金額で借方と貸方を入れ替えた仕訳が 30 日以内(reversalWindowDays)にある組を、両方に相手の伝票番号を添えて返します。1件は1組にしか入れません。期末をまたぐ組は理由にそう書き添えますが、重要度は上げません。期首の洗替仕訳も同じ形になるためです。

「摘要のキーワード」の既定の語は「修正」「訂正」「取消」「調整」「仮計上」「不明」です。descriptionKeywords で差し替えられ、空の配列を渡すと止まります。「仮」1文字は仮払金などを拾いすぎるので、既定には入れていません。

「伝票番号の欠番」は、伝票番号を頭の文字と末尾の数字に分け(JV-0382 なら JV- と 382)、頭の文字ごとに連番の飛びを探します。欠けているのは仕訳そのものなので、前後の仕訳のスコアには入れず、欠番の一覧として返します。範囲の半分以上が欠けている番号は、連番で振られていないとみなして見ません。伝票番号の無い仕訳も見ません。

ベンフォード分析について

MAD の判定境界は Nigrini, M. J. Benford's Law (Wiley, 2012) Table 5.1 の値を使っています。実務で広く引かれている値ですが、法令や監査基準が定めたものではありません。

サンプルが300件を下回る場合、結果に注記が付きます。この判定境界は大標本を前提にしているためです。

そして、ベンフォードは母集団の性質を見る道具であって、個別の仕訳を判定するものではありません。分布が崩れていても、事業の性質(単価が固定の商売、規制価格、少額取引の多い業態)で説明がつくことが普通にあります。

この道具の限界

検出は不正の証拠ではありません。 どのルールも、正当な処理を大量に拾います。重複仕訳の多くは毎月同額の定期計上ですし、期末の大口は期末に売上が立つ商売なら当たり前に出ます。

この道具がやるのは、母集団のどこから見るかを決めることだけです。検出された仕訳をどう評価するかは、依然として人の仕事として残ります。

以下は、この道具ではできません。

  • 監査手続そのものの代替(十分かつ適切な監査証拠は、これでは得られません)

  • 不正の有無の結論づけ

  • 勘定科目の内容的な妥当性の判断

  • 税務上の取扱いの判定

監査意見の形成や、税務申告の根拠として使えるものではありません。

実データの取り扱い

examples/ に入っているのは合成データです。固定シードで生成しているので、node examples/generate.js を何度実行しても同じファイルになります。

.gitignore で *.csv *.xlsx journals.json /data/ を除外しています。実際の仕訳データをコミットしないための保険です。

サーバー自体はネットワークに出ません。読むのは stdin と、起動時に --data-dir で許可したフォルダの中の .json と .csv だけです。書くのは stdout だけで、ファイルには書きません。

テスト

npm test

107件のテストが走ります。MCP サーバーのテストは、子プロセスとして起こして実際に JSON-RPC を投げる経路で書いています。

CI は Node 20 / 22 / 24 で走ります。テストのほかに、依存が増えていないこと、package-lock.json が生まれていないこと、固定シードのサンプルデータが再生成しても一致することを検査しています。依存ゼロはこのリポジトリの前提なので、人の注意ではなく CI で守っています。

解説記事

このサーバーを書いた経緯と設計の判断は、記事にしています。

ライセンス

MIT

作者

星野宇潮(公認会計士・税理士)

Available Tools

5 tools
benford_analysisA

金額の先頭桁の分布をベンフォードの法則と比較し、MAD と χ² を返す。母集団の性質を見るための道具で、個別仕訳の判定には使えない。

ParametersJSON Schema
NameRequiredDescriptionDefault
digitsNo1 = 先頭1桁、2 = 先頭2桁。既定は 1。
amountsNojournals の代わりに金額だけを渡す場合。
journalsNo仕訳の配列。簡易形 { date, debit_account, credit_account, amount } か、明細形 { date, lines: [{ account, debit, credit }] } のどちらでも受ける。

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must disclose behavioral traits. It states that the tool returns MAD and χ² and that it is not for individual judgment, but it does not explicitly mention that it is read-only or describe how it handles edge cases (e.g., zero or negative amounts). Since the name and nature imply a non-mutating analysis, the description is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The core purpose and a critical usage constraint are front-loaded, and every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (3 params, no output schema, no annotations), the description covers the essential usage distinction and return metrics. It does not detail the exact output shape, but for a simple statistical result this is a minor gap. Overall, the description is sufficiently complete for an agent to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% – all three parameters (digits, amounts, journals) have clear descriptions in the schema. The tool description adds little beyond the schema, but it does clarify the overall purpose, which indirectly aids parameter understanding. Baseline 3 is appropriate when schema fully documents parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('compare'), a specific resource (leading-digit distribution of amounts), and the output (MAD and χ²). It also explicitly distinguishes the tool from individual-entry judgment, which differentiates it from siblings like screen_journals and detect_duplicates.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly states when to use it (for population properties) and when not to (for individual journal judgment). It does not name an alternative tool explicitly, but the exclusion is unambiguous and sufficient for an agent to avoid misuse.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_balanceA

貸借が一致しない仕訳だけを返す。取込不良と手入力の混入を最初に落とすために使う。

ParametersJSON Schema
NameRequiredDescriptionDefault
journalsYes仕訳の配列。簡易形 { date, debit_account, credit_account, amount } か、明細形 { date, lines: [{ account, debit, credit }] } のどちらでも受ける。

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It clearly states this is a read-only filter that returns only unbalanced journals and frames it as a data-quality pre-filter, which is the main behavior an agent needs to know.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, purposeful sentences; the core behavior is front-loaded and the second sentence adds a concrete use case. There is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple, accepts one well-documented parameter, and the description states both the input purpose and the output predicate. It is fully usable without needing extra edge-case or output formatting details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds a tool-level purpose but does not expand on the `journals` parameter semantics beyond what the schema already documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('返す') and a precise resource ('貸借が一致しない仕訳だけ' return only unmatched journal entries), clearly distinguishing this from sibling tools like screen_journals or detect_duplicates. It defines the tool's core behavior in one concrete predicate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence explicitly states when to use it: as the first-pass filter to drop import failures and manual-entry contamination. It gives clear usage context but does not explicitly name alternatives or when not to use it, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

detect_duplicatesB

計上日・金額・借方科目・貸方科目が完全に一致する仕訳をグループにして返す。

ParametersJSON Schema
NameRequiredDescriptionDefault
journalsYes仕訳の配列。簡易形 { date, debit_account, credit_account, amount } か、明細形 { date, lines: [{ account, debit, credit }] } のどちらでも受ける。

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It discloses the grouping behavior (returns groups of matching journals) and the exact matching criteria. However, it doesn't disclose what the output format looks like, whether it handles both simple and detailed journal forms consistently, or any edge cases like partial matches or how groups are ordered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, focused sentence in Japanese that states the core behavior and matching criteria. It's efficient and front-loaded with the key information. It could arguably add a note about output format, but as-is it's concise and clear.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one parameter and no output schema, the description covers the input format and matching logic. However, it doesn't describe the return value structure (what a 'group' looks like), which is important for an agent to use the result. Given the tool's simplicity, this is a moderate gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the journals parameter thoroughly, including both simple and detailed forms. The description adds the matching criteria context but doesn't add new parameter-level meaning beyond what the schema provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('detect duplicates'), a resource ('journals'), and the exact matching criteria (date, amount, debit account, credit account all matching). It clearly distinguishes this from sibling tools like benford_analysis or check_balance, though it doesn't explicitly name them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool: when you need to find journal entries that are exact duplicates across four fields. However, it doesn't explicitly state when not to use it or mention alternatives like screen_journals for reviewing entries. The usage context is clear but not fully elaborated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_rulesA

実装されているスクリーニングルールの一覧と、それぞれが何を示すかの説明を返す。

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It conveys the core behavior — a read operation that returns a list with explanatory text — but offers no detail about response shape, ordering, or whether the rule set is complete. For a simple listing tool this is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler, and the action plus object are front-loaded. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter listing tool with no output schema and no annotations, the description covers the essential facts: what is returned and what each list entry means. It could note how this fits the screening workflow (e.g., reading rules before running screen_journals), but nothing critical for invoking the tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema declares zero parameters, so the baseline is 4; there are no arguments the description needs to clarify. The sentence appropriately refers only to the returned list and its explanations, which is all the schema requires.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('returns') and resource ('list of implemented screening rules'), and adds what each item contains ('description of what each indicates'). It is implicitly distinct from siblings like screen_journals, benford_analysis, and detect_duplicates, which appear to run analyses rather than enumerate rules, though this distinction is not stated explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool versus the sibling screening tools, and no exclusions or conditions are mentioned. An agent must infer that list_rules is a discovery or prerequisite step before running the screening tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

screen_journalsA

仕訳データに全ルールを適用し、リスクスコアの高い順に並べ替えて返す。先に人間が見るべき束を絞り込むための一次スクリーニング。検出は不正の証拠ではない。

ParametersJSON Schema
NameRequiredDescriptionDefault
topNo上位何件を返すか。既定は 50。
optionsNoスクリーニングの前提条件。監査対象の実態に合わせて渡す。
journalsYes仕訳の配列。簡易形 { date, debit_account, credit_account, amount } か、明細形 { date, lines: [{ account, debit, credit }] } のどちらでも受ける。

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavior itself. It discloses that detection is not evidence of fraud and that it returns a sorted list. However, it does not state that the operation is read-only/non-destructive, does not describe pagination or truncation via the 'top' parameter, and gives no detail on the output structure (e.g., risk scores, matched rules). The caveat is useful, but more behavioral context is missing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the primary action and outcome, then a critical caveat. There is no redundancy or filler; every sentence earns its place. The caveat is placed at the end, which does not obscure the main purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a complex nested options object and no output schema, the description does not describe the return format (e.g., whether it includes risk scores, rule triggers, or just the sorted journals). It also does not mention that it accepts both simple and detailed journal formats (though the schema covers this). The description is adequate for a basic understanding but lacks details needed for an agent to interpret results correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description mentions applying rules to journal data but does not elaborate on the 'top' or 'options' parameters beyond what the schema already documents. It adds no additional semantic value to parameter usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it applies all rules to journal data, sorts by risk score descending, and returns results. It positions itself as a primary screening tool distinct from sibling tools like check_balance or detect_duplicates, which are more targeted. The phrase '仕訳データに全ルールを適用し' makes the action and resource explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives context as a '一次スクリーニング' (primary screening) for narrowing what humans should review first, implying it is the initial step. However, it does not explicitly mention when to avoid using it or direct the agent to specific siblings for narrower analyses (e.g., detect_duplicates for duplicate checks). Guidance is implied but not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.1.0
    • First observedbenford_analysis
    • First observedcheck_balance
    • First observeddetect_duplicates
    • First observedlist_rules
    • First observedscreen_journals

TDQS

A3.8/5.0

Scored across 5 tools

Disambiguation4/5

Each tool targets a distinct screening concern—aggregate risk, balance errors, Benford distribution, duplicates, and rule metadata. screen_journals conceptually overlaps by applying all rules, but its description clearly frames it as a first-pass triage rather than a replacement for the focused checks.

Naming Consistency4/5

Four tools follow a clear verb_noun pattern (screen_journals, check_balance, detect_duplicates, list_rules), and all names are snake_case. benford_analysis breaks the pattern by using a noun phrase instead of a verb-object form, though it remains readable.

Tool Count5/5

Five tools is a well-scoped size for a specialized journal-entry screening server. Each tool covers a meaningful stage of the workflow without redundancy or bloat.

Completeness4/5

The set covers first-pass screening, unbalanced-entry cleanup, population-level Benford testing, duplicate detection, and rule documentation, which forms a coherent lifecycle. A per-entry rule explanation or a way to inspect raw risk factors would be a minor gap, but agents can work around it via screen_journals and list_rules.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    Audit infrastructure for AI agents to log consequential decisions (invoice, GL, anomaly) and verify attestations via MCP tools.
    6
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables AI agents to securely call MCP tools with risk scoring, checkpoints, rollback, and approval workflows.
    13 npm
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Enables AI assistants to perform financial analysis, budget forecasting, compliance checks, expense categorization, and risk assessment, returning structured JSON with audit-ready governance receipts.
    5
    44 npm
    1
    Business Source 1.1