gristmill-mcp
gristmill-mcp
AIが生成したコードを検査し、構造的・セキュリティ上の違反を決定的に列挙するMCPサーバーです。AIコーディングエージェントがコードを確定させる前に自身の出力を修正できるようにします。
グリストとは製粉所に運ばれる穀物のことです。AIの出力はグリスト — 確かに価値のある原料ですが、未加工です。製粉所がそれに構造を与えます。
AIがグリストを書く。Gristmillがそれを出荷可能なコードにする。
なぜMCPサーバーなのか、スキルではないのか
スキルはモデルのコンテキストに読み込まれるテキストであり、モデルが知っていることを変えます。MCPサーバーはモデルが実行するプログラムであり、モデルができることを変えます。
スタイルのガイダンス(「クラスを緩やかな関数より優先する」)はスキルに属します。検証(「このファイルには12行目、40行目、66行目…に7つのトップレベル関数があります」)はファイルに対してコードを実行する必要があります。モデルが自身の出力を読んで「これには関数が多すぎるように見える」と推論するのは、観察に擬態した推測にすぎません — このファイルにおける「多すぎる」の意味の真実はなく、信頼できるカウント方法もありません。GristmillはASTを解析して実際に数えます。指示と実行のこの区別こそが、これがアドバイスの段落ではなくサーバーとして存在する理由です。
サーバーは決してLLMを呼び出さず、同じ入力に対して実行間で変化することはなく、信頼スコアを出力することもありません。同じ入力 → 毎回バイト単位で同一の出力。その決定性こそが製品のすべてです。AI層はこのサーバーの上にあり、その結果を消費し、何をすべきかを決定します。サーバーの役割は行番号付きの事実を報告することで終わります。
Related MCP server: code-verify-mcp
インストール
git clone <this repo> gristmill-mcp
cd gristmill-mcp
python3 -m venv .venv
.venv/bin/pip install -e .Claude Code
venvのコンソールスクリプトを指すようにCLIで登録します:
claude mcp add gristmill -- /absolute/path/to/gristmill-mcp/.venv/bin/gristmill-mcpまたは、MCP設定に直接追加します(プロジェクト内の.mcp.json、またはグローバルのClaude Code設定):
{
"mcpServers": {
"gristmill": {
"command": "/absolute/path/to/gristmill-mcp/.venv/bin/gristmill-mcp"
}
}
}その他のMCPクライアント
stdioベースのMCPクライアントは同じバイナリを起動できます。gristmill-mcp(またはvenv内でpython3 -m gristmill.server)は標準のMCP stdioトランスポートをクライアント固有の設定なしで使用します。
コマンドライン(MCPクライアントなし)
ローカルテスト用、または以下の実際の例を再現するために、薄いCLIが同じエンジンをラップしています:
.venv/bin/gristmill-verify path/to/file_or_dir [--checks secrets structure comment_slop] [--severity-floor warning] [--json]実際の例
demo/billing.py、Stripe課金ヘルパーの未編集の初稿:
import stripe
# I've added this as you requested — sets up the Stripe client
STRIPE_SECRET_KEY = None # was a literal sk_live_... key — see note below
stripe.api_key = STRIPE_SECRET_KEY
def customer_create(config):
return stripe.Customer.create(**config)
def customer_delete(config):
return stripe.Customer.delete(config["id"])
def customer_find(config):
return stripe.Customer.retrieve(config["id"])
def customer_update(config):
return stripe.Customer.modify(config["id"], **config).venv/bin/gristmill-verify demo/billing.py上記のNoneの代わりに実際のStripeライブキー形式のリテラルを入れた場合の出力:
gristmill: 1 files scanned, 0 skipped (2 error, 4 warning, 0 info) in 1ms
[WARNING] STR002 billing.py:1 4 top-level functions share the prefix `customer_` — consider a `Customer` class or module
[WARNING] STR003 billing.py:1 4 top-level functions take a first parameter named `config` — consider making it instance state
[WARNING] CMT001 billing.py:3 Comment addresses the reader conversationally ('as you requested')
[ERROR ] SEC006 billing.py:4:22 Stripe live key assigned to `STRIPE_SECRET_KEY`
[ERROR ] SEC010 billing.py:4:22 String literal assigned to `STRIPE_SECRET_KEY`, which looks credential-shaped
[WARNING] SEC011 billing.py:4:22 High-entropy string literal (5.1 bits/char) assigned to `STRIPE_SECRET_KEY`(ファイルパスは最も近い.gristmill.tomlからの相対パスで表示されます。demo/は独自の設定を持っているため、この例の出力はトップレベルのプロジェクト設定に依存しません。)
注釈: GitHubのプッシュ保護は、実際の形式のシークレットを含むファイル(コメント内、Markdownコードブロック内、このREADMEを含む)のプッシュをブロックします。現在
demo/billing.pyはキーをNoneに置き換えて初期プッシュを解除しています。これは(許可リストに登録されたシークレットスキャン例外を介して)復元し、デモを再び動作させるためのTODOです。
--jsonフラグ(または両方を返すverify MCPツール)は、完全な構造化形式(ファイル、行、列、静的な提案文字列、および編集されたevidenceフィールド(sk_l…(49文字)、キー自体は決して表示されません))を提供します。
ツール
verify
ソースファイルを検査し、シークレット、構造上の問題、質の低いコメントを検出します。ファイルパスと行番号付きの決定的な結果を返します。コードを生成または編集した後、完成品として提示する前に呼び出してください。
入力:paths(ファイルまたはディレクトリ、必須)、checks(オプション、secrets/structure/comment_slopのサブセット、デフォルトはすべて)、severity_floor(オプション、デフォルトはinfo)。
出力:コンパクトな人間可読サマリーと、それに続く完全な構造化JSON(ファイル、行、列、メッセージ、編集された証拠、ルールごとの静的な提案文字列)。結果は常にpath→line→rule_idでソートされます。この安定性により、実行間でバイト単位の同一性が保証され、モデルが問題に直接移動できるようになります。
explain_rule
rule_id(例:SEC001)を受け取り、その根拠、検出対象、検出できないもの、抑制方法を返します。これはdocs/RULES.mdと同じ内容をオンデマンドで提供し、verifyの出力を簡潔に保つことができます。
ルール
ルール | チェック | タイトル | デフォルトの重要度 |
| secrets | AWSアクセスキーID | エラー |
| secrets | AWSシークレットアクセスキー | エラー |
| secrets | GitHubトークン | エラー |
| secrets | Google APIキー | エラー |
| secrets | Slackトークン | エラー |
| secrets | Stripeライブキー | エラー |
| secrets | 秘密鍵ブロック | エラー |
| secrets | JWT | エラー |
| secrets | インラインパスワード付きデータベースURI | エラー |
| secrets | 一般的な認証情報形式の代入 | エラー |
| secrets | 高エントロピー文字列リテラル | 警告 |
| structure | トップレベル関数が多すぎる(デフォルト上限5) | 警告 |
| structure | 共有された関数名の接頭辞(3関数以上) | 警告 |
| structure | 最初のパラメータ名の繰り返し(3関数以上) | 警告 |
| structure | 関数が長すぎる(デフォルト上限60行) | 警告 |
| structure | 可変モジュールレベルの状態がファイル内の他の場所で変更されている | 警告 |
| comment_slop | コメント内の会話形式の呼びかけ | 警告 |
| comment_slop | 自明なことを説明するコメント | 情報 |
| comment_slop | 短い関数に対する過剰に大きなコメントブロック | 情報 |
| comment_slop | プレースホルダーのスキャフォールドが残っている | 警告 |
| comment_slop | セクション区切りバナーの繰り返し(ファイルあたり4回以上) | 情報 |
各ルールの完全な根拠、偽陰性に関する注意、抑制手順については、docs/RULES.mdを参照してください。
設定
プロジェクトルートに.gristmill.tomlを置きます。すべてのキーはオプションです:
[checks]
enabled = ["secrets", "structure", "comment_slop"]
[structure]
max_top_level_functions = 5
max_function_lines = 60
[secrets]
entropy_threshold = 4.5
[ignore]
paths = ["legacy/**", "vendor/**"]
rules = ["CMT003"].gristmillignoreファイル(gitignore構文)は[ignore] pathsと併用できます。また、フラグが立てられた行またはその上の行でインライン抑制も有効です:
SUPPRESSED = "ghp_" + "..." # gristmill: ignore SEC003// gristmill: ignore SEC003
const suppressed = "ghp_" + "...";言語サポート
Python — 完全サポート(stdlibの
astおよびtokenize)。JavaScript/TypeScript — 完全サポート。
tree-sitterとtree-sitter-javascript、tree-sitter-typescriptのコンパイル済みグラマーを使用します。Nodeベースのパーサーに外部コマンドを発行しません。これにより、コンパイル済みのPython依存関係と引き換えに、ホストにNodeがインストールされている必要がないという独立性を得ています。structureとcomment_slopはnodeがPATHにあるかどうかに依存せずに同じように動作し、テキストのみのフォールバックではなく実際のASTを提供します。その他 —
secretsチェックは引き続き実行されます(正規表現ベースで言語に依存しません)。structureとcomment_slopはそのファイルではスキップされ、skipped_pathsに報告されます。
制限事項
このツールを、その実績以上に信頼する前に読んでください:
secretsは整形された文字列または高エントロピー文字列のみを検出します。hunter2のような低エントロピーの人間のパスワードは決してフラグされません。通常の短い文字列と区別する信頼できる方法がないためです。実行時に組み立てられる認証情報(文字列連結、os.environ.get(...) or "fallback"、Base64デコードされた断片)は、静的なテキストに対する正規表現/エントロピーパスでは不可視です。ファイルをまたがる構造上の問題は不可視です。
structureは一度に1つのファイルを検査します。複数のファイルに分割すべきクラスや、2つの異なるモジュールにある重複ロジックは対象外です。comment_slopのCMT002は意図的に狭い範囲にしています。 これはこのセットの中で最も偽陽性リスクが高いルールであるため、沈黙に大きく偏るように実装されています。実際の説明を見逃す頻度の方が、過剰にフラグするよりもはるかに高くなります。正確な部分一致ルールはdocs/RULES.mdを参照してください。PythonおよびJS/TS以外の言語はsecretsのみのカバレッジになります。 v1ではGo、Rust、Rubyなどの構造分析やコメント分析はありません。
これはgit履歴のシークレットスキャナーではありません。 与えられたワーキングツリーを検査します。コミットされて現在のファイルから削除されたキーは、このツールの関心事ではありません(git履歴スキャナーは別の補完的なツールです)。
自動修正はありません。 Gristmillは報告します。呼び出し元のモデルが何をどのように変更するかを決定します。この分割は意図的です(上記「なぜMCPサーバーなのか、スキルではないのか」を参照)。ただし、
verify呼び出しだけでは何も修正されないことを意味します。
カバレッジを過大評価するツールは、盲点について正直なツールよりも劣ります。ノイズの多い発見と同様に、偽りの自信よりも沈黙が勝ります。
ロードマップ
v1では明示的に範囲外。大まかな優先順位順:
自動修正/パッチ生成(現在は呼び出し元のモデルが
verifyの結果を使用してこれを行っています)依存関係の鮮度とCVEチェック(パッケージレジストリへのネットワーク呼び出しが必要 — 自然なv2)
PythonおよびJavaScript/TypeScriptを超えた言語サポート
コミットされて後で削除されたシークレットのgit履歴スキャン
ホスティングサービス、Web UI、またはダッシュボード
開発
.venv/bin/pip install -e ".[dev]"
.venv/bin/pytest tests/ -qsrc/gristmill/rules.pyを編集した後にdocs/RULES.mdを再生成:
.venv/bin/python3 scripts/generate_rules_doc.pyテストは以下をカバーします(tests/):既知の汚染フィクスチャディレクトリのゴールデンファイル出力、並列実行あり/なしでの10回の決定性、ゼロ発見を生み出さなければならない偽陽性コーパス、編集(生のシークレットが出力フィールドに決して到達しないこと)、および耐障害性(無効な構文、バイナリ、空、巨大なファイルが実行をクラッシュさせないこと)。
ライセンス
MIT — LICENSEを参照してください。
Available Tools
2 toolsexplain_ruleA
Look up a gristmill rule by id (e.g. SEC001, STR002, CMT004): its rationale, what it catches, what it misses, and how to suppress it.
| Name | Required | Description | Default |
|---|---|---|---|
| rule_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It correctly implies a read-only operation ('Look up') and lists the content returned. However, it does not explicitly state that the tool has no side effects, nor does it mention authorization requirements, rate limits, or error handling for invalid rule IDs. A 3 is adequate but leaves gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that is front-loaded with the action and resource, includes concrete examples in parentheses, and conveys the full return intent. There is no wasted text; every part earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the tool's purpose, the parameter, and the core return values. An output schema exists to handle return type details, so the description does not need to reiterate those. However, it omits mention of what happens if the rule ID is invalid or missing (e.g., error or null response), which would make it more complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% (no parameter descriptions in the input schema), so the description must compensate. It adds value by specifying the parameter is a 'rule id' and provides examples (SEC001, STR002, CMT004), hinting at a consistent format. However, it does not fully specify the pattern or acceptable formats, leaving ambiguity for the agent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Look up', identifies the resource as a 'gristmill rule', and specifies exactly what information is returned: rationale, what it catches, what it misses, and how to suppress it. This distinguishes it from the sibling tool 'verify', which likely performs a different function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when you need details about a specific rule (e.g., by its ID), but it does not explicitly state when to use this tool versus the sibling 'verify', nor does it provide guidance on when not to use it or any prerequisites. More specific exclusions or comparisons would improve this dimension.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verifyA
Inspect source files for secrets, structural problems, and low-quality comments. Returns deterministic findings with file paths and line numbers. Call this after generating or editing code, before presenting it as finished.
| Name | Required | Description | Default |
|---|---|---|---|
| paths | Yes | ||
| checks | No | ||
| severity_floor | No | info |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without any annotations, the description must disclose behavioral traits fully. It states that findings are 'deterministic' and include 'file paths and line numbers,' which adds value. However, it does not address permissions, side effects (though likely read-only), rate limits, or what happens when no issues are found.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three efficiently structured sentences: purpose, output nature, and usage timing. No redundant or extraneous content. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The presence of an output schema partially offsets the need to describe return values. However, the description does not explain how the three parameters interact or provide examples for common use cases, leaving gaps for a tool invoked after code generation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only implicitly refers to 'paths' via 'source files' and does not explain 'checks' (the enum options) or 'severity_floor' at all. This forces the agent to rely solely on parameter names, which are insufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('inspect') and resource ('source files') and enumerates three concrete issue types (secrets, structural problems, low-quality comments). With only one sibling tool 'explain_rule', the purpose is clearly distinct.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly tells when to call this tool: 'after generating or editing code, before presenting it as finished.' It provides clear context but does not mention when not to use it or compare to alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.1.0- First observed
explain_rule - First observed
verify
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: one is for static analysis of code, the other for documentation of rules. An agent would not confuse them.
The naming is inconsistent: 'verify' uses a plain verb while 'explain_rule' uses verb_noun pattern. Both are clear in isolation, but the lack of a unified pattern (e.g., 'verify_code' vs 'explain_rule') makes the set feel ad-hoc.
With only 2 tools, the server feels thin for a tool suite called 'gristmill-mcp'. A code analysis server typically needs more tools like listing rules or scanning for specific categories to feel properly scoped.
The server only provides a scan tool and a rule lookup tool, but is missing operations like listing all rules, skipping specific rules, or generating reports. Users cannot discover available rules without knowing their IDs, creating a dead end.
Maintenance
Related MCP Connectors
An MCP server that gives your AI access to the source code and docs of all public github repos
Official DevSpeak MCP server — translate technical text into formal specs from any AI IDE or agent
MCP server teaching AI agents to implement TideCloak: auth, E2EE, IGA, security analysis
Find, compare, and audit software for AI agents. Scored registry of tools and MCP servers.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceAn MCP server that provides AI coding agents with AST-accurate, context-budget-aware codebase querying, safety gates, and team policy integration via structured tools and a local plugin layer.104 npm4MIT
- AlicenseNot gradedqualityCmaintenanceAn MCP server for verifying AI-generated code quality, security, and performance, addressing trust gaps in AI coding assistants.MIT
- FlicenseNot gradedqualityCmaintenanceMCP server that helps AI agents inspect Minecraft project evidence (crash logs, mod files, datapacks) before writing development code.2-
- AlicenseNot gradedqualityDmaintenanceAn MCP server that gives AI coding agents structured access to a project's architecture, rules, modules, and technical decisions.MIT