Skip to main content
Glama
mattshuttle

gristmill-mcp

by mattshuttle

gristmill-mcp

AIが生成したコードを検査し、構造的・セキュリティ上の違反を決定的に列挙するMCPサーバーです。AIコーディングエージェントがコードを確定させる前に自身の出力を修正できるようにします。

グリストとは製粉所に運ばれる穀物のことです。AIの出力はグリスト — 確かに価値のある原料ですが、未加工です。製粉所がそれに構造を与えます。

AIがグリストを書く。Gristmillがそれを出荷可能なコードにする。

なぜMCPサーバーなのか、スキルではないのか

スキルはモデルのコンテキストに読み込まれるテキストであり、モデルが知っていることを変えます。MCPサーバーはモデルが実行するプログラムであり、モデルができることを変えます。

スタイルのガイダンス(「クラスを緩やかな関数より優先する」)はスキルに属します。検証(「このファイルには12行目、40行目、66行目…に7つのトップレベル関数があります」)はファイルに対してコードを実行する必要があります。モデルが自身の出力を読んで「これには関数が多すぎるように見える」と推論するのは、観察に擬態した推測にすぎません — このファイルにおける「多すぎる」の意味の真実はなく、信頼できるカウント方法もありません。GristmillはASTを解析して実際に数えます。指示と実行のこの区別こそが、これがアドバイスの段落ではなくサーバーとして存在する理由です。

サーバーは決してLLMを呼び出さず、同じ入力に対して実行間で変化することはなく、信頼スコアを出力することもありません。同じ入力 → 毎回バイト単位で同一の出力。その決定性こそが製品のすべてです。AI層はこのサーバーの上にあり、その結果を消費し、何をすべきかを決定します。サーバーの役割は行番号付きの事実を報告することで終わります。

Related MCP server: code-verify-mcp

インストール

git clone <this repo> gristmill-mcp
cd gristmill-mcp
python3 -m venv .venv
.venv/bin/pip install -e .

Claude Code

venvのコンソールスクリプトを指すようにCLIで登録します:

claude mcp add gristmill -- /absolute/path/to/gristmill-mcp/.venv/bin/gristmill-mcp

または、MCP設定に直接追加します(プロジェクト内の.mcp.json、またはグローバルのClaude Code設定):

{
  "mcpServers": {
    "gristmill": {
      "command": "/absolute/path/to/gristmill-mcp/.venv/bin/gristmill-mcp"
    }
  }
}

その他のMCPクライアント

stdioベースのMCPクライアントは同じバイナリを起動できます。gristmill-mcp(またはvenv内でpython3 -m gristmill.server)は標準のMCP stdioトランスポートをクライアント固有の設定なしで使用します。

コマンドライン(MCPクライアントなし)

ローカルテスト用、または以下の実際の例を再現するために、薄いCLIが同じエンジンをラップしています:

.venv/bin/gristmill-verify path/to/file_or_dir [--checks secrets structure comment_slop] [--severity-floor warning] [--json]

実際の例

demo/billing.py、Stripe課金ヘルパーの未編集の初稿:

import stripe

# I've added this as you requested — sets up the Stripe client
STRIPE_SECRET_KEY = None  # was a literal sk_live_... key — see note below

stripe.api_key = STRIPE_SECRET_KEY


def customer_create(config):
    return stripe.Customer.create(**config)


def customer_delete(config):
    return stripe.Customer.delete(config["id"])


def customer_find(config):
    return stripe.Customer.retrieve(config["id"])


def customer_update(config):
    return stripe.Customer.modify(config["id"], **config)
.venv/bin/gristmill-verify demo/billing.py

上記のNoneの代わりに実際のStripeライブキー形式のリテラルを入れた場合の出力:

gristmill: 1 files scanned, 0 skipped (2 error, 4 warning, 0 info) in 1ms
  [WARNING] STR002  billing.py:1  4 top-level functions share the prefix `customer_` — consider a `Customer` class or module
  [WARNING] STR003  billing.py:1  4 top-level functions take a first parameter named `config` — consider making it instance state
  [WARNING] CMT001  billing.py:3  Comment addresses the reader conversationally ('as you requested')
  [ERROR  ] SEC006  billing.py:4:22  Stripe live key assigned to `STRIPE_SECRET_KEY`
  [ERROR  ] SEC010  billing.py:4:22  String literal assigned to `STRIPE_SECRET_KEY`, which looks credential-shaped
  [WARNING] SEC011  billing.py:4:22  High-entropy string literal (5.1 bits/char) assigned to `STRIPE_SECRET_KEY`

(ファイルパスは最も近い.gristmill.tomlからの相対パスで表示されます。demo/は独自の設定を持っているため、この例の出力はトップレベルのプロジェクト設定に依存しません。)

注釈: GitHubのプッシュ保護は、実際の形式のシークレットを含むファイル(コメント内、Markdownコードブロック内、このREADMEを含む)のプッシュをブロックします。現在demo/billing.pyはキーをNoneに置き換えて初期プッシュを解除しています。これは(許可リストに登録されたシークレットスキャン例外を介して)復元し、デモを再び動作させるためのTODOです。

--jsonフラグ(または両方を返すverify MCPツール)は、完全な構造化形式(ファイル、行、列、静的な提案文字列、および編集されたevidenceフィールド(sk_l…(49文字)、キー自体は決して表示されません))を提供します。

ツール

verify

ソースファイルを検査し、シークレット、構造上の問題、質の低いコメントを検出します。ファイルパスと行番号付きの決定的な結果を返します。コードを生成または編集した後、完成品として提示する前に呼び出してください。

入力:paths(ファイルまたはディレクトリ、必須)、checks(オプション、secrets/structure/comment_slopのサブセット、デフォルトはすべて)、severity_floor(オプション、デフォルトはinfo)。

出力:コンパクトな人間可読サマリーと、それに続く完全な構造化JSON(ファイル、行、列、メッセージ、編集された証拠、ルールごとの静的な提案文字列)。結果は常にpath→line→rule_idでソートされます。この安定性により、実行間でバイト単位の同一性が保証され、モデルが問題に直接移動できるようになります。

explain_rule

rule_id(例:SEC001)を受け取り、その根拠、検出対象、検出できないもの、抑制方法を返します。これはdocs/RULES.mdと同じ内容をオンデマンドで提供し、verifyの出力を簡潔に保つことができます。

ルール

ルール

チェック

タイトル

デフォルトの重要度

SEC001

secrets

AWSアクセスキーID

エラー

SEC002

secrets

AWSシークレットアクセスキー

エラー

SEC003

secrets

GitHubトークン

エラー

SEC004

secrets

Google APIキー

エラー

SEC005

secrets

Slackトークン

エラー

SEC006

secrets

Stripeライブキー

エラー

SEC007

secrets

秘密鍵ブロック

エラー

SEC008

secrets

JWT

エラー

SEC009

secrets

インラインパスワード付きデータベースURI

エラー

SEC010

secrets

一般的な認証情報形式の代入

エラー

SEC011

secrets

高エントロピー文字列リテラル

警告

STR001

structure

トップレベル関数が多すぎる(デフォルト上限5)

警告

STR002

structure

共有された関数名の接頭辞(3関数以上)

警告

STR003

structure

最初のパラメータ名の繰り返し(3関数以上)

警告

STR004

structure

関数が長すぎる(デフォルト上限60行)

警告

STR005

structure

可変モジュールレベルの状態がファイル内の他の場所で変更されている

警告

CMT001

comment_slop

コメント内の会話形式の呼びかけ

警告

CMT002

comment_slop

自明なことを説明するコメント

情報

CMT003

comment_slop

短い関数に対する過剰に大きなコメントブロック

情報

CMT004

comment_slop

プレースホルダーのスキャフォールドが残っている

警告

CMT005

comment_slop

セクション区切りバナーの繰り返し(ファイルあたり4回以上)

情報

各ルールの完全な根拠、偽陰性に関する注意、抑制手順については、docs/RULES.mdを参照してください。

設定

プロジェクトルートに.gristmill.tomlを置きます。すべてのキーはオプションです:

[checks]
enabled = ["secrets", "structure", "comment_slop"]

[structure]
max_top_level_functions = 5
max_function_lines = 60

[secrets]
entropy_threshold = 4.5

[ignore]
paths = ["legacy/**", "vendor/**"]
rules = ["CMT003"]

.gristmillignoreファイル(gitignore構文)は[ignore] pathsと併用できます。また、フラグが立てられた行またはその上の行でインライン抑制も有効です:

SUPPRESSED = "ghp_" + "..."  # gristmill: ignore SEC003
// gristmill: ignore SEC003
const suppressed = "ghp_" + "...";

言語サポート

  • Python — 完全サポート(stdlibのastおよびtokenize)。

  • JavaScript/TypeScript — 完全サポート。tree-sitterとtree-sitter-javascript、tree-sitter-typescriptのコンパイル済みグラマーを使用します。Nodeベースのパーサーに外部コマンドを発行しません。これにより、コンパイル済みのPython依存関係と引き換えに、ホストにNodeがインストールされている必要がないという独立性を得ています。structureとcomment_slopはnodeがPATHにあるかどうかに依存せずに同じように動作し、テキストのみのフォールバックではなく実際のASTを提供します。

  • その他 — secretsチェックは引き続き実行されます(正規表現ベースで言語に依存しません)。structureとcomment_slopはそのファイルではスキップされ、skipped_pathsに報告されます。

制限事項

このツールを、その実績以上に信頼する前に読んでください:

  • secretsは整形された文字列または高エントロピー文字列のみを検出します。 hunter2のような低エントロピーの人間のパスワードは決してフラグされません。通常の短い文字列と区別する信頼できる方法がないためです。実行時に組み立てられる認証情報(文字列連結、os.environ.get(...) or "fallback"、Base64デコードされた断片)は、静的なテキストに対する正規表現/エントロピーパスでは不可視です。

  • ファイルをまたがる構造上の問題は不可視です。 structureは一度に1つのファイルを検査します。複数のファイルに分割すべきクラスや、2つの異なるモジュールにある重複ロジックは対象外です。

  • comment_slopのCMT002は意図的に狭い範囲にしています。 これはこのセットの中で最も偽陽性リスクが高いルールであるため、沈黙に大きく偏るように実装されています。実際の説明を見逃す頻度の方が、過剰にフラグするよりもはるかに高くなります。正確な部分一致ルールはdocs/RULES.mdを参照してください。

  • PythonおよびJS/TS以外の言語はsecretsのみのカバレッジになります。 v1ではGo、Rust、Rubyなどの構造分析やコメント分析はありません。

  • これはgit履歴のシークレットスキャナーではありません。 与えられたワーキングツリーを検査します。コミットされて現在のファイルから削除されたキーは、このツールの関心事ではありません(git履歴スキャナーは別の補完的なツールです)。

  • 自動修正はありません。 Gristmillは報告します。呼び出し元のモデルが何をどのように変更するかを決定します。この分割は意図的です(上記「なぜMCPサーバーなのか、スキルではないのか」を参照)。ただし、verify呼び出しだけでは何も修正されないことを意味します。

カバレッジを過大評価するツールは、盲点について正直なツールよりも劣ります。ノイズの多い発見と同様に、偽りの自信よりも沈黙が勝ります。

ロードマップ

v1では明示的に範囲外。大まかな優先順位順:

  • 自動修正/パッチ生成(現在は呼び出し元のモデルがverifyの結果を使用してこれを行っています)

  • 依存関係の鮮度とCVEチェック(パッケージレジストリへのネットワーク呼び出しが必要 — 自然なv2)

  • PythonおよびJavaScript/TypeScriptを超えた言語サポート

  • コミットされて後で削除されたシークレットのgit履歴スキャン

  • ホスティングサービス、Web UI、またはダッシュボード

開発

.venv/bin/pip install -e ".[dev]"
.venv/bin/pytest tests/ -q

src/gristmill/rules.pyを編集した後にdocs/RULES.mdを再生成:

.venv/bin/python3 scripts/generate_rules_doc.py

テストは以下をカバーします(tests/):既知の汚染フィクスチャディレクトリのゴールデンファイル出力、並列実行あり/なしでの10回の決定性、ゼロ発見を生み出さなければならない偽陽性コーパス、編集(生のシークレットが出力フィールドに決して到達しないこと)、および耐障害性(無効な構文、バイナリ、空、巨大なファイルが実行をクラッシュさせないこと)。

ライセンス

MIT — LICENSEを参照してください。

Available Tools

2 tools
explain_ruleA

Look up a gristmill rule by id (e.g. SEC001, STR002, CMT004): its rationale, what it catches, what it misses, and how to suppress it.

ParametersJSON Schema
NameRequiredDescriptionDefault
rule_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It correctly implies a read-only operation ('Look up') and lists the content returned. However, it does not explicitly state that the tool has no side effects, nor does it mention authorization requirements, rate limits, or error handling for invalid rule IDs. A 3 is adequate but leaves gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that is front-loaded with the action and resource, includes concrete examples in parentheses, and conveys the full return intent. There is no wasted text; every part earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the tool's purpose, the parameter, and the core return values. An output schema exists to handle return type details, so the description does not need to reiterate those. However, it omits mention of what happens if the rule ID is invalid or missing (e.g., error or null response), which would make it more complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% (no parameter descriptions in the input schema), so the description must compensate. It adds value by specifying the parameter is a 'rule id' and provides examples (SEC001, STR002, CMT004), hinting at a consistent format. However, it does not fully specify the pattern or acceptable formats, leaving ambiguity for the agent.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Look up', identifies the resource as a 'gristmill rule', and specifies exactly what information is returned: rationale, what it catches, what it misses, and how to suppress it. This distinguishes it from the sibling tool 'verify', which likely performs a different function.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when you need details about a specific rule (e.g., by its ID), but it does not explicitly state when to use this tool versus the sibling 'verify', nor does it provide guidance on when not to use it or any prerequisites. More specific exclusions or comparisons would improve this dimension.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verifyA

Inspect source files for secrets, structural problems, and low-quality comments. Returns deterministic findings with file paths and line numbers. Call this after generating or editing code, before presenting it as finished.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathsYes
checksNo
severity_floorNoinfo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without any annotations, the description must disclose behavioral traits fully. It states that findings are 'deterministic' and include 'file paths and line numbers,' which adds value. However, it does not address permissions, side effects (though likely read-only), rate limits, or what happens when no issues are found.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three efficiently structured sentences: purpose, output nature, and usage timing. No redundant or extraneous content. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The presence of an output schema partially offsets the need to describe return values. However, the description does not explain how the three parameters interact or provide examples for common use cases, leaving gaps for a tool invoked after code generation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It only implicitly refers to 'paths' via 'source files' and does not explain 'checks' (the enum options) or 'severity_floor' at all. This forces the agent to rely solely on parameter names, which are insufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('inspect') and resource ('source files') and enumerates three concrete issue types (secrets, structural problems, low-quality comments). With only one sibling tool 'explain_rule', the purpose is clearly distinct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells when to call this tool: 'after generating or editing code, before presenting it as finished.' It provides clear context but does not mention when not to use it or compare to alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.1.0
    • First observedexplain_rule
    • First observedverify

TDQS

A3.5/5.0

Scored across 2 tools

Disambiguation4/5

The two tools have clearly distinct purposes: one is for static analysis of code, the other for documentation of rules. An agent would not confuse them.

Naming Consistency3/5

The naming is inconsistent: 'verify' uses a plain verb while 'explain_rule' uses verb_noun pattern. Both are clear in isolation, but the lack of a unified pattern (e.g., 'verify_code' vs 'explain_rule') makes the set feel ad-hoc.

Tool Count2/5

With only 2 tools, the server feels thin for a tool suite called 'gristmill-mcp'. A code analysis server typically needs more tools like listing rules or scanning for specific categories to feel properly scoped.

Completeness2/5

The server only provides a scan tool and a rule lookup tool, but is missing operations like listing all rules, skipping specific rules, or generating reports. Users cannot discover available rules without knowing their IDs, creating a dead end.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers