Skip to main content
Glama

ocr-mcp-server

给 OpenCodeReview 补上确定性分析这条腿。

OpenCodeReview 是一个 LLM 驱动的代码审查 CLI。它的架构是「确定性管线 + LLM Agent」, 但审查结论本身完全来自 LLM——模型没注意到的缺陷,它就漏掉了。项目 README 自己承认 这是有意的取舍:

Recall is explicitly lower than that of general-purpose agents — a deliberate trade-off favoring precision over noise.

本服务通过 OpenCodeReview 官方支持的 MCP 扩展点暴露 Semgrep,让审查 agent 在推理 之外多一条「规则匹配」的路径。规则匹配不是猜测——命中就是命中。


核心设计:只报 diff 里的行

这是本项目最有价值的部分,也是实测逼出来的。

OpenCodeReview 一次只审一个文件的 diff,且它的 prompt 明确禁止评论 diff 之外的 内容(见其 tools.md:「收集该上下文时发现的任何问题按设计都会被忽略」)。

一个报「整个文件」的扫描器因此是有害的:它把大量 agent 必须丢弃的命中塞进上下文, 而漏过去的那部分会变成对一个 PR 从未改动过的代码的评论。所以本服务的默认工具是 semgrep_scan_diff,它先解析 git diff --unified=0 取出新侧改动行区间,扫完再按区间过滤:

base..head  ──►  git diff --unified=0  ──►  {file: [(start, end), ...]}
                                              │
              semgrep --json  ──► findings ──►│──► 只保留落在区间内的
                                              ▼
                                        可被 agent 直接采用的结果

被过滤掉的数量会一并返回(... N match(es) outside the changed lines were suppressed.), 这样「扫描干净」和「扫描被截断」不会被混为一谈。

Related MCP server: scanline

实测:这个方法在 AACR-Bench 上的边界

OpenCodeReview 的官方基准 AACR-Bench 有 196 个 真实 PR、1505 条由 80+ 资深工程师交叉验证的标注。我用它的前 5 个样本实测了确定性分析的 覆盖面:

仓库

参考标注

semgrep 命中

命中位置

vllm-project/vllm

1

0

—

lvgl/lvgl

5

13

12 条在 .github/workflows,与标注无关

FreeCAD/FreeCAD

3

6

全在构建脚本

elastic/elasticsearch

5

0

—

astral-sh/uv

6

未测(仓库未克隆)

—

参考标注所在的文件,semgrep 一条都没报。

原因不是工具失效——同一套代码在一个受控的两提交仓库上,把 SQL 注入、shell=True、 MD5 存密码三个缺陷全部报出(见下方冒烟测试输出)。原因是分布不重叠:

  • AACR-Bench 的标注是人工 reviewer 级的语义评论:方法命名误导、注释与实现不符、 缓存失效、设计冗余。

  • Semgrep 是模式匹配器:它能找的是语法与数据流层面的已知缺陷模式。

因此在这个基准上,直接接入确定性分析不会提升 Recall,反而会显著拉低 Precision。 这个负结果写在这里,是因为它划定了方法的适用边界——把模式匹配器接到语义级评审基准上 是个错配,而这件事在动手之前并不显然。

未完成:在模式匹配能赢的场景(例如「安全修复提交的父提交」这类缺陷检出任务) 上的量化结果尚未测量。本仓库不声称任何 Recall 改进数字。

工具

工具

用途

semgrep_scan_diff

首选。 只扫 base..head 改动的行

semgrep_scan

扫整个文件,用于 diff 之外的自查

semgrep_status

报告分析器是否可用、版本、规则集与限额

返回格式刻意紧凑——每条命中一行,外加可直接用作评论锚点的源码片段:

app.py
  L9       ERROR   sqlalchemy.security.sqlalchemy-execute-raw-query  Avoiding SQL string concatenation...  [CWE-89]
          | return conn.execute("SELECT * FROM users WHERE name = '" + name + "'").fetchone()
  L17      ERROR   security.audit.subprocess-shell-true  Found 'subprocess' function 'run' with 'shell=True'.  [CWE-78]
          | return subprocess.run(f"echo {user_input}", shell=True)

实现中发现的坑

1. --config=auto 与 --metrics=off 互斥

Semgrep 拒绝在关闭遥测时解析 auto 规则集:

Cannot create auto config when metrics are off.
Please allow metrics or run with a specific config.

审查沙箱里不该向外发遥测,所以默认钉死 p/default。副作用是好的:钉死的规则集让 基准可复现,而整个项目的意义就在于测出可比较的差值。

2. Semgrep 的 extra.lines 字段不可信

实测 semgrep 1.177.0:对同一文件里两条毫不相关的命中(shell=True 与 MD5 存密码), extra.lines 都返回同一个字符串 "requires login"——而这个字符串在被扫描的文件里 根本不存在。

所以源码片段不取 Semgrep 的字段,改为按报出的行号自己读文件。多一次文件读取, 换来不依赖一个未文档化、跨版本可能变化的行为。

安装

uv tool install semgrep
git clone <this repo> && cd ocr-mcp-server
uv sync

接入 OpenCodeReview

ocr config set mcp_servers.semgrep.command uvx
ocr config set mcp_servers.semgrep.args '["--from", "/abs/path/to/ocr-mcp-server", "ocr-mcp-server"]'
ocr config set mcp_servers.semgrep.tools '["semgrep_scan_diff", "semgrep_scan", "semgrep_status"]'

工具名刻意带 semgrep_ 前缀:OpenCodeReview 的内置工具与 MCP 工具共享同一命名空间, 重名会被静默跳过(file_read、code_search 等已被占用),不报错。

配置

环境变量

默认值

说明

OCR_MCP_SEMGREP_BIN

semgrep

分析器可执行文件

OCR_MCP_SEMGREP_CONFIG

p/default

规则集。可用 p/security-audit、p/python 或本地规则文件

OCR_MCP_TIMEOUT

120

单次扫描超时(秒)

OCR_MCP_MAX_FINDINGS

50

单次返回的命中上限

测试

uv run pytest -q                      # 56 个单元测试
uv run python scripts/smoke_test.py   # 端到端:经 MCP 协议真实调用

单元测试用桩替换子进程,覆盖 Semgrep 缺失 / 超时 / 输出损坏 / 正常返回四种情况, 并验证路径穿越(../、绝对路径、空字节)被拒绝。test_diff.py 里的解析测试跑在 真实的临时 git 仓库上,而不是手写的 diff 文本。

冒烟测试走真实 stdio 传输启动服务,并在一个真实的两提交仓库上调用工具——它抓到过 单元测试抓不到的 bug:--config=auto 与 --metrics=off 的冲突只在真实子进程里暴露。

技术栈

Python · MCP (Model Context Protocol) · Semgrep · uv · pytest

Available Tools

3 tools
semgrep_scanA

Run deterministic static analysis (Semgrep) over repository files and return findings with exact line numbers plus the matched source snippet.

Use this to substantiate a suspected defect rather than guessing at it: the results are rule matches, so a hit is a fact. The snippet can be reused directly as existing_code when emitting a comment.

Scope paths to the file currently under review. OpenCodeReview only allows comments on the file being reviewed, so findings in other files cannot be reported even when the tool returns them.

An empty result is not a clean bill of health — it means no configured rule matched. The ruleset is broad but not exhaustive.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathsYes
rulesetNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the burden. It discloses determinism, that results are factual rule matches, that an empty result means 'no configured rule matched' rather than clean code, and that the ruleset is nonexhaustive. It does not discuss performance, permissions, or failure behavior, but the most decision-relevant behavior is covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences cover result, use case, path scoping, and interpretation caveat. There is no filler; each sentence adds actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need no elaboration. The description covers when to use, how to scope paths, how to interpret empty results, and how to reuse the snippet. It would be slightly stronger with explicit guidance on the optional ruleset parameter and differentiation from semgrep_scan_diff, but nothing needed to perform a safe scan is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds real semantics for paths ('Scope paths to the file currently under review') and explains why other-file results are unusable. It gives only indirect context for the ruleset parameter ('The ruleset is broad but not exhaustive') without explaining acceptable values or how the default is treated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names a specific operation ('Run deterministic static analysis (Semgrep) over repository files') and a concrete result ('exact line numbers plus the matched source snippet'), so an agent knows what the tool does. It does not explicitly contrast with semgrep_scan_diff or semgrep_status, though 'over repository files' hints at full-repo scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It says when to reach for it: 'Use this to substantiate a suspected defect rather than guessing at it.' It gives a hard scoping rule ('Scope paths to the file currently under review') and explains why ('findings in other files cannot be reported'). It still doesn't name semgrep_scan_diff or semgrep_status as alternatives or state when not to use this tool, so it stops short of full routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

semgrep_scan_diffA

Run deterministic static analysis (Semgrep) restricted to the lines changed between two revisions, and return findings with exact line numbers plus the matched source snippet.

Prefer this over semgrep_scan when reviewing a pull request. OpenCodeReview reviews a diff and only permits comments on changed lines, so a finding on an untouched line cannot be reported — this tool does not return those at all, which keeps the result short and entirely actionable.

base and head are any git revisions (abc123, main, HEAD~1). Pass paths to narrow the scan to specific changed files.

ParametersJSON Schema
NameRequiredDescriptionDefault
baseYes
headYes
pathsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses determinism, diff-scoping, exclusion of findings on untouched lines, and the output format. It does not explicitly state read-only/no side-effect behavior, but 'static analysis' and 'scan' strongly imply it, so the gap is minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three tightly packed sentences: the core action, the key usage guidance with rationale, and parameter clarification. Every sentence earns its place, with the most important distinguishing detail front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter tool with an output schema present, the description covers purpose, selection criteria, parameter semantics, and business context. Nothing needed for an agent to call and interpret this tool correctly is left unaddressed, and the output schema handles return-value details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the description compensates fully: it explains that base and head are 'any git revisions (`abc123`, `main`, `HEAD~1`)' and that paths is used 'to narrow the scan to specific changed files.' This adds meaning beyond the bare schema titles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Run deterministic static analysis (Semgrep)'), the resource ('restricted to the lines changed between two revisions'), and the output shape ('findings with exact line numbers plus the matched source snippet'). It also names the sibling tool it is not, semgrep_scan, making the distinction explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description directly instructs when to use this tool: 'Prefer this over semgrep_scan when reviewing a pull request.' It explains why with the OpenCodeReview diff-comment constraint and contrasts the tool's behavior ('does not return those at all'), leaving no ambiguity about selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

semgrep_statusA

Report whether deterministic analysis is available: the Semgrep binary in use, its version, the active ruleset, and the per-call limits. Call this once if a scan reports that it is unavailable.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must carry the behavioral disclosure burden. It clearly describes the nature of the operation as reporting status information and lists the specific elements (binary, version, active ruleset, per-call limits). However, it does not explicitly state whether this call has any side effects, consumes scan quota, or requires authentication, though 'report' strongly implies a read-only operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise: two sentences with no fluff. The primary purpose is stated first, followed by a succinct list of the reported items, and a clear usage instruction. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (no parameters), an output schema exists, and the description explains both what the tool reports and the exact situation in which to call it. Nothing essential is missing for an agent to invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the input schema is empty. The baseline for parameterless tools is 4, and the description correctly does not need to document any parameters. It adds no parameter information because none exists.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool reports whether deterministic analysis is available, with a specific verb ('Report') and a well-defined resource (Semgrep status/availability). It enumerates the exact information returned (binary, version, ruleset, per-call limits), making it easy to distinguish from the sibling scan tools semgrep_scan and semgrep_scan_diff.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Call this once if a scan reports that it is unavailable,' providing a clear trigger condition. It implies this tool is for diagnostic status checks, not for performing scans, but does not explicitly name alternatives or state when not to use it beyond the trigger condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observedsemgrep_scan
    • First observedsemgrep_scan_diff
    • First observedsemgrep_status

TDQS

A4.4/5.0

Scored across 3 tools

Disambiguation5/5

All three tools have clearly distinct purposes: one scans an entire repository, one scans only changed lines, and one reports tooling status. The descriptions explicitly clarify when to prefer scan_diff over scan, eliminating meaningful ambiguity.

Naming Consistency5/5

Every tool follows the same snake_case `semgrep_*` prefix pattern, making the family instantly recognizable. The verb-like suffixes `_scan`, `_scan_diff`, and `_status` are consistent and predictable.

Tool Count5/5

Three tools is well-scoped for a focused analysis server: two scanning modes plus a status endpoint cover the core workflow without redundancy. Each tool earns its place; adding more would likely dilute the server's purpose.

Completeness5/5

The tool surface covers full-repo scanning, PR/diff scanning, and availability reporting, with no obvious dead ends. The status tool fills the operational gap for diagnosing unavailable scans, making the workflow self-contained.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Security scanning MCP server. Semgrep integration, SARIF parsing, baseline diffing, framework-aware ruleset selection, and automated finding triage.
    4 npm
    1
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    MCP server for static analysis of multi-tenant SaaS and MCP server code, catching cross-tenant data leakage. Provides tools to scan code, list/explain rules, and manage suppressions.
    6,036 npm
    4
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    MCP server for persistent code exploration and diff-aware review, exposing explore and review tools to map call paths, blast radius, and generate explainable risk-scored reports.
    14 npm
    2
    MIT