Skip to main content
Glama

Culprit — автоматический git bisect с изоляцией и точным доказательством

Знакомая боль: что-то сломалось, а виновный коммит потерялся среди сотен изменений. Ручной git bisect работает, но дёргает вашу рабочую копию туда- сюда на каждом шаге и требует вручную гонять тесты. Culprit делает то же самое — полностью изолированно (через git worktree, ваша рабочая копия не трогается вообще) и резюмируемо между вызовами.

Доказано на настоящем репозитории, не гипотетически

15 реальных коммитов, баг внедрён ИМЕННО в commit 9

start_bisect(good=v-good, bad=HEAD, check="<your interpreter> -c '...'")
run_bisect(session_id)

  …  GOOD / BAD …

CULPRIT FOUND after ≤5 test(s): …  "commit 9"

Вместо перебора кандидатов — ⌈log₂ n⌉ проб. HEAD и рабочая копия после прогона не меняются — это проверяет интеграционный тест, а не README.

Related MCP server: culprit

Ключевая инженерная деталь: git worktree, а не checkout

Наивная автоматизация делает git checkout <commit> в рабочей копии — переключает ветку, конфликтует с грязным деревом, оставляет detached HEAD. Culprit для каждого кандидата поднимает изолированный git worktree и сразу сносит его. Ваш checkout в процессе не участвует.

Алгоритм: бинарный поиск с обработкой SKIP

Ядро — чистая BisectAlgorithm.advance() без git и I/O:

  • SKIP на середине → берётся ближайший непротестированный сосед;

  • весь диапазон SKIP → честный ambiguous, без бесконечного цикла;

  • результаты можно копить в любом порядке — границы пересчитываются из полной истории outcomes.

Код возврата 125 = SKIP, как у git bisect run. Таймаут check-команды тоже становится SKIP.

Архитектура (Clean Architecture)

src/culprit/
├── entities/          # CommitRef, BisectAlgorithm, BisectSession
├── use_cases/         # start / step / run-to-completion / status
├── interfaces/        # порты (Git, Runner, SessionRepository, Clock, Id)
├── infrastructure/    # real git worktree, subprocess+timeout, sqlite
└── presentation/      # composition root, MCP tools, response formatting

Почему здесь есть subprocess

В отличие от Loom/Ward/Covenant, subprocess — суть инструмента (как у git bisect run). Trust boundary: check_command задаёт человек (тест/скрипт, который вы уже доверяете). Агент/LLM не должен изобретать shell-пейлоады — только передать вашу команду в start_bisect.

Линейность истории

Кандидаты берутся через git log --first-parent good..bad: бинарный поиск идёт по mainline, а не по произвольному merge-DAG. Для типичного «сломалось на main» этого достаточно; полный обход всех merge-веток не поддерживается.

Установка

git clone <этот репозиторий>
cd culprit
pip install -e ".[dev]"

Проверка качества (lint, типы, тесты, покрытие, стена Мартина):

.\make.ps1

Подключение к Claude Desktop / Claude Code

{
  "mcpServers": {
    "culprit": {
      "command": "culprit-mcp",
      "env": {
        "CULPRIT_DB": "~/.culprit/sessions.db",
        "CULPRIT_CHECK_TIMEOUT_SECONDS": "300"
      }
    }
  }
}

~ в CULPRIT_DB раскрывается через os.path.expanduser (и для env, и для default). Невалидный timeout тихо откатывается к 300 секундам.

На Windows укажите полный путь к culprit-mcp.exe или к python -m culprit.presentation.mcp_server, если Scripts не в PATH.

check_command: кроссплатформенно

Используйте интерпретатор, который точно есть на машине:

# Linux / macOS
python3 -c "…"

# Windows (или везде, если venv активирован)
python -c "…"

Контракт: 0 — бага нет (GOOD), ненулевой — бага есть (BAD), 125 — SKIP. Таймаут check-команды тоже = SKIP. Команду передаёт человек, не модель.

Пример разговора

Вы: Начиная с v2.3.0 тест test_login_timeout стал падать. Найди, какой коммит это сломал.

Claude: [start_bisect → run_bisect] Нашёл: a3f8c21 "Refactor session timeout handling". Рабочая копия не трогалась.

MCP Tools

Tool

Назначение

start_bisect(repo_path, good_ref, bad_ref, check_command)

Начать сессию

step_bisect(session_id)

Один кандидат в worktree

run_bisect(session_id)

До конца (≤50 шагов)

bisect_status(session_id)

Состояние сессии

list_bisect_sessions()

Все сессии

Ошибки инструментов приходят строкой error: … — удобно читать в чате.

Тесты

pytest -v

Покрытие: чистый алгоритм, use cases на фейках (включая timeout→SKIP), форматирование MCP-ответов, subprocess timeout, реальный git с внедрённым багом и доказательством, что HEAD/файлы не изменились.

Лицензия

MIT — см. файл LICENSE.

Available Tools

5 tools
bisect_statusC

Show status, candidate count, and culprit if found.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It never states that the operation is read-only, whether polling affects bisect state, or any permission/rate constraints; the only content it offers (returned fields) is already covered by the output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, no filler, and the most useful information (status/candidate count/culprit) is front-loaded. It is efficient, though arguably too terse to carry the missing behavioral and parameter context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because an output schema exists, the description need not explain return values, and for a one-parameter status check this is minimally viable. However, with no annotations and no usage or parameter context, it leaves meaningful gaps for an agent deciding when and how to call it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single parameter, so the description is responsible for compensating and does not mention session_id or what a valid session id is. The field name is self-explanatory, which softens the gap, but no added meaning is provided.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (show) and concrete bisect concepts (status, candidate count, culprit), making the tool's role legible without opening the schema. It does not, however, differentiate itself from siblings like list_bisect_sessions or step_bisect.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is offered. With siblings such as step_bisect, run_bisect, and list_bisect_sessions, an agent would benefit from knowing this is a non-mutating poll of a single session, but nothing is stated about timing or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_bisect_sessionsB

List all stored bisect sessions.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden and does not meet it. It says nothing about ordering, pagination, whether completed/aborted sessions are included, or memory/scale constraints for a listing operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short, front-loaded sentence with no filler. It is efficient, though its brevity is part of why behavioral and usage details are absent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the parameter surface is empty. Still, for a discovery tool surrounded by four related bisect tools, the description omits the routing/usage context an agent needs to pick it over bisect_status.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; the schema baseline of 4 applies. No parameter-level meaning is missing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List all stored bisect sessions'), which is clear and unambiguous. It does not, however, contrast itself with the siblings (start_bisect, step_bisect, run_bisect, bisect_status), so the agent gets no explicit differentiation beyond the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of alternatives such as bisect_status, which is a potential sibling for inspecting bisect progress. The agent must infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_bisectB

Drive step_bisect until done/failed (max 50 steps).

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose one meaningful behavioral trait: the 50-step cap. It omits other key traits, such as whether the call blocks, what happens when the cap is reached, and whether it mutates or locks the session.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence with the terminal condition and step bound front-loaded; no wasted words. It is arguably over-terse, but that gap is captured under completeness rather than structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation. But for a driver-loop tool in a multi-tool bisect workflow, the description leaves critical questions unanswered: blocking behavior, post-cap outcome, and session state expectations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The sole required parameter (session_id) has 0% schema description coverage and is not explained anywhere in the description. An agent must guess that it refers to a session created by start_bisect; the description adds no compensating meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action (drive step_bisect to completion) with an explicit terminal condition (done/failed), which clearly separates it from the single-step sibling step_bisect. It does not explicitly contrast itself with start_bisect or bisect_status, but the resource and verb are unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'until done/failed' implies this is a convenience loop that replaces calling step_bisect repeatedly, so the usage is inferable. However, there is no explicit when-to-use statement, no mention of when to prefer manual step_bisect calls, and no prerequisites noted.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_bisectA

Start an isolated first-parent bisect. check_command MUST be supplied by the human (never invent shell). Exits: 0=good, non-zero=bad, 125=skip. Working tree untouched — each probe uses a fresh git worktree.

ParametersJSON Schema
NameRequiredDescriptionDefault
bad_refYes
good_refYes
repo_pathYes
check_commandYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does it well: it discloses exit-code interpretation (0=good, non-zero=bad, 125=skip), a key side-effect guarantee ('Working tree untouched'), and the implementation detail that each probe uses a fresh git worktree. It omits auth/permission or concurrency caveats, but for this tool the safety-relevant behavior is well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, each earning its place: identity, hard constraint, exit-code contract, and side-effect guarantee. The most decision-relevant information (what it starts, the human-supplied check rule) is front-loaded with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the isolation guarantee and exit semantics. What is missing is the workflow context: it never states that a session is created and must then be driven via step_bisect/run_bisect or inspected with bisect_status, which is important for a stateful session-start tool with no annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% across 4 required parameters, so the description must compensate. It gives strong semantics for the riskiest parameter (check_command must be human-supplied, never invented shell) but says nothing about good_ref/bad_ref resolution or repo_path expectations, leaving three parameters defined only by their titles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Start ... bisect') plus the qualifier 'isolated first-parent', which signals the git bisect --first-parent behavior and its worktree-based isolation. It does not explicitly contrast itself with step_bisect/run_bisect, so sibling differentiation is only implied by the word 'Start'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance: nothing says this must be called before step_bisect/run_bisect, nor when a full run_bisect would be preferable. The only constraint given is about check_command provenance, which is a safety rule rather than alternative-selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

step_bisectC

Test one candidate in an isolated worktree; resume-safe across calls.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose two meaningful traits: execution happens in an isolated worktree (containment) and the call is resume-safe (idempotent/re-runnable). It stops short of saying what state it mutates, what the worktree isolation protects, or any limits, so the disclosure is partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence with the key qualifiers front-loaded and no filler. It is efficiently sized, though its brevity is also the source of the coverage gaps noted elsewhere.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the isolation/resume notes cover the core behaviour. What is missing is workflow positioning against the four sibling tools, which for a multi-step bisect process is material context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single parameter (session_id) with 0% schema description coverage, and the description adds nothing about its expected value, format, or provenance. The parameter is self-evident by name, which keeps this from being worse, but the description contributes no semantic value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a concrete action (test one candidate) and implies the resource (a bisect session), so the purpose is recognisable. However, it does not distinguish this step operation from the sibling run_bisect or explain its place in the start_bisect/run_bisect/bisect_status workflow, leaving the exact role ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Resume-safe across calls' hints that it can be re-invoked, but there is no explicit guidance on when to use step_bisect versus run_bisect or start_bisect, nor any prerequisites or ordering constraints. The agent must infer the workflow entirely from sibling names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.1.0
    • First observedbisect_status
    • First observedlist_bisect_sessions
    • First observedrun_bisect
    • First observedstart_bisect
    • First observedstep_bisect

TDQS

B3.4/5.0

Scored across 5 tools

Disambiguation5/5

Each tool has a clearly distinct role: start_bisect initializes a bisect, step_bisect tests one candidate, run_bisect drives stepping to completion, bisect_status reports the current session, and list_bisect_sessions enumerates stored sessions. There is no overlap that would cause misselection.

Naming Consistency4/5

start_bisect, step_bisect, and run_bisect follow a consistent verb_bisect pattern, but bisect_status inverts to noun_verb and list_bisect_sessions inserts bisect in the middle. The deviations are minor and the names remain readable and predictable.

Tool Count5/5

Five tools is well-scoped for git bisect automation, with each tool earning its place across initialization, stepping, driving, status, and session listing. No obvious bloat or thinness.

Completeness4/5

The core bisect lifecycle is covered: start, step, run, status, and list sessions. Minor gaps exist, such as no explicit abort/cleanup or delete-session operation, but agents can work around this by ignoring or overwriting sessions.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    A
    maintenance
    Version-control-native project memory for AI agents: an encrypted knowledge graph whose beliefs are staleness-checked deterministically against repo history — belief bisect names the commit that invalidated a fact, and agents write their own evidence-anchored beliefs via propose_belief. Also exposes changeset proposals and per-changeset token cost reporting; bridges to plain Git.
    9
    2
    MIT