Skip to main content
Glama

Alibi

CI PyPI License: MIT Python 3.12+

We check whether your agent has an alibi for what it did.

Alibi finds where a long AI-agent run went wrong. It borrows the playbook of automotive driver-assistance systems: treat the run as a time series, filter it forward, detect the fault, then smooth backward to the moment it began. The only sensor is Jev, TypeSafe AI's System One model, which answers with calibrated probabilities instead of text.

It is built for long traces, where reading everything at once breaks down: on real coding failures of 57K to 120K tokens, a single whole-trace read found the root-cause step in 0% of cases; Alibi found it in 16% and 7.8%, and points to the right neighbourhood (within 3 steps) in 21% and 17%.

How it works

Stage

What happens

Driver-assistance analogue

Chapters

Overlapping token windows, read one at a time. The size is ALIBI_JUDGE_WINDOW_TOKENS, default 10,000 - what every measurement below used

Measurement frames

Forward filter

Jev rates health and 4 warning signs per chapter; a memory card of numbers carries state forward

Recursive filter (predict + update)

CUSUM

Accumulates health drift; the first alarm marks the failure chapter

Fault detection

Look-back

Re-reads the alarm chapter and every earlier one in parallel, with hindsight

Fixed-interval (RTS-style) smoothing

Pick

P(chapter) x evidence per step; the earliest step within 80% of the top score

Fault-onset estimation

Related MCP server: AgentDelta

What you get

A short trace is refused, without spending a call:

$ alibi diagnose examples/sample_trace.json
Trace is 219 tokens, under the 50,000-token threshold: a single direct read is
enough; Alibi adds value on long traces.

A long one comes back as three steps to read, in order. This is a real run from the locked TrajErrBench set (a Claude Opus coding agent failing on a qutebrowser issue), replayed from its recorded result:

$ alibi diagnose trace.json
Read these 3 steps first, in order.
80,778 tokens, 10 chapters, alarm at chapter 6; 17 Jev calls, 35 s, $0.0096
1. step 74 (assistant), chapter 6, score 0.277
   'Now I see the full picture. The test on line 458 expects
    `str(proc.outcome) == 'Testprocess crashed.'` for SIGSEGV...'
2. step 76 (assistant), chapter 6, score 0.202
   '## Phase 5: FIX ANALYSIS\n\nNow I have a clear understanding. Let me
    implement the changes to `guiprocess.py`...'
3. step 78 (assistant), chapter 6, score 0.178
   'Now let me implement all the changes:\nTool calls:\nstr_replace_editor(...'

Step 74 is the labelled root cause: the agent reads the test wrong and every later edit builds on that reading. Being right at rank 1 happens on 16% of these traces; the honest claim is that three steps out of 118 is a much smaller haystack.

Results

Locked test sets, each run once against a pass bar written down beforehand:

Test set

Traces

Median length

Alibi exact step

Alibi within 3 steps

Whole-trace read, exact

TrajErrBench SWE-Bench Pro (real coding failures)

56

57K tokens

16%

21%

0% (p = 0.004)

LongRCA SWE-bench Pro (real coding failures)

90

120K tokens

7.8%

16.7%

0% (p = 0.016)

LongRCA WebArena (real web-task failures)

48

38K tokens

12.5%

20.8%

16.9% published (no clear difference)

Against published methods on the LongRCA leaderboard (all on DeepSeek-V4-Flash; exact root step, same 128 SWE-bench Pro failures):

Method

Exact root step

RCTA

38.3%

Alibi (Jev)

10.2%

ECHO

7.8%

FALAT

2.3%

All-at-once (whole trace)

1.6%

Step-by-step

0.8%

Binary search

0.8%

Alibi ties ECHO (Fisher p = 0.66) and trails RCTA (p < 0.000001), which traces each suspect back to the handoff instruction between agents.

When to use it, and when not to

  • Use it for traces of tens of thousands of tokens or more (default gate: 50K tokens).

  • Do not use it for short traces: a single direct read by any capable model is enough, and Alibi says so without spending a call.

  • Treat the output as "read these 3 steps first", not as a verdict.

Speed and cost

  • Each chapter is one Jev call, and chapters are judged in parallel on the way back. A 38K-token trace is 4.6 chapters and 15 seconds of judge time; a 120K-token trace is 17 chapters and about 48 seconds. No reasoning tokens are generated.

  • Cost scales with trace length: about $0.08 per million trace tokens at Jev's price of $0.042 per million input tokens, since every token is read about twice (forward, then the look-back). That is roughly $0.01 for a 100K-token trace, and $0.02 for a 250K one.

Install and run

Claude Code.

/plugin marketplace add ahmedezz26/alibi
/plugin install alibi@alibi

Set TYPESAFE_API_KEY in the shell you start Claude Code from. The plugin does the rest: it starts the MCP server with uvx and sets the other two variables itself. Then ask it "why did my last session go wrong?"

Command line. The PyPI distribution is agent-alibi; the import package is alibi.

uv tool install agent-alibi
export ALIBI_JUDGE_BACKEND=typesafe ALIBI_ALLOW_PAID_MODELS=1 TYPESAFE_API_KEY=...
alibi diagnose path/to/trace.json

All three variables are required - Jev is a paid API and ALIBI_ALLOW_PAID_MODELS is the spend guard. Miss one and the run stops at exit 2 naming what is missing, having sent nothing. A .env beside the directory you run from is read for any of them you did not set.

Short traces need no key at all, so you can check the install before you have one:

$ printf '[{"type":"user","inputs":{"q":"hi"}}]' > /tmp/t.json
$ alibi diagnose /tmp/t.json
Trace is 9 tokens, under the 50,000-token threshold: a single direct read is enough;
Alibi adds value on long traces.

What you can point it at

One trace per run. --source auto, the default, picks by file extension.

A JSON file (.json) - a list of steps, oldest first. Every field is optional, but include outputs and error where you have them: a misread observation or an unfixed error is most of the signal the method looks for.

[
  {"type": "llm",  "name": "plan",           "inputs": {"task": "Book the cheapest direct flight."},
                                             "outputs": {"text": "Plan: search, filter, pick."}},
  {"type": "tool", "name": "search_flights", "inputs": {"from": "CAI", "to": "BER"},
                                             "outputs": {"flights": []}},
  {"type": "tool", "name": "book",           "inputs": {"flight": "TK33"}, "error": "not direct"}
]

Other keys read: name, step_id, timestamp (ISO 8601; an unparseable one is rejected with exit 2). Steps are numbered by position, so step 74 is the 75th entry in the file. examples/sample_trace.json is a working example. For any other format, convert it to this shape or add a TraceSource adapter (see CONTRIBUTING.md).

A Claude Code session (.jsonl) - the transcripts under ~/.claude/projects/, in a folder named after your project path with /, _ and . turned into -:

alibi diagnose ~/.claude/projects/-Users-me-LLM-projects-app/<session-id>.jsonl

For the most recent session of the project you are standing in, without looking the name up:

alibi diagnose "$(ls -t ~/.claude/projects/"${PWD//[\/_.]/-}"/*.jsonl | head -1)"

A real working session is often past the 250K ceiling, and is then refused with its cost rather than analysed. If the estimate is fine, say so: --max-trace-tokens 400000. Parsing is best effort, since the format is internal to Claude Code: assistant text, tool calls and tool results become steps, and thinking blocks are skipped.

A LangSmith trace - the trace id and the project, with LANGSMITH_API_KEY set:

alibi diagnose <trace-id> --source langsmith --project "my-project"

There is no folder mode: the method is a filter over one time series. To sweep a directory, loop:

for f in traces/*.json; do alibi diagnose "$f" --json > "${f%.json}.diagnosis.json"; done

Reading the result

In the output above:

  • alarm at chapter N - where CUSUM first saw the agent's health break down. The cause is usually at or before it, which is why the look-back starts there. None means no alarm fired and the run's ending was used as the anchor instead.

  • score - fused evidence for that step, not a probability. Only the ordering is meaningful.

  • steps are 0-based positions in the trace you passed in.

--json prints the same result as an object, for piping into something else.

Exit codes: 0 on success, a gated trace included. 2 for anything refused before a judge was built, which always means nothing was sent and nothing was spent. 3 if the run failed once the judge had started reading, where the message says how many calls were billed.

The two gates, and what they cost

Nothing is sent anywhere, and nothing is spent, unless the trace falls between them:

Trace size

What happens

Override

Under 50,000 tokens

Not analysed: "a single direct read is enough"

--min-trace-tokens, ALIBI_MIN_TRACE_TOKENS

50,000 to 250,000 tokens

Analysed; about $0.01 per 100K tokens

Over 250,000 tokens

Refused with an estimated cost, so a huge transcript cannot spend unannounced

--max-trace-tokens, ALIBI_MAX_TRACE_TOKENS

The MCP tool

The server (alibi-mcp, stdio) exposes exactly one tool for any MCP client, not just Claude Code:

diagnose_trace(trace: str, source: str = "auto", project: str | None = None) -> Diagnosis

trace is the same path or id the CLI takes. To wire it up by hand:

{ "mcpServers": { "alibi": {
  "command": "uvx",
  "args": ["--from", "agent-alibi>=0.1,<0.2", "alibi-mcp"],
  "env": { "TYPESAFE_API_KEY": "...", "ALIBI_JUDGE_BACKEND": "typesafe",
           "ALIBI_ALLOW_PAID_MODELS": "1" } } } }

Privacy

Traces above the length gate are sent to TypeSafe's API; below it, nothing leaves your machine. There is no telemetry. Claude Code session parsing is best effort: the transcript format is internal to Claude Code and may change. Session transcripts often contain source code and secrets, so read SECURITY.md before diagnosing one.

Contributing

See CONTRIBUTING.md. Plumbing, adapters, bug fixes and docs are welcome as ordinary pull requests; changes to the estimation method need a pre-registered measurement, for the reason the research log makes obvious.

Background

Alibi started as an experiment by an ADAS engineer: can the tracking filters used in cars work on AI-agent traces? The research log with every pre-registered test, including the ones that failed, is in docs/research-log.md.

License

MIT

Available Tools

1 tool
diagnose_traceA

Find where a long AI-agent run went wrong. Reads the trace chapter by chapter (never all at once), raises a CUSUM alarm, looks back, and returns the 3 steps to read first. trace: path to a .json trace or a Claude Code session .jsonl, or a LangSmith trace id. A trace under the length gate (default 50K tokens) or over the cost ceiling (default 250K tokens) is returned with gated true and a message saying which: it is not analysed and nothing is spent.

ParametersJSON Schema
NameRequiredDescriptionDefault
traceYes
sourceNoauto
projectNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
gatedYesTrue when the trace was not analysed and nothing was spent: it is empty, under the length gate, or over the cost ceiling. `message` says which.
anchorNo
messageYes
n_stepsYes
cost_usdNoBilled for this trace in USD; null if the judge logs none.
suspectsNo
n_chaptersNo
judge_callsNoJudge calls made for this trace; null if the judge logs none.
trace_tokensYes
alarm_chapterNo
judge_secondsNo

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure, and it delivers. It reveals that the tool reads incrementally ('chapter by chapter, never all at once'), uses a CUSUM alarm, performs a lookback, returns exactly 3 steps, and gates traces by token length/cost with a 'gated' flag and no spend. This is unusually rich behavioral context for a description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: purpose, algorithm behavior, parameter explanation, and gating rules are all covered without redundancy. It is slightly long, but the complexity of the tool justifies the length. Front-loading the purpose sentence is ideal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Considering the tool's complexity and the absence of annotations, the description is remarkably complete: it explains the algorithm, the gating thresholds, cost behavior, and expected output concept. An output schema exists, so return-value details need not be spelled out. Missing only minor clarifications around optional parameter semantics and normal (non-gated) cost behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It fully explains the required `trace` parameter, including accepted formats (.json, Claude Code .jsonl, LangSmith trace id). However, it provides no semantic guidance for `source` or `project`; the enum on `source` is somewhat self-explanatory given the trace formats, but `project` remains opaque. Partial compensation warrants a middle score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with a specific, outcome-oriented verb and resource: 'Find where a long AI-agent run went wrong.' It goes beyond a mere label by explaining the analysis approach (chapter-by-chapter reading, CUSUM alarm, lookback) and the concrete deliverable (the 3 steps to read first), leaving no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly frames when the tool is appropriate: for diagnosing long AI-agent runs that went wrong. It also tells the agent about gating behavior for traces that are too short or too costly, which effectively describes conditions under which the tool will not perform analysis. No sibling tools exist to contrast against, so exclusion guidance is unnecessary; the context is clear enough for a 4.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.1.4
    • Changeddiagnose_trace9 fields changed
      • addedInput schema / properties / source / enum
        Added value: +[
        +  "auto",
        +  "json",
        +  "claude-code",
        +  "langsmith"
        +]
      • addedOutput schema / properties / cost_usd / anyOf
        Added value: +[
        +  {
        +    "type": "number"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • addedOutput schema / properties / cost_usd / description
        Added value: +"Billed for this trace in USD; null if the judge logs none."
      • removedOutput schema / properties / cost_usd / type
        Removed value: -"number"
      • addedOutput schema / properties / judge_calls / anyOf
        Added value: +[
        +  {
        +    "type": "integer"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • addedOutput schema / properties / judge_calls / description
        Added value: +"Judge calls made for this trace; null if the judge logs none."
      • removedOutput schema / properties / judge_calls / type
        Removed value: -"integer"
      • addedOutput schema / properties / judge_seconds / anyOf
        Added value: +[
        +  {
        +    "type": "number"
        +  },
        +  {
        +    "type": "null"
        +  }
        +]
      • removedOutput schema / properties / judge_seconds / type
        Removed value: -"number"
  2. 1 tool updatev0.1.3
    • First observeddiagnose_trace

TDQS

A4.3/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is zero possibility of confusion or misselection. The tool's purpose is clearly distinct by virtue of being the sole tool in the server.

Naming Consistency5/5

The single tool name 'diagnose_trace' follows a clear verb_noun pattern and accurately describes its action and target. There is no inconsistent naming to flag.

Tool Count3/5

A single tool feels thin, but the server's apparent scope is narrowly focused on trace diagnosis. The tool consolidates several internal steps (reading, alarming, looking back, returning recommendations), making one tool borderline acceptable.

Completeness4/5

For its stated purpose, the tool covers the full diagnostic workflow in one call. Minor gaps exist (e.g., no configurable gates for the length/cost thresholds, no follow-up actions), but agents can achieve the core outcome without dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers