agent-eval
Analyze coding-agent session transcripts to measure behavior, efficiency, and cost.
list_sessions: Find session transcripts under a root directory, with summaries (id, working directory, tool-call count, human turn count, cost).
analyze_session: Analyze one transcript end to end: tool usage/failure rates, loops, test outcomes, permission posture, cost, and human-readable findings.
find_loops: Detect repeated stateful commands or edits (same action repeated at least a threshold number of times), excluding read-only actions.
cost_report: Aggregate total spend, lines changed, and spend per 100 lines across all sessions under a root.
Use inside an agent session: MCP tools let an agent check whether it has been going in circles (e.g., repeated failing commands) and stop or adjust.
CI-friendly analysis: CLI supports
--json --strictfor automated checks.Privacy-conscious: prompt text and file contents are not included in results; shell commands are echoed only in loop findings and truncated to 400 characters.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@agent-evalAnalyze my most recent session for loops and cost."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
agent-eval-mcp
Measure what a coding agent actually did. Tool efficiency, loops, test outcomes, permission posture and cost, read out of session transcripts. Ships as a CLI and as an MCP server.
uv venv --python 3.11
uv pip install -e .
agent-eval sessions # what sessions exist
agent-eval analyze path/to/session.jsonl # what happened in one
agent-eval analyze path/to/session.jsonl --json --strict # for CIOne of three. agentic-harness-jvm gates the code an agent writes. agent-egress-gate is a deny-by-default proxy constraining where a headless agent can reach, with a tamper-evident audit log. This one gates the agent's own behaviour.
Why
You can read a diff and see what an agent produced. You cannot easily see that it spent forty minutes running the same failing build, that it ended the session on a red test, or that it burned nine dollars to change eleven lines. That information is in the transcript, and nobody reads transcripts.
So: read them mechanically, and report only what a person should look at.
The findings list is the point. Everything else is supporting detail.
Session 4f2a...
Findings
- Possible loop: Bash ran the same action 7 times (every attempt failed) - Bash:mvn -q verify
- Session ended with a failing test run
Tools
118 calls, 31 failed (26%)
Bash: 74, 28 failed
Edit: 31, 3 failed
Read: 13
Tests
9 runs, 6 failed, ended red
Cost
$8.44 over 51m
312 lines changed, $2.71 per 100 linesRelated MCP server: agentops-mcp
What it measures
Signal | What it answers |
Tool usage | Where did the effort go, and which tools kept failing? |
Loops | Did it issue the same command or edit over and over? Read-only tools are excluded, because re-reading a file is navigation, not thrash, and so are shell commands that change nothing ( |
Test events | Did the suite run, how often did it fail, and did the session end green? |
Gates | What permission posture was it running under, and did a hook ever refuse a call? |
Cost | Total spend, and spend per 100 lines changed. |
Findings | The short list: loops, a red ending, a tool failure rate above 25% over a meaningful sample, hook blocks. |
Loop detection matches on a normalized call signature rather than on adjacency. An agent that alternates between the same failing build and the same failing edit is looping just as surely as one that repeats a single command, and the interleaving is exactly what makes it hard to see by eye.
Use it from an agent, not just after the fact
The useful moment for this data is inside a session, when the question is "have I been going in circles for the last twenty minutes." A CLI answers that afterwards, to a human. An MCP tool answers it during, to the agent, which can then stop.
// .mcp.json
{
"mcpServers": {
"agent-eval": {
"command": "uv",
"args": ["run", "--directory", "/path/to/agent-eval-mcp", "python", "-m", "agent_eval.mcp_server"]
}
}
}Four tools: list_sessions, analyze_session, find_loops, cost_report. Schemas and
descriptions are derived from the function signatures and docstrings, so what a model reads and what
it gets cannot drift apart.
Results come back as MCP structured content. A tool returning an object gives you that object; a
tool returning a list is wrapped as {"result": [...]}. A missing transcript or a root that is not
a directory raises ToolError naming the path, so a model that guessed wrong can correct itself
from the error alone rather than concluding that the corpus is empty. ~ is expanded.
list_sessions and cost_report walk top-level sessions, newest first. A session's delegated
subagent transcripts live under its subagents/ directory and are that session's work, not
sessions of their own; pass include_subagents=true to see them as well.
Privacy is a design constraint, not a setting
Transcripts contain prompts, file contents, paths and occasionally credentials that were pasted into a terminal. A tool whose job is to analyse them should not become a second, less careful copy of them.
So:
Prompt text is counted and discarded.
human_turnsis a number. The text never enters aSessionobject, which is asserted by a test.File contents never leave the parser. An edit or write is identified by its path plus a fingerprint of the whole input, so identical edits still match as a loop while the edited text, which may be a whole file, is never carried into a finding, a report, or an MCP result. A test writes a fake credential three times and asserts it appears nowhere in any output.
A repeated shell command is echoed in its loop finding, because the command is what a reviewer needs to see. That is the one place raw transcript text reaches the output, and it is capped at 400 characters. A command can carry content inline - a heredoc body, an
echo secret >redirect - so an uncapped echo would put a whole file in a finding. The cap bounds that. It does not make a short secret written inline invisible, which is why the advice below is to grep the JSON.Tool results are truncated to 400 characters, and those characters are read:
suite_eventsscans a test run's own output for failure markers, because the transport's error flag cannot be trusted on its own. A shell pipeline exits with the status of its last command, sopytest ... 2>&1 | head -30exits 0 however the tests went - and that shape is 85% of the test-run commands in the corpus below. The retention now pays for itself; it used to be held against a parsing feature that did not exist.Nothing is written to a transcript directory. The parser opens files read-only, and there is no code path in the package that writes there.
No real transcript is committed. Every test fixture is hand-written synthetic JSONL, and the suite never reads a real transcript directory.
The output is counts and classifications. Run --json and grep it for anything you consider
sensitive; the tests do exactly that.
Run against a real corpus
The analyser was run over a personal archive of 3,000+ transcript files (several hundred top-level sessions plus the subagent transcripts they spawned) containing 200,000+ tool calls, in 27 seconds on a laptop.
The floors are deliberate. Transcripts rotate away and new ones arrive, so any exact figure here is stale within days; the table below is one dated measurement, not a standing claim. A re-run on 2026-09-17 gave 3,469 transcripts and 217,197 tool calls: fewer files than the 2026-09-08 set, more calls, because the sessions that aged out were lighter than the ones that replaced them.
Measured 2026-09-08:
Top-level sessions | Including subagents | |
Transcripts parsed | 611 (537 with at least one tool call) | 3,542 (3,463) |
Tool calls | 85,883, of which 3.2% returned an error | 208,035, 2.7% |
Sessions with a detected loop | 108, or 17.7% | 266, or 7.5% |
Sessions ending green / red on a recognized test run | 137 / 6 | 646 / 17 |
Findings raised | 221 loops, 6 red endings | 486 loops, 17 red endings, 5 high-failure-rate sessions |
Two of this tool's own defects were found by running it, not by reading it.
These figures were re-measured on 2026-09-08 and replace an earlier set. Two things moved them, and it is worth separating them, because only one is a change to the tool.
The corpus itself shrank: transcripts are rotated away over time, so the archive is a living thing and 629 top-level sessions became 611. Nothing can be concluded from a comparison across that.
The red column is the real change. It was 0 / 7 and is now 6 / 17, because suite_events
stopped taking the exit status at face value. To attribute that honestly rather than confound it
with the corpus drift, the classifier change was measured against a frozen set of the 135,618
distinct shell commands in the archive, run through both the old and the new implementation:
3,303 commands classified as test runs before, 3,300 after - a 0.09% change, so the "recognized test run" denominator is essentially the same population.
The corrections were
dotnet test --help(twice) and apytest --cocollect-only, all three genuinely not test runs; and one command the old code missed entirely, annpm testhidden behind a glued);token.Two commands are classified wrongly by both implementations, in opposite directions, for the same reason: the parser has no notion of a heredoc body or a quoted string spanning lines, so prose in a commit message can look like a command and an embedded script can hide one. Two in 135,618, and recorded rather than papered over.
So the green/red shift is not a re-baselining artefact. It is the same commands, judged correctly: those sessions really did end on a failing suite, and the tool used to say they were green.
That run is also how the loop detector's own bug was found. On the first pass, a seven-hour session reported five separate edits to one file as a loop, because the signature for an editing tool keyed on the file path alone. Editing one file repeatedly is ordinary iterative work; only the identical edit repeated is thrash. The signature now keys on the whole tool input, three regression tests pin the behaviour, and the false positives disappeared while the genuine ones (the same shell command re-run four times) stayed.
The second pass, a code review after publication, found that the fix had swung too far: keying on
the whole input put the whole input into the loop's signature, so a repeated Write reported its
file contents verbatim in the findings, contradicting the privacy claim below it. The signature is
now a path plus a fingerprint, and the test that would have caught it exists. The same review found
that the first published figure of 3,904 sessions counted every subagent transcript as a session of
its own, inflating the count about sixfold; the table above separates them.
Worth stating plainly: a tool that measures agent behaviour is only trustworthy if you have pointed it at messy real data and fixed what it got wrong. One synthetic fixture would never have surfaced either.
Reading the numbers honestly
A command that changes nothing is not a loop. Repeating a read-only tool was already
navigation rather than thrash; from 2026-09-17 the same holds one level down, for a shell command
whose every segment is a no-op or a question: true, sleep, echo, date, cd, ls, cat,
grep, git status, git log. One stateful segment anywhere makes the whole line count again, so
cd repo && git status && dotnet build is still reported, and so is any redirect, because
echo secret > .env writes a file while echo secret does not.
This came out of running the tool on the archive rather than reading it. Of 468 loop findings,
three had every attempt fail; the rest were overwhelmingly orchestrators polling with sleep
and echo while they waited on background work. A count that large made the metric unreadable and
buried the three. Measured over one snapshot of 3,470 transcripts, both implementations in the same
pass:
Sessions with a loop | Loop findings | Of those, every attempt failed | |
Before | 233 | 468 | 3 |
After | 167 | 298 | 3 |
The findings that matter are all still there. What went is noise.
What still gets through: a shell until/while ... do ... done polling construct is reported,
because its head is until or while rather than a program this can classify, and reading control
flow is a different job from reading a command. That is the next obvious improvement, alongside the
time window below.
Loop detection is a heuristic. Three identical npm run build calls may be fine. The
threshold is a parameter (--loop-threshold) because the right value depends on your work. What the
tool is actually good at is surfacing failing repetition, which is why all_failed is reported
separately, and why the quiet-command rule above matters: it is the difference between three
findings worth reading and 468 nobody will.
Loops are counted across the whole session, not within a time window. The same command run four times over seven hours reads the same as four times in five minutes, and the first is usually fine. Adding a window is the obvious next improvement.
Test detection parses the shell command. Each simple command is tokenised, launcher prefixes
(uv run, python -m, poetry run, npx, timeout, environment assignments, cd x &&) are
peeled off, and the program is looked up in a table: pytest, tox, nox, jest, vitest, mocha, npm/yarn/
pnpm/bun/deno test, mvn test/verify/install, gradle test/check/build, go, cargo, dotnet, make, mix,
rake and swift test, rspec, phpunit, ctest. A commit message that mentions pytest does not count,
and neither does ls tests/. It will still miss a suite invoked through a custom script.
"Ended green" means the last recognized test run passed. It is not a statement about whether the
work was correct. A session with no test runs, or whose last run never returned a result because the
session was cut off, reports null rather than pretending.
Human turns are the prompts a person typed. The runtime also writes user-role records nobody typed (command output, task notifications, compaction summaries, meta injections); those are recognised by their flags or their opening tag and excluded, which matters on sessions that delegate to many subagents.
Cost comes from the runtime's own rollup, not from re-pricing tokens. Dollars per 100 lines is a crude ratio: a session that spends its budget on reading and reasoning will look expensive per line and may have been the right call.
None of this measures whether the change was any good. It measures process. A session can be clean on every metric here and still produce the wrong feature.
Development
uv venv --python 3.11
uv pip install -e ".[dev]"
pytest # 163 tests, coverage floor at 95%
ruff check . && ruff format --check . # lint and format gatesBuilt test-first. The coverage floor is in pyproject.toml and sits at 95%; measured coverage is
above it. The floor exists to notice a test file going dark, so it tracks the real figure closely.
Tests are organized by module: transcript parsing, the analyses, rendering, the CLI, and the MCP surface. The MCP tests exercise the server's own dispatch rather than a stdio transport, because transport framing is the SDK's job and is tested there.
License
MIT. See LICENSE.
Available Tools
4 toolsanalyze_sessionA
Analyse one session transcript end to end.
Returns tool usage and failure rates, detected loops, test-run outcomes, permission posture, cost, and a short list of findings worth a human's attention. Prompt text and file contents are never included.
Args: path: Path to a .jsonl session transcript. loop_threshold: How many repeats of one action count as a loop.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| loop_threshold | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the transparency burden. It discloses that prompt text and file contents are never included and indicates a read-only analysis function, giving reasonable insight into behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and organized with a short summary followed by parameter explanations. No redundant or filler content is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given an output schema exists, the description adequately summarizes the returned analysis items without needing exhaustive detail. It provides enough context for an agent to understand the tool's purpose and key inputs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Both parameters are described in the description text despite having no schema annotations: path is a .jsonl transcript path and loop_threshold defines the repeat count for loop detection. The explanations are brief but sufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool analyzes a single session transcript and lists the types of results returned, distinguishing it from simple listing or focused sub-analyses. It does not explicitly name sibling tools, but the scope is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It indicates this is for end-to-end analysis and enumerates outputs, but it does not explicitly state when to choose this over the sibling tools find_loops or cost_report, which overlap in detected loops and cost reporting.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cost_reportA
Aggregate cost across every session under a root.
Returns total spend, total lines changed, and spend per 100 lines changed.
Args: root: Directory to search. Defaults to the standard transcript location.
| Name | Required | Description | Default |
|---|---|---|---|
| root | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of disclosing behavior. It states the tool returns total spend and lines changed, implying a read-only aggregation, but does not explicitly confirm it is non-destructive or describe any side effects. It also does not mention any limitations or prerequisites beyond the root parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the core purpose, followed by a clear Args section. Every sentence adds value and there is no redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the purpose, parameter semantics, and the output metrics. Since an output schema is provided, the description does not need to detail the return structure. However, it lacks explicit usage guidance or prerequisites, which would make it more complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description explains the single parameter 'root' as a directory to search, and specifies its default to the standard transcript location. This adds meaning beyond the schema, which only defines the type and default null. The description fully compensates for the 0% schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool aggregates cost across sessions under a root and lists the specific metrics returned (total spend, total lines changed, spend per 100 lines changed). This distinguishes it from sibling tools like list_sessions and analyze_session, which serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly state when to use this tool versus alternatives like analyze_session or list_sessions. The purpose is clear but there is no guidance on exclusions or conditions that would make a different tool more appropriate. This is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_loopsA
Return only the repeated stateful actions in a session.
A loop is the same command or edit issued at least threshold times. Reads and
other read-only tools are excluded, because re-reading a file is navigation
rather than thrash.
Args: path: Path to a .jsonl session transcript. threshold: Repeats of one action that count as a loop.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| threshold | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the core behavior (returning repeated stateful actions), excludes read-only operations, and clarifies the threshold parameter. However, it does not mention side effects, error handling, or output format, which are typical behavioral details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, well-organized, and free of fluff. It opens with a one-sentence summary, provides a brief definition, and lists parameters in a clear Args format. No unnecessary information is included.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the essential aspects for using the tool: what it returns, what it excludes, and the meaning of both parameters. It does not specify the exact return format or edge cases, but for a simple analysis tool the information is sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has no parameter descriptions, but the tool description fully compensates by explaining both parameters: 'path' is described as a path to a .jsonl transcript, and 'threshold' is defined as the repeat count that qualifies as a loop. Both are clear and align with the earlier definition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function ('Return only the repeated stateful actions in a session') with a specific verb and resource. It also defines a 'loop' precisely, but does not explicitly distinguish it from sibling tools like analyze_session or cost_report.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear definition of what constitutes a loop, but it lacks explicit guidance on when to use this tool versus alternatives such as list_sessions or analyze_session. No direct comparison or 'use this when' statement is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_sessionsA
List session transcripts under a root directory.
Returns one summary per session: id, working directory, tool-call count, human turn count and cost. Use it to find the session worth analysing.
Args: root: Directory to search. Defaults to the standard transcript location. limit: Maximum number of sessions to return.
| Name | Required | Description | Default |
|---|---|---|---|
| root | No | ||
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of disclosing behavior. The description implies a read-only listing operation, but it does not explicitly state that it has no side effects, does not modify anything, or what happens in error cases. It is clear in intent but lacks explicit behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: the purpose is stated in the first sentence, the return format is clarified in the second, and the usage hint is appended without redundancy. No unnecessary words or vague filler are present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple listing tool, the description covers the essential context: what it does, what it returns, and why an agent would use it. Even though an output schema is not shown, the description enumerates the return fields, making the tool self-contained and sufficient for an agent to decide when and how to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Both parameters are explained in the description: 'root' is identified as the directory to search with a default to the standard location, and 'limit' is described as the maximum number of sessions. This fully complements the schema, which only provides types and defaults, giving the agent all necessary context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('List') and the resource ('session transcripts under a root directory'), and it explains the return value (summaries with specific fields). It also conveys its primary use case ('find the session worth analysing'), which distinguishes it from the sibling tools that analyze or report on sessions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description includes an explicit usage hint ('Use it to find the session worth analysing'), which tells the agent when to invoke this tool. However, it does not explicitly contrast it with the sibling tools (analyze_session, find_loops, cost_report) or mention when not to use it, so it falls slightly short of full guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
analyze_session - First observed
cost_report - First observed
find_loops - First observed
list_sessions
TDQS
Scored across 4 tools
The tools are mostly distinct: list_sessions, analyze_session, find_loops, and cost_report each have clear responsibilities. find_loops overlaps somewhat with analyze_session since analyze_session already reports detected loops, but the focused purpose keeps them distinguishable.
Three tools follow a verb_noun pattern (list_sessions, analyze_session, find_loops), but cost_report breaks the pattern by leading with a noun rather than an action verb. This is a minor inconsistency in an otherwise readable set.
Four tools is well-scoped for a session-analysis server. Each tool earns its place by covering listing, deep analysis, loop detection, and cost aggregation without unnecessary bloat.
The tool surface covers the core workflow: enumerate available sessions, analyze an individual session, inspect loops specifically, and aggregate costs. No obvious necessary operation is missing for the stated domain.
Maintenance
Related MCP Connectors
Analytics for MCP servers. Query your tool calls, first-call success, retries and schema cost.
- AgentCatOAuthcom.agentcat
Analytics and debugging for your MCP server — explore usage and sessions, then root-cause errors.
Analytics your AI agent can actually use. Track, experiment, and optimize via MCP.
Analytics for MCP servers. Find out which of your tools agents get wrong. MCPulse shows you which tools AI agents retry, which come back empty, and which they never call at all. Two lines inside your own server. It never sees your arguments or your results. getmcpulse.com
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceMCP server that analyzes AI agent session logs to find token waste and optimization opportunities.4 npmMIT
- AlicenseAqualityCmaintenanceExposes analytics from Claude Code transcripts as MCP tools, enabling cost, audit, safety, and efficiency queries through natural language.4MIT
- AlicenseAqualityBmaintenanceA read-only MCP server that exposes local coding-agent session logs as three tools for introspection of recent work, debugging tool failures, and tracking token usage and estimated cost without parsing log files.3MIT
- AlicenseAqualityAmaintenanceEnables search, analytics, and visualization of Claude Code sessions with MCP tools for session management, recovery, and insights.1029 PyPI1MIT