TokenTrust
TokenTrust
Install • What it measures • Commands • Agent-native / MCP • FAQ
Vendor-neutral CLI that independently verifies the token and cost savings AI-coding-agent context-reduction proxies actually deliver, by running the proxy for real against a labeled task corpus instead of trusting the maintainer's own number.
Install
TokenTrust ships as two complementary, equally first-class distributions: an npm package for
Node.js toolchains and a PyPI package for Python toolchains. Both install a tokentrust command
with the identical CLI surface, the same TT01-TT05 verification categories, and the same bundled
task corpus, so pick whichever matches your existing stack.
npm (Node.js):
npx tokentrust-cli verify --proxy rtkNo clone, no local build. npx fetches the published package and runs it directly. To install
it as a dependency instead: npm install -g tokentrust-cli.
pip (Python):
pip install tokentrust-cli
tokentrust verify --proxy rtkSee python/README.md for the Python package's full documentation,
including a note on the one real behavioral difference between the two: the npm package's
js-tiktoken dependency bundles its tokenizer data for fully offline use, while the Python
package's tiktoken dependency fetches and caches that same public data on first use.
Real output from that exact command, run against this repo's own bundled task corpus:
$ npx tokentrust-cli verify --proxy rtk
TokenTrust v0.1 -- Token/Context-Reduction Claims Verification
Proxy: rtk 0.43.0 | Repo: TokenTrust | Task corpus: 23 labeled tasks
[MEASURED] TT01 Compression Ratio
Claimed (rtk README): up to 70% context reduction
Measured (this repo, this corpus): 60.6% average reduction across 23 tasks
Range: 0.0% ("verify-go-build-filter") to 95.4% ("verify-git-log-filter")
[MEASURED] TT02 Cost-Savings Delta
Baseline (uncompressed): $0.02 across 23 tasks @ claude-5-sonnet pricing
Compressed (rtk-proxied): $0.00 across 23 tasks
Actual savings: 77.0% ($0.01) -- vs. claimed 70% ceiling
[FAIL] TT03 Never-Worse Output Guard
2/23 tasks regressed in task-completion diff vs. uncompressed baseline
[PASS] TT05 Version-Drift Regression Check
No prior verified baseline for rtk on this repo -- this run establishes the first baseline.
Summary: 77.0% measured cost savings (claimed: up to 70%) -- see full reportThat's a real run's output, not a hand-typed example -- npx tokentrust-cli invokes the exact
same dist/cli.js entry point (the tokentrust command name is unchanged), so it reproduces on
your machine with no clone required.
Related MCP server: agent-eval-mcp
What it measures
TT01: Compression Ratio. Actual token reduction, measured with a local tokenizer (
js-tiktoken), against every task in the corpus.TT02: Cost-Savings Delta. Dollar-cost savings computed from TT01's measured token delta at published model pricing. Optional
--livemode verifies the estimate against a real, provider-billed sample (opt-in, your own API key, gated behind--confirm-cost, capped at 5 tasks by default).TT03: Never-Worse Output Guard. Checks whether a proxy's compressed output dropped content a task marks as required to survive compression.
TT04: Cross-Tool Comparative Benchmark. Pass
--proxymore than once and TokenTrust runs the identical task corpus through every named proxy side by side.TT05: Version-Drift Regression Detection. Compares a run's measured savings against the last-verified baseline for the same proxy/repo pair, so a silent regression across a version bump (like
rtk#582) gets caught automatically.
Commands
tokentrust verify --proxy <name> [options]Flag | Description |
| Proxy to verify. Repeatable, pass it more than once to run TT04's cross-tool comparison. Supported: |
| Repo to measure against. Defaults to the current directory. |
| Task corpus YAML file. Defaults to the bundled 23-task corpus. |
| Sample real, provider-billed tokens for the first proxy instead of estimating from pricing tables. Requires |
| Confirms the estimated spend |
| Max tasks sampled in |
| Report output format. Defaults to |
| Show the help message and exit. |
This table (and the tokentrust mcp reference below) is verified against the actual --help
output of the published tokentrust-cli package, not an old copy. One known gap between
that live --help text and the table above is called out directly in the FAQ, instead of
silently repeating it.
--format json gives every category's claimed-vs-measured numbers as structured data, so a
script or agent can pull the comparison straight out with jq instead of parsing terminal text:
Exit code is 0 when the run completes with no gated failure, non-zero otherwise. The bundled
GitHub Action's --fail-on-regression maps that straight to a failed CI step, so a version-drift
regression breaks the build instead of shipping silently.
Add it to CI with the bundled GitHub Action (action/action.yml) so verification reruns
automatically whenever a proxy's version bumps:
- uses: RudrenduPaul/TokenTrust-CLI@main
with:
proxy: rtk
fail-on-regression: 'true'
cli-version: '<pin to the exact tokentrust-cli version you have verified against>'Pin cli-version explicitly rather than relying on the Action's default. The Action deliberately
does not support latest for this input: pinning is a supply-chain safeguard, so a compromised
npm publish can't reach every workflow using this Action on its very next run without a
deliberate version bump on your part. The Action's own default for this input is an early
pre-TT04/TT05/MCP release -- omitting cli-version runs that old version, not the one this
README documents.
Agent-native / MCP
TokenTrust ships in the same dual CLI + MCP-server mode as Semgrep, Trivy, Snyk, and
SonarQube: one binary, one underlying verification engine, and a second, thin front door for
agents that speak MCP (Model Context Protocol) instead of a
shell. tokentrust mcp starts an MCP server over stdio, exposing a single tool,
verify_proxy_savings, that calls straight into the same runVerify() engine tokentrust verify uses -- no verification logic is duplicated, and the tool returns the exact structured
JSON report --format json already produces.
npx tokentrust-cli mcpRegister it with an MCP client
Point the client's server config at this binary with the mcp argument. For Claude Code,
Claude Desktop, or any other client that reads an mcpServers block:
{
"mcpServers": {
"tokentrust": {
"command": "npx",
"args": ["tokentrust-cli", "mcp"]
}
}
}The tool
Field | Description |
| Tool name. Mirrors |
| A single proxy name ( |
| Same as |
| Same as |
| Same |
| Same as |
This is the tool's real, unedited tools/list schema, captured from a running tokentrust mcp
server (inputSchema trimmed of per-field descriptions here for length; the live server returns
them in full):
{
"name": "verify_proxy_savings",
"title": "Verify proxy token/cost savings",
"inputSchema": {
"type": "object",
"properties": {
"proxy": { "anyOf": [{ "type": "string", "enum": ["rtk", "headroom"] }, { "type": "array", "items": { "type": "string", "enum": ["rtk", "headroom"] }, "minItems": 1 }] },
"repo": { "type": "string" },
"tasks": { "type": "string" },
"live": { "type": "boolean" },
"confirmCost": { "type": "boolean" },
"liveMaxTasks": { "type": "integer", "exclusiveMinimum": 0 }
},
"required": ["proxy"]
}
}A real tools/call against this repo, {"name": "verify_proxy_savings", "arguments": {"proxy": "rtk"}}, returns the same shape as the CLI's --format json output (trimmed here; the live
call returns the full records array with TT01/TT02/TT05 entries):
{
"content": [
{
"type": "text",
"text": "{\n \"run_id\": \"tt_2026-07-18_f88644\",\n \"repo\": \"...\",\n \"task_corpus_size\": 23,\n \"proxies\": [\"rtk\"],\n \"records\": [ /* TT01, TT02, TT05 -- same shape as `verify --format json` */ ],\n \"tt03\": { \"rtk\": { \"pass\": false, \"regressed_count\": 2, \"task_corpus_size\": 23 } },\n \"tt05\": { \"rtk\": { \"pass\": true, \"message\": \"No regression vs. last-verified rtk 0.43.0 baseline (stored 2026-07-18).\", \"prior_run_id\": \"tt_2026-07-18_608b74\", \"degraded\": false } }\n}"
}
],
"isError": false
}isError is true (with no report) whenever the underlying runVerify() call itself would
have exited non-zero on the CLI -- a missing proxy binary, an invalid task corpus, or the
--live safety gate refusing an under-confirmed call. Progress output and the trace log
tokentrust verify normally prints to stdout are rerouted to stderr in MCP mode, since stdout
is the live JSON-RPC wire once a stdio transport is connected.
Proxy support
Proxy | Status |
| Fully supported: real subprocess-based verification ( |
| Recognized ( |
How it compares
What it does | Ongoing / self-serve | Verifies a specific claim | |
TokenTrust | Runs a named proxy against a labeled task corpus, measures real compression, cost, and output-quality regression, prints claimed vs. measured | Yes, runs in your own CI, on your own repo, every time a proxy version bumps | Yes, that's the whole point |
Independent pilot benchmark of rtk, headroom, and lean-ctx on real agentic SDLC tasks, with raw transcripts and a pre-registered protocol | No, a single-repo pilot report, replication in progress | Yes, and rigorously: credit where it's due | |
Langfuse, Vantage, Finout, Amnic, Revenium | LLM/AI cost observability and FinOps. Track your actual API spend across models and providers, allocate it across teams | Yes, hosted or self-hosted, ongoing | No, these track what you spent; they don't check whether a specific proxy's specific savings claim holds up |
tokbench is the closest prior art and deserves real credit. It's rigorous and disclosed, but its pilot scope is narrower than a first read suggests: one repository, one task, N=1 per arm, replication runs in progress. Its own numbers on that pilot are worth reading directly. Provider-billed input tokens against a 2.28M-token native baseline came in at 2.89M for rtk (+27%) and 3.24M for headroom (+43%, despite headroom genuinely compressing 342K tokens on the wire), because the agent's turn count grew even as the per-turn payload shrank. That's a real, independently useful data point, and it's exactly the kind of gap between "compressed" and "cheaper" TokenTrust exists to keep catching, continuously, in your own repo rather than a single published pilot.
Why this exists
Context-reduction proxies (rtk, headroom, and others) publish compression and
cost-savings numbers in their own READMEs. Those numbers come from the maintainer's own
benchmark, on the maintainer's own workload, with nobody outside the project checking the math.
That's not an accusation. It's just how every proxy in this space currently reports its own
numbers, and a maintainer benchmarking their own tool isn't running an adversarial test.
The gap shows up in the proxies' own issue trackers:
rtk#839, an open, 5-repo, 2,100-measurement empirical benchmark thread asking how rtk's actual savings compare to what it claims.rtk#1935, "rtk gain hallucinates massive token usage and savings" (open).rtk#582, "RTK Hook Increases Claude Code Costs by 18%," a cost regression a maintainer's own test suite didn't catch on its own. TT05 exists specifically to catch this class of regression before a user does.
TokenTrust doesn't compete with these proxies. It verifies them. It has no stake in whether a proxy's claimed number holds up, and every category run prints the claimed number right next to the measured one, so the comparison is never hidden or averaged away.
We also found and fixed a bug in our own measurement: one fixture's baseline had accidentally
been captured with git log --oneline instead of a true raw git log, which understated rtk's
real compression on that task by roughly 42 percentage points. Recapturing it honestly is why
verify-git-log-filter now measures 95.4%, the highest reduction in the corpus, and a real one.
Commit e42246c has the fix --
no measurement number ships without a fixture-run behind it. That same commit expanded the
corpus from 15 tasks to the current 23, which is why the actual runtime output above says "23
labeled tasks" even though one leftover help string still mentions the old count (see the FAQ).
What is TokenTrust, and why does it exist
TokenTrust is a command-line tool that measures whether an AI-coding-agent context-reduction proxy's advertised token and cost savings hold up against a real, labeled task corpus, run with a local tokenizer instead of a spreadsheet estimate. It exists because compression proxies currently self-report their own savings numbers, and there is no independent, repeatable, CI-native way to check one before adopting it. TokenTrust is not a proxy itself and does not compress anything. It verifies proxies that do.
Real-world validation
TokenTrust's own validation work has already fed back into a real, independently tracked GitHub
issue: rtk-ai/rtk#1313 (filed by @ChrisEdwards,
asking rtk for a lossless-only mode and an honest account of the silent failures truncation
causes in agent contexts) was originally verified as only partially addressed by rtk's existing
mechanism, because TokenTrust's own fixtures didn't yet carry the quality markers needed to prove
it either way. Extending three of TokenTrust's pipe --filter fixtures with real, verified
quality markers closed that gap in TokenTrust's own instrumentation, not in rtk, and let the tool
confirm, against the real rtk 0.43.0 binary and not a claim, that rtk's existing never-worse guard
mechanism already does what the issue asked for. The issue's verdict moved from partial to a
genuine, re-verified pass as a direct result. TokenTrust never touched rtk's own repository; it
got sharp enough to prove what was already true there.
Python package
pip install tokentrust-cli installs the same tokentrust CLI as a genuine Python port, not a
wrapper around the Node binary: real Python source under python/src/tokentrust/,
its own pytest suite, and the identical bundled 23-task corpus, copied verbatim into the wheel.
Both distributions run the same cl100k_base tokenizer encoding, verified to produce identical
token counts on real sample text, and both are maintained together going forward, including
tokentrust mcp: the Python port exposes the same verify_proxy_savings MCP tool, with a
byte-identical wire schema, as the npm package (see python/README.md's
"Agent-native / MCP" section). See
python/README.md for install instructions, python/docs/getting-started.md
for a walkthrough, and python/docs/concepts.md for the verification
methodology shared by both packages.
FAQ
What is TokenTrust, and how is it different from a context-reduction proxy like rtk or headroom?
TokenTrust is not a proxy itself and does not compress anything. It is a vendor-neutral
verification layer: it runs a proxy like rtk or headroom as a real subprocess against a
fixed, labeled 23-task corpus, measures the actual token and dollar savings with a local
tokenizer, and prints that measured number next to the number the proxy's own README claims. The
differentiator is independence: TokenTrust has no stake in whether a proxy's claimed number holds
up, so it never averages the gap away.
Which platforms does TokenTrust run on, and are the npm and PyPI packages the same tool?
Both are genuine, separately maintained ports of the same tool, not one wrapping the other.
npm install -g tokentrust-cli (or npx tokentrust-cli) installs the Node.js build; pip install tokentrust-cli installs a real Python port under
python/src/tokentrust/, with its own pytest suite. Both expose the
same tokentrust command, the same TT01-TT05 categories, the same bundled task corpus, and the
same cl100k_base tokenizer encoding.
Does TokenTrust work with AI agents directly, not just from a shell?
Yes. tokentrust mcp (or npx tokentrust-cli mcp) starts an MCP (Model Context Protocol) server
over stdio that exposes one tool, verify_proxy_savings, backed by the same runVerify() engine
the CLI uses. Any MCP-compatible client, including Claude Code and Claude Desktop, can call that
tool and get back the same structured JSON report --format json produces on the command line.
How does TokenTrust compare to tokbench, the other independent proxy benchmark? tokbench is real prior art and deserves credit: a rigorous, disclosed pilot with raw transcripts and a pre-registered protocol. Its current scope is narrower than a first read suggests, one repository, one task, N=1 per arm, with replication in progress. TokenTrust instead runs a 23-task corpus continuously, in your own CI, on your own repo, every time a proxy version bumps, rather than as a single published pilot report.
The tokentrust verify --help text on the npm package says the default task corpus has 15 tasks. Which is correct, 15 or 23?
23 is correct. The bundled corpus was expanded from 15 to 23 tasks in
commit e42246c, and every real
run (both npm and PyPI) reports "Task corpus: 23 labeled tasks" at the top of its output, matching
the actual fixtures/tasks.yml file. The npm package's --tasks flag help text is a leftover
string from before that expansion and hasn't been updated to say 23; the PyPI package's help text
already says 23 correctly. This affects only what the --help text displays, not what tasks the
CLI actually runs.
What if pip install tokentrust-cli fails on my Python version?
The PyPI package declares requires-python = ">=3.10" in its pyproject.toml, so pip will
refuse to install it on Python 3.9 or older. Upgrade to Python 3.10, 3.11, 3.12, or 3.13 (the
versions the package is tested against), or use the npm package instead, which only requires
Node.js 18 or newer.
Can I use TokenTrust commercially, and do I need to attribute it? Yes. TokenTrust is licensed Apache-2.0 (see LICENSE), which permits commercial use, modification, and distribution, including inside closed-source products, as long as you keep the license and copyright notice and note any changes you made to the source itself.
Does TokenTrust modify my code or my proxy's compressed output? No. It runs the proxy as a real subprocess against fixture tasks, captures the output, and measures it. Nothing in your repo or the proxy's configuration is changed.
Can TokenTrust verify a proxy's live, provider-billed cost instead of an estimate?
Yes, with --live --confirm-cost, capped at 5 tasks by default via --live-max-tasks. It uses
your own API key and never runs a real charge without printing the estimated spend first.
Is TokenTrust importable as a library, or is it CLI-only?
CLI-only today. Neither the npm package nor the PyPI package ships a documented, importable
public API: the npm package.json lists a main entry that points at a file the published
package doesn't actually contain, and the PyPI package's top-level module only exports a version
string. Use the tokentrust command or the verify_proxy_savings MCP tool; there is no
supported way to import/require this package's verification logic directly yet.
Contributing
See CONTRIBUTING.md for the project layout, how to add a verification category or fixture task, and the coverage bar every category change is held to, for both the npm and PyPI packages.
License
Apache-2.0. See LICENSE.
Available Tools
1 toolverify_proxy_savingsVerify proxy token/cost savingsA
Runs an independent, adversarial verification of an AI-coding-agent context-reduction proxy's (rtk, headroom) claimed token and cost savings, using the same TT01-TT05 engine as tokentrust verify on the command line: a real local tokenizer (tiktoken, cl100k_base) and the bundled 23-task labeled corpus, with the proxy invoked as a real subprocess rather than re-running the vendor's own benchmark script. Call this when an agent needs a trustworthy, third-party number for a proxy's actual compression ratio, cost delta, or output-safety guard -- e.g. before recommending a proxy, evaluating a version upgrade, or checking a CI regression -- not for general token counting or for proxies outside {rtk, headroom} (headroom is recognized but not yet runnable; the report notes this rather than failing). No API key or network access is required in the default (non-live) mode, which estimates cost from published pricing tables; it only needs the proxy binary and a task corpus to already be present on disk.
Side effects: read-only against repo (the proxy runs against the task corpus; the target repo itself is never modified) and it appends a versioned run record keyed by run_id to local on-disk history so a later TT05 call can diff a new run against the prior one for the same proxy/repo pair -- safe to retry, but not idempotent output-wise, since every run gets a fresh run_id. Setting BOTH live and confirmCost to true additionally makes real, provider-billed API calls against your own key (env-configured, never passed as a parameter) for up to liveMaxTasks tasks; omit either one and zero network calls are made. A failed or refused run (missing proxy binary, invalid task corpus, or the live safety gate declining the call) still returns a CallToolResult with isError=true and a JSON {ok: false, exit_code, message} body explaining why, instead of throwing.
Parameters: proxy (required) is a proxy name or array of names from {rtk, headroom} -- pass an array to run TT04's cross-tool comparison in one call. repo (optional) is a filesystem path, defaulting to this server's own working directory. tasks (optional) is a path to a tokentrust-tasks.yml corpus, defaulting to the bundled 23-task set. live/confirmCost (optional booleans, both false by default) gate real billed sampling as described above. liveMaxTasks (optional integer, default 5) caps how many tasks live mode samples. Example calls: {"proxy": "rtk"} for a standard estimated-cost run against the bundled corpus; {"proxy": ["rtk", "headroom"], "repo": "/path/to/target-repo"} for a side-by-side TT04 comparison against a specific repo; {"proxy": "rtk", "live": true, "confirmCost": true, "liveMaxTasks": 3} to verify the cost estimate against 3 real, billed samples.
Returns: the same structured JSON tokentrust verify --format json produces on success -- run_id, timestamp, repo, task_corpus_size, proxies, a records array (one entry per proxy/category with claimed_savings_pct vs measured_savings_pct), plus tt03 (never-worse guard pass/fail per proxy) and tt05 (version-drift regression pass/fail per proxy) maps. Run tokentrust verify --help on the command line for the full flag reference this schema mirrors.
| Name | Required | Description | Default |
|---|---|---|---|
| live | No | Sample real, provider-billed tokens for the first proxy instead of estimating from local pricing tables. Requires confirmCost=true in the SAME call, exactly like the CLI's --live/--confirm-cost safety gate -- setting only one of the two makes zero API calls and reports the refusal instead. Defaults to false. | |
| repo | No | Filesystem path to the repo to measure against. Defaults to the MCP server process's current working directory, same as the CLI's --repo default. | |
| proxy | Yes | Proxy name to verify. Pass a single name (e.g. "rtk") or an array of names to run the TT04 cross-tool comparison across all of them in one call -- mirrors the CLI's repeatable --proxy flag. Supported: rtk, headroom. | |
| tasks | No | Path to a task corpus YAML file. Defaults to the bundled task corpus shipped with the package, same as the CLI's --tasks default. | |
| confirmCost | No | Confirms the estimated spend `live` mode would print before any real, billed API call is made. Defaults to false. Has no effect unless `live` is also true. | |
| liveMaxTasks | No | Max tasks sampled in live mode. Defaults to 5. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and delivers: it discloses read-only repo behavior, the append of a versioned run record to local history, non-idempotent output, live network access gated by live+confirmCost, and failure/refusal returning isError=true. This far exceeds basic safety disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but organized into purpose, usage, side effects, parameters, and return sections. Every paragraph carries necessary context, though it repeats some schema parameter details. Given the tool's complexity and the absence of annotations/output schema, it is appropriately sized and structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description correctly enumerates the return structure (run_id, records, tt03/tt05 maps). It also covers prerequisites, side effects, failure modes, and live-mode safety, making the tool fully self-contained for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description adds value by explaining the array form for TT04 cross-tool comparison, the live/confirmCost gating semantics, example calls, and the headroom-not-yet-runnable nuance. However, it largely restates the schema details rather than introducing substantial new parameter-level semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Runs an independent, adversarial verification') and names the target resource (AI-coding-agent context-reduction proxy) and the method (TT01-TT05 engine, real tokenizer, bundled corpus). It clearly distinguishes this from a vendor benchmark rerun and from general token counting, which satisfies the top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states 'Call this when an agent needs a trustworthy...' and provides concrete examples ('before recommending a proxy, evaluating a version upgrade, or checking a CI regression'). It also gives exclusions: 'not for general token counting or for proxies outside {rtk, headroom}' and explains the headroom limitation. This is model usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
1 tool update
v0.1.0- First observed
verify_proxy_savings
TDQS
Only one tool exists, so there is no possibility of confusing it with another tool. The tool's purpose is clearly described, and its parameters are well-defined.
The single tool name follows a clear verb_noun pattern (verify_proxy_savings), and with only one tool there is no inconsistency to evaluate.
The server exposes only one tool, which feels too thin for a platform named TokenTrust that implies a broader verification suite. While the tool is complex, a single tool limits the server's ability to support related operations.
The tool covers the core verification workflow (standard, cross-tool comparison, live mode, regression diff), but lacks supporting operations such as listing supported proxies, retrieving historical runs, or accessing task corpus details. These are notable gaps that an agent might need.
Maintenance
Related MCP Connectors
MCP server providing access to the Scorecard API to evaluate and optimize LLM systems.
A paid remote MCP for OpenAI Codex context compressor, built to return verdicts, receipts, usage log
MCP-native AI evaluation: rubric audits, eval suites, and proof reports for AI/LLM output.
MCP server for progressive tool usage at any scale (see https://klavis.ai)
Related MCP Servers
- AlicenseAqualityBmaintenanceMCP server for verifying AI agent claims vs reality — single-transcript inline grounding-check that flags when an agent's response states facts not in the input context, when its code silently swallows exceptions and substitutes mock data, or when its multi-turn transcript contains contradictions or unverified completion claims. Sub-second, local, free, no API calls.41MIT
- AlicenseBqualityBmaintenanceAn MCP-style stdio server for evaluating AI agent outputs, enabling CI-friendly quality gates, regression comparisons, and canary promotion decisions.3MIT
- AlicenseNot gradedqualityBmaintenanceA transparent MCP proxy that independently re-verifies tool-call claims instead of trusting them, paired with devmcp — the git/CI server it's proven against.1MIT
- AlicenseAqualityBmaintenanceMCP server providing lossless file reads with deduplication, per-repo token metering, and HMAC-signed context receipts. Enables auditable, vendor-neutral measurement of what an AI agent saw.63223MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/RudrenduPaul/TokenTrust-CLI'
If you have feedback or need assistance with the MCP directory API, please join our Discord server