Skip to main content
Glama

ctxjev

Keep what matters when your agent's context gets compacted.

In five lines: ctxjev trims an AI agent's history; by default the CLI and the Claude Code plugin send nothing anywhere. The CLI keeps the newest entries (recency, plain truncation); the Claude Code plugin uses keyword overlap (local). Jev scoring is opt-in there (--scorer jev, CTXJEV_SCORER=jev) and needs TYPESAFE_API_KEY; the MCP tools ask Jev unless a call passes scorer: "local" or "recency". On held-out tasks, no scorer has yet beaten plain truncation: see Does It Work?. Contributors: run pnpm build before this repo's .mcp.json can start the MCP server.

ctxjev ranks the entries of an AI agent's history and decides what to keep, drop, or summarize. By default it ranks by position alone (newest kept, the same as plain truncation), with no key and nothing sent; opt in to Jev, TypeSafe AI's typed-decision model, or to an offline keyword heuristic. In Claude Code, the plugin carries the top few entries through compaction (scored offline by default; Jev if you opt in). In an agent loop you write yourself, pruneMessages() removes what ranked lowest and keeps the request valid.

npm (ctxjev-cli) npm (ctxjev-core) npm (ctxjev-mcp) CI License: MIT Node TypeScript pnpm

Why · Choosing a Package · Claude Code Plugin · How Scoring Works · Does It Work? · Quick Start · MCP Hosts · Design Notes · Changelog


Why

Long-running agents accumulate context faster than it stays useful, and hosts deal with that by compacting: summarizing the whole history at once. A summary is lossy by nature. The one line that turned out to be the bug can get smoothed away along with everything that didn't matter.

Most of that history is easy to judge: "is this old tool result still relevant to the current task?" is the kind of fast, cheap, structured decision Jev is built for. It returns typed judgments (a yes/no probability, a choice, a score) instead of writing a sentence about it, so ctxjev can ask it about every entry cheaply. Whether its answers beat simpler rankings is what Does It Work? measures; on unseen tasks, so far, they don't.

What ctxjev does with the answer depends on where it runs. A host like Claude Code doesn't let anything remove entries from its context, so there ctxjev works alongside compaction and hands the most relevant entries back once it's done. An agent loop you write yourself owns its message list, so there you can drop what scored low before it's ever sent to the model again.

Jev can't see images, do arithmetic, or generate text. ctxjev never asks it to: token counting happens in code, and the keep/drop/summarize decision is a plain threshold applied to Jev's typed output. See CLAUDE.md for the full list of things this project deliberately never asks Jev to do.

Where things stand. A preregistered comparison on 6 tasks Jev's ranking had never seen found no difference in whether the agent finished the job: Jev and plain truncation both passed every task, with Claude Haiku 4.5 and with Claude Sonnet 5. So the default scorer is 'recency' (plain truncation) in ctxjev-core, ctxjev-cli, and pruneMessages(); Jev is opt-in via scorer: 'jev'. On the other measure — how much of what a task needs survives a tight budget — Jev beat truncation only because truncation scores nothing there by construction; it kept less than a random ordering of the same entries, and less than plain keyword overlap, the opposite of what the pre-holdout numbers below showed. Read Does It Work? before deciding whether to opt in; ctxjev-mcp, whose only job is exposing Jev, keeps calling it regardless of this default.

Related MCP server: @yavdaanalytics/context-optimiser

Choosing a Package

You want to…

Use

What it actually does

Keep key details through Claude Code's compaction

ctxjev-claude plugin

Scores the session right before compaction and re-injects the top few entries right after. Adds a short reminder; doesn't remove anything.

Drop stale history in an agent loop you control

ctxjev-core

pruneMessages() takes an Anthropic Messages conversation and returns it with stale entries removed (or cut to fit a token budget), every tool_use/tool_result pair kept intact, and reports what that costs a prompt cache. For other formats, pruneContext() returns keep/drop/summarize per entry.

See how a transcript would score, or prune a saved one

ctxjev-cli

analyze prints a report; prune writes the transcript back out with drops removed.

Expose scoring as a tool to an MCP host

ctxjev-mcp

Returns scores to whoever calls the tool. Read the caveat before expecting it to save tokens.

The Claude Code Plugin

Claude Code hooks can read the conversation transcript but can't rewrite it: there is no API for a hook to remove old entries before compaction summarizes them away. ctxjev-claude works within that constraint, using the one pattern Claude Code supports: score every entry at PreCompact, cache the highest-relevance ones, and re-inject a digest of them at SessionStart (matcher: "compact"), the documented way a hook can put content back into context after compaction has already run.

PreCompact             → score entries since the last compaction; cache the top few (~/.claude/ctxjev/)
  (Claude Code's own compaction runs, untouched)
SessionStart (compact) → print that cache as a digest; Claude Code adds it as a system reminder

Tool entries carry what was called (Bash(npm test): 12 passed, Read(src/payments.ts): …), so both the scoring and the reminder know which command or file a result came from. Only what's still in context counts: anything before the previous compaction is skipped.

The goal to score against is either set explicitly with /ctxjev:set-goal <text>, which applies to the current session only and lasts through compactions, or inferred from your first request plus your latest instruction. /ctxjev:status shows that goal and what the last run did, including why if it skipped, failed, or fell back to offline scoring. The plugin's measured effect so far is within the noise; see Does It Work?.

Privacy: by default the plugin scores offline by keyword overlap and sends nothing anywhere. Only with CTXJEV_SCORER=jev (and TYPESAFE_API_KEY) does it send excerpts of your session to TypeSafe AI's Jev API, with common secret formats masked to [REDACTED] first (best-effort, not exhaustive). Its cache lives in ~/.claude/ctxjev/, private to you, never in your project. See the plugin's Privacy section for exactly what's sent.

Install it (Claude Code desktop app or CLI):

/plugin marketplace add x96x64/ctxjev

This repo carries a .claude-plugin/marketplace.json at its root, pointing at the packages/claude-plugin subdirectory at the latest release's tag (never unreleased main), so the desktop app can install it directly with no local clone needed. ctxjev-claude isn't on npm; Claude Code plugins install through the marketplace, not npm install. See the plugin README for setup details, including how to make the key visible to the desktop app.

How Scoring Works

With Jev, every entry becomes its own question, and the questions for up to 50 entries are evaluated in parallel against one shared state, as a single request. Jev bills input tokens, and those still grow with the entries in a request (each entry's excerpt is part of the state); at Jev's published price that stays small, and the Jev run below shows its own usage.

import { pruneMessages } from 'ctxjev-core'

// `messages` is the Anthropic Messages conversation your agent loop sends each turn.
const { messages: pruned, removed } = await pruneMessages(
  messages,
  'Fix a bug where checkout charges customers twice on a slow network retry.',
)
// `pruned` is still a valid request: tool_use/tool_result pairs are removed together, and the
// first message and the latest turn (your last instruction and everything after it) are never
// touched. If your loop's only instruction is the first message, pass `protectLastTurn: false`,
// or that whole loop is the latest turn and nothing is removed.

Any other history shape works through pruneContext(entries, goal), which takes plain { id, role, toolName?, content, timestamp } entries and returns a decision per entry.

With no options, ctxjev analyze uses the default scorer, recency: position alone, no Jev call.

$ ctxjev analyze examples/sample-transcripts/checkout-bug.json
score: 0–1, position in the transcript (oldest 0, newest 1), not relevance: the goal isn't used · keep = leave as-is, summarize = worth shortening, drop = worth removing

  e1  bash       drop       score 0.00  ran: npm test -- checkout.test.ts — 12 passed, 0 failed
  e2  read       drop       score 0.17  read package.json — saw the dependency list and script names
  e3  grep       summarize  score 0.33  grep "charge" in src/payments.ts — found chargeCustomer() c…
  e4  bash       summarize  score 0.50  ran: git log --oneline -5 — recent commits about unrelated …
  e5  read       keep       score 0.67  read src/payments.ts — the retry handler re-calls chargeCus…
  e6  assistant  keep       score 0.83  Found it: the retry path doesn't check for an in-flight or …
  e7  bash       keep       score 1.00  ran: ls public/audio — unrelated, was checking something el…

3 keep, 2 summarize, 2 drop (of 7 entries)
prune would remove the 2 entries marked drop, ~27 / 154 tokens (18%); the 2 entries marked summarize (~45 tokens) stay as they are unless you shorten them yourself
⚠ would remove the first entry (e1): this transcript has no user entry to protect as the original request
Scored by position alone (newest kept, like plain truncation) — no Jev call, nothing sent.

That's plain truncation, and it shows: the unrelated ls public/audio is kept because it's newest, and the grep that found the bug is only marked for summarizing. With recency the goal isn't used at all: every entry's score is its position (oldest 0, newest 1), so the default thresholds (dropBelow 0.3, summarizeBelow 0.6) drop roughly the oldest 30% of entries and mark the next 30% for summarizing, whatever they say, the first request included. What prune then removes is protected, though: in ctxjev's own format it keeps the first user entry and the last two entries (--no-protect-first and --protect-last change that), and pruneMessages() never touches the first message, the latest turn, or by default anything the user wrote. The same sample with Jev:

$ ctxjev analyze examples/sample-transcripts/checkout-bug.json --scorer jev
score: 0–1, Jev's judgment of relevance to your goal, blended with recency · keep = leave as-is, summarize = worth shortening, drop = worth removing

  e1  bash       summarize  score 0.51  ran: npm test -- checkout.test.ts — 12 passed, 0 failed
  e2  read       drop       score 0.28  read package.json — saw the dependency list and script names
  e3  grep       keep       score 0.84  grep "charge" in src/payments.ts — found chargeCustomer() c…
  e4  bash       drop       score 0.14  ran: git log --oneline -5 — recent commits about unrelated …
  e5  read       keep       score 0.88  read src/payments.ts — the retry handler re-calls chargeCus…
  e6  assistant  keep       score 0.90  Found it: the retry path doesn't check for an in-flight or …
  e7  bash       drop       score 0.15  ran: ls public/audio — unrelated, was checking something el…

3 keep, 1 summarize, 3 drop (of 7 entries)
prune would remove 2 of the 3 entries marked drop, ~28 / 154 tokens (18%); the 1 entry marked summarize (~16 tokens) stays as it is unless you shorten it yourself
1 marked drop but kept: 1 protected in the last 2 entries
Jev cost: 1,290 input tokens, 123 output tokens (free) — ~$0.000054

Both are real output against the sample transcript in this repo, captured 2026-09-25. Jev is probabilistic, so its numbers vary between runs, and its cost line is computed from the usage Jev's API reported for that request, not estimated. "Score" is the ranking's relevance blended with each entry's recency within the batch, described further in Design Notes. Only dropped entries count as saved. Jev doesn't generate text, so what summarize saves depends on what you do with those entries: ctxjev prune --summarize-excerpts cuts them to their head and tail, and pruneMessages() can also hand them to your own summarizer.

The entries array above is the one shape every agent's history maps onto, regardless of host. ctxjev analyze also auto-detects a real Claude Code session .jsonl and infers the goal from your first request plus your latest instruction unless --goal overrides it. See examples/sample-transcripts/claude-code-session.jsonl for a synthetic one. Be careful pointing it at a real session log with --scorer jev: entry content is sent to the live Jev API, and although common secret formats are masked first, that masking can't catch everything. The default scorer, and --scorer local, send nothing.

Does It Work?

The numbers below come from the dev and holdout splits. Dev (15 of the 21 sessions: 5 written by hand and 10 recorded in examples/eval-sessions, and the tasks they were recorded on) informed 0.5.0's design — keepUserText, the removal note, and the plugin's goal inference were all built from failures seen on it, so read those numbers as optimistic. Holdout (6 further tasks, recorded and labeled after that design was frozen, preregistered before any of them were scored) is what actually decided the default scorer, below. It has now been run and looked at closely, so it can't confirm anything new; a fresh holdout is part of the Round 2 design proposal (in Japanese).

Every number in this section is generated from the saved results in eval/results/ by eval/check-docs.mjs, and CI fails if they disagree. So does a number typed into the prose here instead of generated.

Holdout: does Jev's ranking beat plain truncation?

No, on task success; and on what survives a tight budget, no better than chance. 6 tasks (rate-limit-window, coupon-stacking, upload-size-limit, audit-retention, shipping-fee, room-booking), each with a mid-session constraint and another changed or added later, run through eval/tasks.mjs and eval/run.mjs exactly as preregistered:

Task success (3 runs/task/model, budget 25%)

Claude Haiku 4.5

Claude Sonnet 5

Everything

100%

not run

Pruned by Jev, as shipped

100%

100%

Pruned by plain truncation, same options

100%

100%

Goal only

0%

not run

Jev minus truncation: 0 points, 95% CI [0, 0] with Claude Haiku 4.5 and 0 points, 95% CI [0, 0] with Claude Sonnet 5. Every task passed under every history condition tested, on every run; only removing the history entirely (goal-only) failed. That's the preregistered primary rule, and its answer is that Jev's ranking made no difference here.

The preregistered secondary measure asks what the ranking alone keeps: the share of each session's labeled facts (probes) still present after pruning to a 25% budget, with keepUserText and the removal note off. Alongside Jev and truncation are keyword overlap, a random order, and the labels themselves:

What survives a 25% budget (ranking alone, no keepUserText; Jev: mean of 3 runs)

Whole session (v1, preregistered)

Up to the fix request (v2, exploratory here)

Jev

23.6%

41.1%

Plain truncation (newest kept)

0.0%

31.0%

Keyword overlap (scorer: 'local', offline, free)

28.3%

37.5%

Random order (mean of 20 seeds)

26.5%

33.5%

The labels themselves (relevant entries first)

32.7%

31.0%

On the holdout, Jev kept less of what the tasks needed than a random ordering of the same entries did (23.6% vs. 26.5%), and less than keyword overlap (28.3%). It beat plain truncation only because truncation scores 0% on the preregistered measure by construction (see below).

Preregistered measure: Jev minus plain truncation is +23.6 points, 95% CI [+9.5, +37.7]. Exploratory comparisons on the same measure: Jev minus random order is −2.9 points [−12.4, +6.4], and Jev minus keyword overlap −4.7 [−19.1, +11.7]. Up to the fix request (v2): Jev minus plain truncation is +10.1 points [+0.7, +22.0], and Jev minus random order +7.6 [−2.4, +19.7].

Read the preregistered column with these in mind, both found by an independent audit:

  • It measures the whole recorded session, including the implementation after the fix request. On the holdout sessions that part is a large share of the tokens but holds none of the labeled facts, so plain truncation, which keeps the newest entries, can't score anything by construction. (Each dev session labels a fact there, which is part of why the splits disagree.) The "up to the fix request" column measures what tasks.mjs actually prunes instead. It was added after the results were seen, is labeled exploratory in PREREGISTRATION.md, and decides nothing.

  • Jev's answers vary between runs. This is a saved re-run of the preregistered command; earlier runs, whose output wasn't saved, gave different values (all of them are in PREREGISTRATION.md). With this few sessions, differences of a few points are within that variation.

On the dev sessions, Jev beat keyword overlap by a wide margin (85.6% vs. 65.3% at a 25% budget); on the holdout it didn't.

What this changes: the default scorer for ctxjev-core, ctxjev-cli, and pruneMessages() is now 'recency' (plain truncation). Pass scorer: 'jev' / --scorer jev to opt in. PREREGISTRATION.md's Results section has the full numbers and commands.

Holdout: does the Claude Code plugin's digest help?

No demonstrated effect. The preregistered plugin comparison (eval/plugin.mjs --split holdout: the summary of a simulated compaction alone, against the same summary plus the plugin's digest, decided on tasks passed) errored on its first attempt because the hook couldn't reach Jev from the recording sandbox. It was then run to completion on 2026-09-24 in two separate sessions on the same code, but both results sat on branches that were never merged until they were recovered on 2026-09-25. Both runs are shown; neither was chosen in advance as the run.

After a simulated compaction (holdout, Claude Haiku 4.5)

Run d8aa0b1: tasks passed

answers right

Run 042cf4c: tasks passed

answers right

Summary alone

100%

78%

100%

79%

Summary + digest (goal inferred by the plugin)

94%

88%

83%

82%

Summary + digest (the same goal, set with /ctxjev:set-goal)

100%

88%

89%

85%

Digest (inferred goal) − summary alone, 95% CI

−6 [−17, 0]

+10 [+5, +14]

−17 [−39, 0]

+3 [−3, +9]

Digest (set goal) − summary alone, 95% CI

0 [0, 0]

+10 [+1, +16]

−11 [−22, 0]

+6 [−1, +15]

Run d8aa0b1: 6 tasks × 3 runs, and the hook scored with Jev in 18 of 18; Run 042cf4c: 6 tasks × 3 runs, and the hook scored with Jev in 18 of 18. In both runs the two digest conditions scored against the identical goal (18 of 18, 18 of 18), so they are the same configuration measured twice (the goal supplied two ways), not two different goals. No tasks-passed interval clears zero in either run, so by the preregistered rule the digest has no demonstrated effect on unseen tasks. Answers right isn't the registered measure; it's shown because the two runs disagree there too.

What this changes: with no demonstrated effect from the Jev-scored digest, and Jev's ranking below keyword overlap on the retention measure above, the plugin now scores offline by keyword overlap by default and sends nothing; Jev is opt-in with CTXJEV_SCORER=jev. That isn't evidence the offline digest helps either: it hasn't been shown to.

Dev sessions (optimistic — see above)

  • Written sessions. Written by hand, with raw tool output, dead ends, and distractors that share the goal's words.

  • Recorded sessions. Real Claude Code sessions, recorded on the throwaway task repos in examples/eval-tasks. Each task has a planted bug. The user states constraints partway through, and in some of them changes the plan. The session ends with the fix.

Intervals are 95% bootstrap intervals that resample whole sessions or tasks. With this few tasks they're wide, and they're shown so you can see how wide.

1. Can an agent still finish the job? eval/tasks.mjs cuts each recorded session before "now implement the fix" and prunes the history to 25% of its tokens. Claude Haiku 4.5 then does the fix with real tools in a fresh copy of the repo. It passes if the hidden acceptance tests pass, and those tests include the constraints the user stated mid-session. 10 tasks × 2 runs:

History given to the agent

Tasks passed

95% CI

Everything

100%

[100, 100]

Pruned to 25% by Jev, as shipped (keeps user text, marks the gap)

100%

[100, 100]

Newest kept, with the same two options

90%

[70, 100]

Pruned by Jev ranking alone

90%

[75, 100]

Newest kept (plain truncation)

90%

[70, 100]

Pruned by keyword overlap

75%

[55, 95]

Only the task

40%

[15, 70]

Against truncation with the same options (user text kept, gap marked), the fair comparison, Jev is +10 points [0, +30]: truncation failed 2 of its 20 runs and Jev 0, all on webhook-dedupe. On the other 9 tasks both passed every run.

The misses were agents that lost something the user said and filled the gap with their own guess. An agent lost "keep the mark for 24 hours" and kept a later "maybe 25h for margin". Another lost the exact masking format. That's why pruneMessages() now keeps what the user wrote and adds a short note where it removed history. Both were designed from these failures, which is why these tasks can't also be the test of them.

With a stronger agent (Claude Sonnet 5), the difference goes away. Same 10 tasks × 2 runs, effort low: Jev as shipped passed 90%, and truncation with the same options passed 95% (−5 points [−15, 0]). Both passed every run on 9 of the 10 tasks. The rest is the "24 hours vs. maybe 25h" task (webhook-dedupe: Jev lost it in 2 of 2 runs, truncation in 1), where the recorded assistant's later suggestion contradicts the user. So at a 25% budget on tasks this size, a strong agent recovers from pruning whichever way it's done. What made the difference for Haiku was the weaker agent, not the ranking alone.

2. Can a model still answer from what's left? eval/outcome.mjs has Claude Haiku 4.5 answer each needed fact as a question from the pruned conversation, and Claude Sonnet 5 grade it (102 questions, 2 runs):

Context

25% budget

50% budget

Everything

95%

95%

Pruned by Jev

79%

91%

Plain truncation

71%

83%

Keyword overlap

67%

74%

The differences: Jev minus truncation is +8 points [−1, +18] at 25% and +8 [−1, +19] at 50%. Jev minus keyword overlap is +13 [+7, +18] and +17 [+10, +24]. The truncation intervals include zero. Ranking alone (eval/run.mjs, the dev sessions, Jev: mean of 3 runs) keeps 85.6% / 95.6% of the facts under a 25% / 50% budget, against 66.3% / 80.4% for truncation, 65.3% / 79.8% for keyword overlap, and 62.0% / 69.3% for a random order. Every release is gated on Jev not falling below truncation or keyword overlap there at a 50% budget (eval/run.mjs --gate --runs 3, which fails without a Jev key).

3. Does the Claude Code plugin help after a compaction? eval/plugin.mjs runs the shipped hooks on each recorded history. Claude Haiku 4.5 simulates the compaction summary; Claude Code's own compaction prompt isn't public, so ours only approximates it, and it asks for every user instruction. The agent then finishes the task from the summary, with or without the plugin's digest (10 tasks × 2 runs):

After compaction

Tasks passed

Answers right

Summary alone

95%

89%

Summary + digest (goal = latest message, as in 0.4.0)

90%

91%

Summary + digest (goal = first request + latest instruction, 0.5.0)

100%

90%

Honest reading: against a summary that already keeps every user instruction, the digest adds little. The 0.4.0 plugin aimed its scoring at the latest message ("also check the tests"), not the task, and did no better than the summary alone. 0.5.0 infers the goal from the first request plus the latest instruction. That passed every task, but +5 points [0, +15] is within the noise. The plugin's value depends on how much the compaction summary drops, and this eval can't measure Claude Code's real compaction.

What this doesn't show:

  • Any of it on unseen material. The dev tasks and sessions informed the design, and the holdout has since been run and analyzed, so neither can confirm a new claim.

  • The tasks are small, and 10 tasks × 2 runs is a small sample.

  • Most runs use Claude Haiku 4.5. Claude Sonnet 5 was checked on the task eval only, with fewer conditions.

  • The recordings come from the same recording model on tasks written for this eval.

  • Token counts come from gpt-tokenizer, an approximation of Claude's tokenizer.

Every answer, verdict, digest, summary, and agent run is in eval/results/.

Quick Start

npm install -g ctxjev-cli

# A small synthetic Claude Code session from this repo; or point it at one of your own.
curl -O https://raw.githubusercontent.com/x96x64/ctxjev/main/examples/sample-transcripts/claude-code-session.jsonl
ctxjev analyze claude-code-session.jsonl
$ ctxjev analyze claude-code-session.jsonl
score: 0–1, position in the transcript (oldest 0, newest 1), not relevance: the goal isn't used · keep = leave as-is, summarize = worth shortening, drop = worth removing

  u1         user       drop       score 0.00  fix the memory leak that crashes the server after a few hou…
  call1      Bash       summarize  score 0.32  Bash(node --heap-prof server.js): heap snapshot saved — ret…
  call2      Read       keep       score 0.66  Read({"file":"README.md"}): general project setup instructi…
  a3:text:0  assistant  keep       score 1.00  Found it — the websocket disconnect handler never calls rem…

2 keep, 1 summarize, 1 drop (of 4 entries)
prune can't write back a Claude Code transcript, so nothing is removed (entries marked drop hold ~14 / 81 tokens, 17%)
Scored by position alone (newest kept, like plain truncation) — no Jev call, nothing sent.

No key is needed: the default scorer, recency, ranks by position alone and sends nothing. --scorer local scores by keyword overlap instead, also offline. Want Jev's judgment? export TYPESAFE_API_KEY=... (console.typesafe.ai/settings/keys, no waitlist) and add --scorer jev — see Does It Work? for what that currently buys you.

ctxjev prune writes a ctxjev-format or Anthropic Messages transcript back out with the drops removed (to stdout, or --out <file>). With the defaults, on the same sample as the first report in How Scoring Works:

$ ctxjev prune examples/sample-transcripts/checkout-bug.json --out pruned.json
removed 2 of 7 entries, ~27 tokens · scored by position alone
⚠ removed the first entry (e1): this transcript has no user entry to protect as the original request

e1 and e2, the two entries that report marks drop, are gone from pruned.json, and every other field of the file (here groundTruth) is written back as it was. This sample has no user entry (it starts with a test run), so nothing stood in for the original request, and prune says so. In ctxjev's own format prune never removes the first user entry or the last two entries (--no-protect-first, --protect-last <n>; it warns if the first user entry goes). In an Anthropic Messages transcript, prune by default never touches the first message or the latest turn (your last instruction and every tool call after it; --no-protect-last-turn lets that turn be pruned), and skips any removal that would save less than the note it leaves in its place, so a short conversation can come back with nothing removed.

Or from a clone of main, to run the exact sample transcript above:

git clone https://github.com/x96x64/ctxjev.git
cd ctxjev
pnpm install && pnpm build

node packages/cli/dist/index.js analyze examples/sample-transcripts/checkout-bug.json               # recency
node packages/cli/dist/index.js analyze examples/sample-transcripts/checkout-bug.json --scorer jev  # needs TYPESAFE_API_KEY

The Claude Code plugin isn't on npm: the marketplace installs it from this repository, at the tag of the latest release (v0.7.1), so plugin users get released code only.

Packages

This is a pnpm workspace monorepo: one host-agnostic engine, and a thin adapter for each place that engine gets used.

Package

What it is

Status

ctxjev-core (npm)

The engine: scoreEntries()/pruneContext()/pruneMessages(), plus the Claude Code transcript parser, secret masking, and the offline scorers. Everything else wraps this.

✅ published (npm: 0.7.1)

ctxjev-cli (npm)

ctxjev analyze (a report) and ctxjev prune (the transcript with drops removed).

✅ published (npm: 0.7.1)

ctxjev-mcp (npm)

MCP server exposing score_relevance/prune_history as tools.

✅ published (npm: 0.7.1)

ctxjev-claude

Claude Code plugin: scores at PreCompact, re-injects a digest at SessionStart, plus /ctxjev:set-goal and /ctxjev:status (answered by a UserPromptSubmit hook without a model turn).

✅ working (not on npm; installed from the release tag)

Using It from an MCP Host

Read this first. An MCP tool returns data to whoever called it; it can't remove anything from the host's own context. And to score its history, the agent has to send that history as tool arguments, which the host model pays for in its own output tokens. So in a host like Claude Code, Codex, or Copilot, calling prune_history doesn't save tokens by itself. It's useful when something acts on the scores: an agent framework that manages its own context and exposes tools, or a workflow where knowing what's stale matters more than the cost of asking. For Claude Code specifically, the plugin is the integration that actually helps.

ctxjev-mcp is a plain stdio server with two tools, score_relevance (scores per entry) and prune_history (a keep/drop/summarize decision per entry, plus a savings report). Unlike the core default ('recency', see Does It Work?), both ask Jev unless a call passes scorer: "local" or scorer: "recency", which run offline; Jev needs TYPESAFE_API_KEY. Setup for Claude Code, Codex (including this repo's Agent Plugins bundle), and GitHub Copilot, and an example response, are in the ctxjev-mcp README.

Design Notes

  • Claude Code's transcript format stays in one module. claudeCodeTranscript.ts parses Claude Code's undocumented session log, so a format change is a one-file fix. It returns only what's still in context (a compaction boundary resets the list), skips subagent sidechains (their result already appears as a tool call), and skips text Claude Code writes into the user turn itself: meta records, local command output, interrupt notices.

  • Every chunk sees the latest activity. Entries are scored in chunks of 50, and each chunk's shared state also carries the batch's most recent entries, so an old failing test is judged knowing a later run fixed it. The score cache is keyed on that context too.

  • Secrets are masked before anything leaves the machine. Every Jev request is built in buildJevRequest(), which runs goal and content through redactSecrets() first. Best-effort, not a guarantee. Measured blind, on lines a separate agent wrote without seeing the masking code, each corpus's holdout half once: Round 3's corpus, 0.7.0: 84 of 91 lines with secrets masked (92.3%), and 6 of 42 harmless lines changed (14.3%). Round 4's corpus, 0.7.1: 80 of 92 lines with secrets masked (87.0%), and 3 of 41 harmless lines changed (7.3%). Misses and false alarms that remain are listed in issue #16.

  • The scorer is pluggable, and the baselines ship. scorer takes 'jev', 'recency' (plain truncation), 'local' (keyword overlap), or your own function, called per chunk with content already masked. The eval compares against the first two on every run.

  • Savings count what's actually removed, at its real size: each entry's sourceTokens, the full payload, not the 600-character excerpt that gets scored. What summarize saves is reported separately, since it depends on your summarizer.

  • Pruning a conversation keeps it a valid request. pruneMessages() removes a tool_use and its tool_result together, never touches the first message or the latest turn (from the last user message with text of its own onward: protectLastTurn), and reports how much of a prompt cache the change invalidates. See prompt caching.

  • Recency is relative to the batch, oldest 0 to newest 1, not to Date.now(), so a transcript analyzed after the fact scores the same as a live one. combinedScore blends it in linearly at recencyWeight (default 0.1, from a sweep over labeled fixtures, one built so the root cause is early). dropBelow is 0.3, the highest value that lost no relevant entry on either fixture set.

  • Jev is never asked to count. Token counts come from gpt-tokenizer, and costs from the usage Jev's API reports for each request (onUsage).

Contributing

Issues and pull requests are welcome; see CONTRIBUTING.md, including the release policy. Report a vulnerability privately, as SECURITY.md describes.

pnpm install && pnpm build && pnpm test

The packages run on Node 20 or later; the eval scripts under packages/core/eval and examples/eval-tasks need Node 22 (.nvmrc), and say so if run on anything older.

pnpm test runs the full pure-logic suite with no API key and no network access. Tests that call the live Jev API end in .live.test.ts and are skipped automatically unless TYPESAFE_API_KEY is set. If you change anything under packages/core/src or packages/claude-plugin/src, commit the rebuilt packages/claude-plugin/dist/ too; CI checks that it matches. When developing, use the synthetic transcripts in examples/sample-transcripts, never a real session log (see CLAUDE.md).

Acknowledgments

Built on Jev, TypeSafe AI's System One model, via the official @typesafe-ai/sdk. ctxjev-mcp is built on Anthropic's @modelcontextprotocol/sdk. ctxjev is an independent, unofficial project, not affiliated with or endorsed by TypeSafe AI or Anthropic.

License

This project is released under the MIT license: free to use, modify, and distribute, including in a commercial product, as long as the license text and copyright notice in LICENSE ship with it. It comes with no warranty of any kind; see the license text for the full disclaimer.

The packages ctxjev depends on directly are all under MIT or ISC:

ISC and MIT are both short, permissive licenses with no material difference in what they let you do. Their own dependencies (what npm install pulls in beneath them, mostly under @modelcontextprotocol/sdk) are permissive too, but not all MIT or ISC: fast-uri and qs are BSD-3-Clause and json-schema-typed is BSD-2-Clause, which also ask you to keep their notices. pnpm licenses list --prod lists every one. The Claude Code plugin bundles code from @typesafe-ai/sdk into its dist/, and ships its notice in packages/claude-plugin/THIRD_PARTY_NOTICES.

Related MCP Connectors

Related MCP Servers