Skip to main content
Glama
AMEOBIUS-space

mcp-context-guard

MCP Context Guard — Context Window Management for AI Agents

PyPI Tests Dependencies License

Compress tool outputs, manage token budgets, deduplicate content, and filter by relevance. 14 tools. Zero dependencies. Pure Python stdlib.

Install

pip install mcp-context-guard

Requirements: Python 3.10+. Zero runtime dependencies (stdlib only).

Type checking: Ships with py.typed marker (PEP 561). Compatible with mypy, pyright, and pyrefly.

Related MCP server: compresh-mcp

The Problem

AI agents waste context window tokens on:

  • Verbose tool outputs (file reads, search results, logs)

  • Duplicate content across tool calls

  • Irrelevant passages that don't match the task

A single read_file on a large config can dump 3,200 tokens into context. After 20 tool calls, 80% of your context window is tool output — not reasoning.

The Solution

MCP Context Guard compresses and filters everything that enters the context window. 84% average token reduction across 5 strategies:

Strategy

When to use

Savings

Head-tail truncation

Large outputs with useful start/end

60-80%

Deduplication

Agent reads same file multiple times

50-90%

Relevance filtering

Search results, log output

70-95%

Semantic bucketing

Structured data (stack traces, logs)

65-85%

Budget enforcement

Runaway agent loops

Prevents overflow

Quick Start

from src.context_guard_engine import ContextGuard

guard = ContextGuard(max_tokens=500)

# Compress a large tool output
compressed = guard.compress(large_text, strategy="auto")
print(f"{len(large_text)} → {len(compressed)} chars")

# Deduplicate across calls
guard.deduplicate("content from call 1")
guard.deduplicate("content from call 2")  # detects overlap

# Filter by relevance
filtered = guard.filter_relevant(search_results, query="auth bug")

# Check budget
remaining = guard.get_stats()
print(f"Budget: {remaining['used']}/{remaining['total']} tokens")

14 Tools

Tool

What it does

compress

Extractive summarization to N tokens

set_budget

Set a total token budget

check_budget

Check if text fits remaining budget

consume_budget

Deduct tokens from budget

deduplicate

Remove near-duplicate texts (Jaccard similarity)

extract_key

Extract top-N key sentences

truncate_smart

Truncate at sentence boundaries

chunk

Split into token-sized chunks with overlap

token_count

Estimate token count (word-based heuristic)

summarize_history

Compress conversation messages

filter_relevant

BM25 relevance scoring, return top-K passages

merge_context

Combine sources with dedup + compression

get_stats

Context usage statistics

reset

Reset all state

MCP Server Setup

{
  "mcpServers": {
    "context-guard": {
      "command": "python3",
      "args": ["-m", "src.server"]
    }
  }
}

Real-World Results

Metric

Before

After

Avg tokens per tool call

4,200

680

Max tool calls before context full (16K)

12

50+

Duplicate content ratio

34%

2%

Agent task completion rate

71%

89%

Tests

python -m pytest tests/ -v  # 36 tests, all passing

Inspiration

License

MIT — see LICENSE

Freelance portfolio: https://ameobius-space.github.io/kwork-portfolio/

A
license - permissive license
-
quality - not tested
C
maintenance

Maintenance

Maintainers
Response time
Release cycle
Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

  • A
    license
    -
    quality
    C
    maintenance
    Provides production-grade context compression for LLM agent conversations with Q-protective ranking, epistemic markers, and semantic store, reducing token usage while preserving equivalence.
    3
    Business Source 1.1
  • F
    license
    A
    quality
    D
    maintenance
    Reduces token consumption by 73-87% by cleaning web and API data before it reaches the LLM context window. Supports fetching URLs, searching the web, optimizing JSON, and more.
    6
    1

View all related MCP servers

Related MCP Connectors

  • Universal memory for AI agents and tools. Save, organize and search context anywhere.

  • SaaS intelligence for AI agents. 5 unified tools cover 1,000+ services with 91-96% token savings.

  • See, price, and control every tool call your AI agents make: policy checks, cost, and audit tools.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/AMEOBIUS-space/mcp-context-guard'

If you have feedback or need assistance with the MCP directory API, please join our Discord server