Skip to main content
Glama
README.md
# agentfit-mcp

**MCP server for [`@mukundakatta/agentfit`](https://www.npmjs.com/package/@mukundakatta/agentfit).** Lets Claude Desktop, Cursor, Cline, Windsurf, Zed, or any other MCP client estimate token counts and fit a chat history into a model's context budget on demand.

```bash
npx -y @mukundakatta/agentfit-mcp
```

Three tools:

- **`count_tokens`** — estimate tokens in a string or chat-message array, with per-model estimator families (openai, anthropic, google, llama, default).
- **`fit_messages`** — drop messages from a chat history until under a `maxTokens` budget. Supports drop-oldest, drop-middle, and priority strategies; honors `preserveSystem`, `preserveFirstN`, `preserveLastN`.
- **`list_estimators`** — list the built-in estimator families.

## Add to your client

### Claude Desktop

Edit `~/Library/Application Support/Claude/claude_desktop_config.json` (macOS) or `%APPDATA%\Claude\claude_desktop_config.json` (Windows):

```json
{
  "mcpServers": {
    "agentfit": {
      "command": "npx",
      "args": ["-y", "@mukundakatta/agentfit-mcp"]
    }
  }
}
```

### Cursor

`~/.cursor/mcp.json`:

```json
{
  "mcpServers": {
    "agentfit": {
      "command": "npx",
      "args": ["-y", "@mukundakatta/agentfit-mcp"]
    }
  }
}
```

### Cline / Windsurf / Zed

Same shape as above. The server speaks plain MCP over stdio, so any client that supports stdio MCP servers will work.

## Tool examples

**`count_tokens`:**

```json
{ "input": "hello world", "model": "claude-sonnet-4-6" }
```

Returns:

```json
{ "tokens": 4, "model": "claude-sonnet-4-6" }
```

**`fit_messages`:**

```json
{
  "messages": [
    { "role": "system", "content": "You are precise." },
    { "role": "user", "content": "long context..." },
    { "role": "assistant", "content": "..." },
    { "role": "user", "content": "final question" }
  ],
  "maxTokens": 8000,
  "model": "claude-sonnet-4-6",
  "preserveSystem": true,
  "preserveLastN": 2,
  "strategy": "drop-oldest"
}
```

Returns:

```json
{
  "messages": [...],
  "dropped": [...],
  "tokens": { "before": 12000, "after": 7800, "budget": 8000 },
  "fit": true
}
```

`fit_messages` always returns a structured result and never throws across the wire: if the budget is unreachable even after dropping all non-protected messages, you get `fit: false` with the partial result so the caller can decide what to do.

## Why a separate MCP server

`@mukundakatta/agentfit` is a zero-dependency JavaScript library. This package wraps it as an MCP server so it's accessible from inside any MCP-aware AI assistant: ask Claude "how many tokens is this transcript?" or "trim this chat to 8k tokens preserving the system prompt and last 2 turns" and the assistant calls these tools directly.

## Sibling MCP servers

Part of the agent-stack series, all `@mukundakatta/*-mcp`:

- [`@mukundakatta/agentfit-mcp`](https://www.npmjs.com/package/@mukundakatta/agentfit-mcp) — *Fit it.* (this)
- [`@mukundakatta/agentguard-mcp`](https://www.npmjs.com/package/@mukundakatta/agentguard-mcp) — *Sandbox it.*
- [`@mukundakatta/agentsnap-mcp`](https://www.npmjs.com/package/@mukundakatta/agentsnap-mcp) — *Test it.*
- [`@mukundakatta/agentvet-mcp`](https://www.npmjs.com/package/@mukundakatta/agentvet-mcp) — *Vet it.*
- [`@mukundakatta/agentcast-mcp`](https://www.npmjs.com/package/@mukundakatta/agentcast-mcp) — *Validate it.*

## License

MIT

TDQS

A4.4/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: counting tokens, fitting messages under a token limit, and listing available estimator families. No overlap or confusion.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern in snake_case (count_tokens, fit_messages, list_estimators), making them predictable and easy to understand.

Tool Count5/5

With 3 tools, the server is tightly scoped to token estimation and message fitting. Each tool earns its place, and the count is appropriate for this focused domain.

Completeness4/5

The tool surface covers core operations: counting tokens, fitting messages with multiple strategies, and listing estimators. A minor gap is the lack of a tool for direct model-specific tokenization, but the per-family estimator covers most use cases.

Maintenance

ActivityInactive
ResponsivenessNo issues