sidecar-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@sidecar-mcpUse bulk_read to summarize src/utils.ts and src/helpers.ts"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
sidecar-mcp
An MCP server that delegates bulk file reading to a cheap worker LLM, so the main agent's context stays small.
When the main agent needs to understand 2+ files or any file over ~100 lines, it calls bulk_read. The full file content never enters the main agent's context — only the worker's summary does.
Install
sidecar-mcp is a small stdio subprocess that sits next to your coding agent. Reads happen out-of-band to a cheap worker LLM, so the main agent's context stays small. Two install paths — both wired in under a minute.
Any MCP client (canonical)
The MCP primitive is the same everywhere — a subprocess with a command and some env vars. Every client wraps it in its own config syntax. Pick your worker backend, set SIDECAR_BACKEND (and any keys), then plug this into your client's MCP config:
{
"command": "npx -y @aemrezorlu/sidecar-mcp",
"env": { "SIDECAR_BACKEND": "ollama" }
}For a from-source install, swap npx -y @aemrezorlu/sidecar-mcp for node /absolute/path/to/sidecar-mcp/dist/index.js. See Wire into any MCP client below for client-specific config locations.
A. Claude Code CLI (once the package is on the npm registry)
If you're on Claude Code, claude mcp add is the shortest path — it writes the same config for you. The SIDECAR_BACKEND=* env vars are independent of the client; claude here is just the CLI that registers the MCP server.
# Ollama — local, free, no API key:
claude mcp add sidecar -e SIDECAR_BACKEND=ollama -- npx -y @aemrezorlu/sidecar-mcp
# OpenAI:
claude mcp add sidecar -e SIDECAR_BACKEND=openai -e SIDECAR_OPENAI_KEY="$OPENAI_API_KEY" -- npx -y @aemrezorlu/sidecar-mcp
# Anthropic (or any Anthropic-compatible provider):
claude mcp add sidecar -e SIDECAR_BACKEND=anthropic -e SIDECAR_ANTHROPIC_KEY="$ANTHROPIC_API_KEY" -- npx -y @aemrezorlu/sidecar-mcp
# Add -e SIDECAR_ANTHROPIC_URL=https://your-host for Anthropic-compatible proxies.npx -y @aemrezorlu/sidecar-mcp downloads and runs the published package on first call. No clone, no build, no node_modules to manage.
B. From source (works today, no publish needed)
git clone https://github.com/dEMonaRE/sidecar-mcp.git
cd sidecar-mcp
pnpm install --frozen-lockfile
pnpm build
claude mcp add sidecar -e SIDECAR_BACKEND=ollama -- node "$PWD/dist/index.js"Pick a backend
Backend | Cost | Needs |
| free, local |
|
| $$ |
|
| $$ |
|
Verify
In your MCP client (Claude Code shown), ask: "Use bulk_read to summarize README.md." You should see bulk_read fire and return a tight summary. Every reply ends with a usage footer (tokens: <prompt> in / <completion> out):
---
sidecar: model=llama3.1:8b, backend=ollama, tokens=412 in / 87 outIf bulk_read doesn't show up:
Claude Code:
claude mcp listshould showsidecaras connected.VS Code Copilot: Command Palette → "MCP: List Servers".
Codex CLI:
codex mcp list.Cursor / Zed: check the MCP panel in settings.
Related MCP server: Ultra-Eye
Configuration
sidecar-mcp reads everything from env vars. No config file. No CLI flags (MCP stdio can't pass them).
Var | Default | Notes |
|
|
|
| per-backend default | see below |
|
| |
|
| any OpenAI-compatible endpoint |
| required for openai | |
|
| any Anthropic-compatible endpoint |
| required for anthropic | also accepts Anthropic-compatible providers |
|
| files larger are skipped + reported |
|
| bulk_read errors if total bytes across all readable files exceeds cap |
| cwd (with stderr warning) | comma-separated absolute paths; see Security |
|
| |
|
|
|
Default models:
ollama →
llama3.1:8bopenai →
gpt-4o-minianthropic →
claude-3-5-haiku-latest
Per-backend quick config
Ollama (local, free, no API key)
ollama serve &
ollama pull llama3.1:8b
export SIDECAR_BACKEND=ollamaOpenAI
export SIDECAR_BACKEND=openai
export SIDECAR_OPENAI_KEY=sk-...Anthropic (or any Anthropic-compatible provider)
export SIDECAR_BACKEND=anthropic
export SIDECAR_ANTHROPIC_KEY=sk-ant-...
# Optional — point at a proxy that speaks the Anthropic Messages API:
# export SIDECAR_ANTHROPIC_URL=https://your-anthropic-compatible-hostWire into any MCP client
The MCP spec is the same everywhere — sidecar-mcp is a subprocess with a command and some env vars. Every MCP client wraps that primitive in its own config syntax, but the primitive itself doesn't change.
Canonical shape (the only thing you actually need to know):
{
"command": "npx -y @aemrezorlu/sidecar-mcp",
"env": { "SIDECAR_BACKEND": "ollama" }
}For a from-source install, swap npx -y @aemrezorlu/sidecar-mcp for node /absolute/path/to/sidecar-mcp/dist/index.js.
Where each client stores it:
Client | Config location | Key |
Claude Code |
|
|
VS Code Copilot |
|
|
Codex CLI |
|
|
Cursor |
|
|
Zed |
|
|
any other MCP client | — |
Example — VS Code Copilot (.vscode/settings.json):
{
"github.copilot.chat.mcp.servers": {
"sidecar": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@aemrezorlu/sidecar-mcp"],
"env": { "SIDECAR_BACKEND": "ollama" }
}
}
}Example — Codex CLI (~/.codex/config.toml):
[mcp_servers.sidecar]
command = "npx"
args = ["-y", "@aemrezorlu/sidecar-mcp"]
[mcp_servers.sidecar.env]
SIDECAR_BACKEND = "ollama"The command/args split varies by client (some take a single string, some take an array); the primitive above is what every client is configuring.
Tool reference
bulk_read
Field | Type | Required | Notes |
| string[] | yes | 1–50 paths. Relative or absolute. |
| string | yes | what to ask the worker |
| string | no | override configured default for this call |
Returns: the worker's text summary. The full file content never appears in the caller's context.
Response footer: every bulk_read reply ends with a one-line footer for cost spot-checks:
---
sidecar: model=<model>, backend=<backend>, tokens=<in> in / <out> out(or tokens: n/a if the backend didn't report usage). Footer is included automatically; no flag to disable in MVP.
Example call:
{
"paths": ["src/Service.java", "src/Handler.java"],
"question": "What does this service do and what are its key methods?"
}Worker prompt shape (for debugging):
<files>
<file path="src/Service.java">…</file>
<file path="src/Handler.java">…</file>
<!-- unreadable: path/to/binary.bin — binary file -->
</files>
<question>
What does this service do?
</question>How it works
main agent (whatever MCP client you wired it into)
│
│ MCP stdio JSON-RPC:
│ {"method":"tools/call","params":{"name":"bulk_read", ...}}
▼
sidecar-mcp subprocess
│
│ reads files, wraps in XML, POSTs to:
▼
worker model (any configured backend — Ollama / OpenAI / Anthropic / proxy — cheap)
│
│ returns ~600 token summary
▼
back to main agent as tool resultThe full file content stays between sidecar-mcp and the worker. The main agent only ever sees the summary.
Security
bulk_read reads files from the local filesystem and ships their content to the worker LLM. Two layers scope what the worker can see:
SIDECAR_ALLOW_ROOTSis a comma-separated allowlist of absolute paths. Anything outside is skipped withpath outside allowed roots. If unset, sidecar-mcp defaults to the current working directory and prints a one-line warning to stderr at boot. Set it explicitly for any non-dev use.Symlink escape is blocked. Every file is checked both lexically (path prefix) and via
realpath(canonical target). A symlink insideallowRootsthat points outside is rejected withpath resolves outside allowed roots (symlink escape).allowRootsare also realpath'd, so/varvs/private/varstyle mounts collapse consistently.
Per-file safety: SIDECAR_FILE_MAX_BYTES (default 512 KB) skips oversize files, NUL-byte sniff skips binaries, and the worker's reply is the only thing that returns to the MCP client — the main agent's context never holds raw file content.
Cwd default is a footgun. Launching from $HOME exposes ~/.ssh, ~/.aws, .env to the worker. For any deployment, set SIDECAR_ALLOW_ROOTS to a tight scope (e.g. the project root) instead of relying on the default.
Limitations
Non-streaming. Worker replies are returned as one text blob.
No
code_write. Bulk boilerplate generation is out of scope for MVP. If you need it, build a separate tool.Claude Code hook layer is opt-in. Auto-redirect of large
Read/Bash cat|head|tailcalls tobulk_readlives inextras/claude-hooks/and is not installed by default. Seeextras/claude-hooks/README.mdto wire it into~/.claude/settings.json.SIDECAR_ALLOW_ROOTSdefaults to cwd. Files outside are skipped with a warning. Set explicitly for stricter scoping.
Development
git clone https://github.com/dEMonaRE/sidecar-mcp.git
cd sidecar-mcp
pnpm install
pnpm dev # run with tsx watch
pnpm test # vitest
pnpm demo # offline self-check (fake backend, prints prompt + reply)
pnpm build # tsc → dist/
pnpm pack # build a tarball to verify before publishingManual smoke test with real Ollama:
ollama serve &
ollama pull llama3.1:8b
SIDECAR_BACKEND=ollama sidecar-mcp & # in one terminal
# Wire into your MCP client (Claude Code, Copilot, Codex) and call bulk_read.License
MIT — see LICENSE.
Available Tools
1 toolbulk_readA
Read multiple files and ask the worker LLM to summarize them in the context of question. Use instead of Read for 2+ files or any file >100 lines. Returns the worker summary, never raw file content.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | ||
| paths | Yes | ||
| question | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It meaningfully discloses that the tool delegates summarization to a worker LLM and that it returns the worker summary, never raw file content. While it doesn't cover auth or error behavior, the key transformation and output characteristics are clearly stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with zero filler. The core behavior and usage guidance are front-loaded, followed by the return-value caveat. Every sentence earns its place and the description is appropriately sized for the tool's simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema alerts and no sibling list, the description covers the essential aspects: what it does, when to use it, and what it returns. The only notable gap is the undocumented model parameter, but that is optional and less critical, so the definition is still mostly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It gives context for 'paths' (multiple files) and 'question' (summarization context), but does not address the optional 'model' parameter at all. This is a partial but not complete compensation for the schema's lack of descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the action ('Read multiple files'), the resource (files), and the specific behavior (asking the worker LLM to summarize them in the context of the question). It explicitly differentiates itself from the alternative 'Read' tool, so an agent can distinguish it without needing additional context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides concrete, actionable criteria: 'Use instead of `Read` for 2+ files or any file >100 lines.' This is an explicit when-to-use instruction with a named alternative trigger condition, leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.2.1- First observed
bulk_read
TDQS
Scored across 1 tool
With only one tool, there is no possibility of confusion or overlap with other tools. The tool's purpose is clearly described.
A single tool name cannot demonstrate a pattern, but 'bulk_read' is a clear verb_noun construction. The naming is sensible, though consistency cannot be assessed across a set.
A single tool is extremely thin for a server, even if it serves a specific niche. The server's purpose appears to be reading and summarizing files, which likely requires additional supporting tools for full utility.
The tool covers bulk reading and summarization, but there are no complementary tools for file listing, individual reads, or other common operations. The surface is severely limited for a file-access server.
Maintenance
Related MCP Connectors
Shared memory for coding agents. Stop re-explaining your codebase every session.
Codebase intelligence for agents: 152 structured artifacts across 21 programs, one call.
- uploads.shOAuthsh.uploads
Host files from coding agents; stage on a branch and attach to GitHub PRs.
- AxisOAuthdev.useaxis
Coding agents from Claude Code, Cursor and Codex claim jobs and lock files on one shared board.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables Codex to delegate bulk code reading, patching, and testing to an async worker using cheaper AI models, while receiving compact results.161 npm2MIT
- AlicenseNot gradedqualityCmaintenanceEnables coding agents to delegate file reads, command output triage, page fetching, and image inspection to cheap flash models, returning concise answers and verified pointers while keeping raw dumps out of the main model's context.MIT
- AlicenseNot gradedqualityAmaintenanceProvides a bulk_read_files tool that reads specified files and forwards them to a remote backend for summarization, so agents receive concise summaries instead of raw file contents.13 npmMIT
- FlicenseAqualityBmaintenanceEnables a primary agent to offload bounded, high-volume analysis of explicitly selected workspace files to an auxiliary LLM through a single read-only tool, returning text, Markdown, JSON, or unified-diff output with token usage while retaining all planning, verification, and edits.1-