Skip to main content
Glama

mcp-sage

Ein MCP-Server (Model Context Protocol), der Tools zum Senden von Eingabeaufforderungen an OpenAIs O3-Modell oder Googles Gemini 2.5 Pro basierend auf der Token-Anzahl bereitstellt. Die Tools betten alle referenzierten Dateipfade (rekursiv für Ordner) in die Eingabeaufforderung ein. Dies ist nützlich, um Zweitmeinungen oder detaillierte Codeüberprüfungen von einem Modell einzuholen, das große Mengen Kontext präzise verarbeiten kann.

Begründung

Ich nutze Claude Code intensiv. Es ist ein großartiges Produkt, das sich gut für meinen Workflow eignet. Neuere Modelle mit großem Kontextumfang scheinen jedoch besonders nützlich für komplexere Codebasen zu sein, die mehr Kontext benötigen. Dadurch kann ich Claude Code weiterhin als Entwicklungstool verwenden und gleichzeitig die umfangreichen Kontextfunktionen von O3 und Gemini 2.5 Pro nutzen, um den begrenzten Kontext von Claude Code zu erweitern.

Related MCP server: Claude Code Review MCP

Modellauswahl

Der Server wählt automatisch das entsprechende Modell basierend auf der Token-Anzahl und den verfügbaren API-Schlüsseln aus:

  • Für kleinere Kontexte (≤ 200.000 Token): Verwendet das O3-Modell von OpenAI (wenn OPENAI_API_KEY festgelegt ist)

  • Für größere Kontexte (> 200.000 und ≤ 1 Mio. Token): Verwendet Google Gemini 2.5 Pro (wenn GEMINI_API_KEY festgelegt ist)

  • Wenn der Inhalt 1 Million Token überschreitet: Gibt einen informativen Fehler zurück

Fallback-Verhalten:

  • API-Schlüssel-Fallback :

    • Wenn OPENAI_API_KEY fehlt, wird Gemini für alle Kontexte innerhalb seines 1M-Token-Limits verwendet

    • Wenn GEMINI_API_KEY fehlt, können nur kleinere Kontexte (≤ 200K Tokens) mit O3 verarbeitet werden

    • Wenn beide API-Schlüssel fehlen, wird ein informativer Fehler zurückgegeben

  • Fallback auf Netzwerkkonnektivität :

    • Wenn die OpenAI-API nicht erreichbar ist (Netzwerkfehler), greift das System automatisch auf Gemini zurück

    • Dies bietet Widerstandsfähigkeit gegen vorübergehende Netzwerkprobleme bei einem Anbieter

    • Erfordert, dass GEMINI_API_KEY festgelegt ist, damit der Fallback funktioniert

Inspiration

Dieses Projekt ist von zwei anderen Open-Source-Projekten inspiriert:

Überblick

Dieses Projekt implementiert einen MCP-Server, der drei Tools bereitstellt:

sage-opinion

  1. Nimmt eine Eingabeaufforderung und eine Liste von Datei-/Verzeichnispfaden als Eingabe entgegen

  2. Packt die Dateien in ein strukturiertes XML-Format

  3. Misst die Token-Anzahl und wählt das entsprechende Modell aus:

    • O3 für ≤ 200.000 Token

    • Gemini 2.5 Pro für > 200.000 und ≤ 1 Mio. Token

  4. Sendet die kombinierte Eingabeaufforderung + Kontext an das ausgewählte Modell

  5. Gibt die Antwort des Modells zurück

sage-review

  1. Nimmt eine Anweisung für Codeänderungen und eine Liste von Datei-/Verzeichnispfaden als Eingabe entgegen

  2. Packt die Dateien in ein strukturiertes XML-Format

  3. Misst die Token-Anzahl und wählt das entsprechende Modell aus:

    • O3 für ≤ 200.000 Token

    • Gemini 2.5 Pro für > 200.000 und ≤ 1 Mio. Token

  4. Erstellt eine spezielle Eingabeaufforderung, die das Modell anweist, Antworten mithilfe von SEARCH/REPLACE-Blöcken zu formatieren

  5. Sendet den kombinierten Kontext + die Anweisung an das ausgewählte Modell

  6. Gibt Bearbeitungsvorschläge zurück, die als Such-/Ersetzungsblöcke formatiert sind, um die Implementierung zu vereinfachen

sage-plan

  1. Nimmt eine Eingabeaufforderung entgegen, in der ein Implementierungsplan und eine Liste von Datei-/Verzeichnispfaden als Eingabe angefordert werden

  2. Packt die Dateien in ein strukturiertes XML-Format

  3. Orchestriert eine Multi-Modell-Debatte, um einen hochwertigen Implementierungsplan zu erstellen

  4. Die Modelle kritisieren und verfeinern die Pläne der anderen in mehreren Runden

  5. Gibt den erfolgreichen Implementierungsplan mit detaillierten Schritten zurück

sage-plan - Multi-Modell- und Selbstdebatten-Workflows

Das Tool sage-plan fragt nicht ein einzelnes Modell nach einem Plan. Stattdessen orchestriert es eine strukturierte Debatte , die über eine oder mehrere Runden läuft, und bittet dann ein separates Richtermodell (oder dasselbe Modell im CoRT-Modus), den Gewinner auszuwählen.


1. Multi-Modell-Debattenfluss

flowchart TD
  S0[Start Debate] -->|determine models, judge, budgets| R1

  subgraph R1["Round 1"]
    direction TB
    R1GEN["Generation Phase<br/>*ALL models run in parallel*"]
    R1GEN --> R1CRIT["Critique Phase<br/>*ALL models critique others in parallel*"]
  end

  subgraph RN["Rounds 2 to N"]
    direction TB
    SYNTH["Synthesis Phase<br/>*every model refines own plan*"]
    SYNTH --> CONS[Consensus Check]
    CONS -->|Consensus reached| JUDGE
    CONS -->|No consensus & round < N| CRIT["Critique Phase<br/>*models critique in parallel*"]
    CRIT --> SYNTH
  end

  R1 --> RN
  JUDGE[Judgment Phase<br/>*judge model selects/merges plan*]
  JUDGE --> FP[Final Plan]

  classDef round fill:#e2eafe,stroke:#4169E1;
  class R1GEN,R1CRIT,SYNTH,CRIT round;
  style FP fill:#D0F0D7,stroke:#2F855A,stroke-width:2px
  style JUDGE fill:#E8E8FF,stroke:#555,stroke-width:1px

Wichtige Phasen der Multimodell-Debatte:

Einrichtungsphase

  • Das System ermittelt die verfügbaren Modelle, wählt einen Richter aus und weist Token-Budgets zu

Runde 1

  • Generierungsphase - Jedes verfügbare Modell (A, B, C usw.) schreibt parallel seinen eigenen Implementierungsplan

  • Kritikphase - Jedes Modell überprüft alle anderen Pläne (niemals seine eigenen) und erstellt parallel dazu strukturierte Kritiken

Runden 2 bis N (N ist standardmäßig 3)

  1. Synthesephase – Jedes Modell verbessert seinen vorherigen Plan anhand der erhaltenen Kritik (Modelle arbeiten parallel)

  2. Konsensprüfung - Das Richtermodell bewertet die Ähnlichkeit zwischen allen aktuellen Plänen

    • Bei einem Ergebnis ≥ 0,9 wird die Debatte vorzeitig beendet und es geht weiter zum Urteil.

  3. Kritikphase - Wenn kein Konsens erreicht wird UND wir nicht in der Endrunde sind, kritisiert jedes Modell alle anderen Pläne erneut (parallel).

Urteilsphase

  • Nach Abschluss aller Runden (oder Erreichen eines frühen Konsenses) gilt für das Richtermodell (standardmäßig O3):

    • Wählt den besten Plan aus ODER führt mehrere Pläne zu einem besseren zusammen

    • Bietet einen Vertrauenswert für die Auswahl/Synthese


2. Selbstdebattenfluss – Einzelmodell verfügbar

flowchart TD
  SD0[Start Self-Debate] --> R1

  subgraph R1["Round 1 - Initial Plans"]
    direction TB
    P1[Generate Plan 1] --> P2[Generate Plan 2<br/>*different approach*]
    P2 --> P3[Generate Plan 3<br/>*different approach*]
  end

  subgraph RN["Rounds 2 to N"]
    direction TB
    REF[Generate Improved Plan<br/>*addresses weaknesses in all previous plans*]
    DEC{More rounds left?}
    REF --> DEC
    DEC -->|Yes| REF
  end

  R1 --> RN
  DEC -->|No| FP[Final Plan = last plan generated]

  style FP fill:#D0F0D7,stroke:#2F855A,stroke-width:2px

Wenn nur ein Modell verfügbar ist, wird ein Chain of Recursive Thoughts (CoRT) -Ansatz verwendet:

  1. Initial Burst - Das Modell generiert drei verschiedene Pläne, die jeweils einen anderen Ansatz verfolgen

  2. Verfeinerungsrunden – Für jede nachfolgende Runde (2 bis N, Standard N=3):

    • Das Modell überprüft alle bisherigen Pläne

    • Es kritisiert sie intern und identifiziert Stärken und Schwächen

    • Es entsteht ein neuer, verbesserter Plan, der die Einschränkungen früherer Pläne berücksichtigt.

  3. Endgültige Auswahl – Der zuletzt erstellte Plan wird zum endgültigen Implementierungsplan


Was tatsächlich im Code passiert (Kurzreferenz)

Phase / Funktionalität

Code-Speicherort

Hinweise

Eingabeaufforderungen zur Generierung

prompts/debatePrompts.generatePrompt

Fügt die Überschrift „# Implementierungsplan (Modell X)“ hinzu.

Kritikanregungen

prompts/debatePrompts.critiquePrompt

Verwendet die Abschnitte „## Kritik an Plan {ID}“

Syntheseaufforderungen

prompts/debatePrompts.synthesizePrompt

Model überarbeitet eigenen Plan

Konsensprüfung

debateOrchestrator.checkConsensus

Judge-Modell gibt JSON mit consensusScore zurück

Urteil

prompts/debatePrompts.judgePrompt

Richter gibt „# Endgültiger Implementierungsplan“ + Vertrauen zurück

Aufforderung zur Selbstdebatte

prompts/debatePrompts.selfDebatePrompt

Kette-von-Rekursiven-Gedanken- Schleife

Leistungs- und Kostenüberlegungen

⚠️ Wichtig: Das Sage-Plan-Tool kann:

  • Die Fertigstellung nimmt viel Zeit in Anspruch (5–10 Minuten bei mehreren Modellen).

  • Verbrauchen Sie erhebliche API-Token aufgrund mehrerer Diskussionsrunden

  • Verursachen höhere Kosten als Einzelmodellansätze

Typische Ressourcennutzung:

  • Multi-Modell-Debatte: 2-4x mehr Token als ein Single-Modell-Ansatz

  • Bearbeitungszeit: 5-10 Minuten, abhängig von der Komplexität und der Modellverfügbarkeit

  • API-Kosten: 0,30–1,50 $ pro Planerstellung (variiert je nach verwendeten Modellen und Plankomplexität)

Voraussetzungen

  • Node.js (v18 oder höher)

  • Ein Google Gemini API-Schlüssel (für größere Kontexte)

  • Ein OpenAI-API-Schlüssel (für kleinere Kontexte)

Installation

# Clone the repository
git clone https://github.com/your-username/mcp-sage.git
cd mcp-sage

# Install dependencies
npm install

# Build the project
npm run build

Umgebungsvariablen

Legen Sie die folgenden Umgebungsvariablen fest:

  • OPENAI_API_KEY : Ihr OpenAI-API-Schlüssel (für O3-Modell)

  • GEMINI_API_KEY : Ihr Google Gemini API-Schlüssel (für Gemini 2.5 Pro)

Verwendung

Fügen Sie Ihrer MCP-Konfiguration nach dem Erstellen mit npm run build Folgendes hinzu:

OPENAI_API_KEY=your_openai_key GEMINI_API_KEY=your_gemini_key node /path/to/this/repo/dist/index.js

Sie können auch an anderer Stelle festgelegte Umgebungsvariablen verwenden, beispielsweise in Ihrem Shell-Profil.

Eingabeaufforderung

Um eine zweite Meinung zu etwas zu bekommen, fragen Sie einfach nach einer zweiten Meinung.

Um eine Codeüberprüfung zu erhalten, fordern Sie eine Codeüberprüfung oder eine Expertenüberprüfung an.

Beide profitieren von der Bereitstellung von Pfaden zu Dateien, die Sie in den Kontext aufnehmen möchten. Wenn diese jedoch weggelassen werden, leitet das Host-LLM wahrscheinlich ab, was aufgenommen werden soll.

Debuggen und Überwachen

Der Server stellt detaillierte Überwachungsinformationen über die MCP-Protokollierungsfunktion bereit. Diese Protokolle umfassen:

  • Token-Nutzungsstatistiken und Modellauswahl

  • Anzahl der in der Anfrage enthaltenen Dateien und Dokumente

  • Metriken zur Anforderungsverarbeitungszeit

  • Fehlerinformationen bei Überschreitung des Token-Limits

Protokolle werden über die notifications/message des MCP-Protokolls gesendet, um sicherzustellen, dass sie die JSON-RPC-Kommunikation nicht beeinträchtigen. MCP-Clients mit Protokollierungsunterstützung zeigen diese Protokolle entsprechend an.

Beispielprotokolleinträge:

Token usage: 1,234 tokens. Selected model: o3-2025-04-16 (limit: 200,000 tokens)
Files included: 3, Document count: 3
Sending request to OpenAI o3-2025-04-16 with 1,234 tokens...
Received response from o3-2025-04-16 in 982ms
Token usage: 235,678 tokens. Selected model: gemini-2.5-pro-preview-03-25 (limit: 1,000,000 tokens)
Files included: 25, Document count: 18
Sending request to Gemini with 235,678 tokens...
Received response from gemini-2.5-pro-preview-03-25 in 3240ms

Verwenden der Tools

Sage-Opinion-Tool

Das Tool sage-opinion akzeptiert die folgenden Parameter:

  • prompt (Zeichenfolge, erforderlich): Die Eingabeaufforderung, die an das ausgewählte Modell gesendet werden soll

  • paths (Array von Zeichenfolgen, erforderlich): Liste der Dateipfade, die als Kontext einbezogen werden sollen

Beispiel für einen MCP-Toolaufruf (mit JSON-RPC 2.0):

{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "sage-opinion",
    "arguments": {
      "prompt": "Explain how this code works",
      "paths": ["path/to/file1.js", "path/to/file2.js"]
    }
  }
}

Sage-Review-Tool

Das Tool sage-review akzeptiert die folgenden Parameter:

  • instruction (Zeichenfolge, erforderlich): Die spezifischen Änderungen oder Verbesserungen, die erforderlich sind

  • paths (Array von Zeichenfolgen, erforderlich): Liste der Dateipfade, die als Kontext einbezogen werden sollen

Beispiel für einen MCP-Toolaufruf (mit JSON-RPC 2.0):

{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "sage-review",
    "arguments": {
      "instruction": "Add error handling to the function",
      "paths": ["path/to/file1.js", "path/to/file2.js"]
    }
  }
}

Die Antwort enthält SEARCH/REPLACE-Blöcke, die Sie zum Implementieren der vorgeschlagenen Änderungen verwenden können:

<<<<<<< SEARCH
function getData() {
  return fetch('/api/data')
    .then(res => res.json());
}
=======
function getData() {
  return fetch('/api/data')
    .then(res => {
      if (!res.ok) {
        throw new Error(`HTTP error! Status: ${res.status}`);
      }
      return res.json();
    })
    .catch(error => {
      console.error('Error fetching data:', error);
      throw error;
    });
}
>>>>>>> REPLACE

Sage-Plan-Tool

Das Tool sage-plan akzeptiert die folgenden Parameter:

  • prompt (Zeichenfolge, erforderlich): Beschreibung, wofür Sie einen Implementierungsplan benötigen

  • paths (Array von Zeichenfolgen, erforderlich): Liste der Dateipfade, die als Kontext einbezogen werden sollen

Beispiel für einen MCP-Toolaufruf (mit JSON-RPC 2.0):

{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "sage-plan",
    "arguments": {
      "prompt": "Create an implementation plan for adding user authentication to this application",
      "paths": ["src/index.js", "src/models/", "src/routes/"]
    }
  }
}

Die Antwort enthält einen detaillierten Implementierungsplan mit:

  1. Allgemeine Architekturübersicht

  2. Konkrete Umsetzungsschritte

  3. Erforderliche Dateiänderungen

  4. Teststrategie

  5. Mögliche Herausforderungen und Abhilfemaßnahmen

Dieser Plan profitiert von der kollektiven Intelligenz mehrerer KI-Modelle (oder einer gründlichen Selbstüberprüfung durch ein einzelnes Modell) und enthält in der Regel robustere, durchdachtere und detailliertere Empfehlungen als ein Single-Pass-Ansatz.

Ausführen der Tests

So testen Sie die Tools:

# Test the sage-opinion tool
OPENAI_API_KEY=your_openai_key GEMINI_API_KEY=your_gemini_key node test/run-test.js

# Test the sage-review tool
OPENAI_API_KEY=your_openai_key GEMINI_API_KEY=your_gemini_key node test/test-expert.js

# Test the sage-plan tool
OPENAI_API_KEY=your_openai_key GEMINI_API_KEY=your_gemini_key node test/run-sage-plan.js

# Test the model selection logic specifically
OPENAI_API_KEY=your_openai_key GEMINI_API_KEY=your_gemini_key node test/test-o3.js

Hinweis : Die Ausführung des Sage-Plan-Tests kann 5–15 Minuten dauern, da er eine Debatte mehrerer Modelle orchestriert.

Projektstruktur

  • src/index.ts : Die Hauptimplementierung des MCP-Servers mit Tooldefinitionen

  • src/pack.ts : Tool zum Packen von Dateien in ein strukturiertes XML-Format

  • src/tokenCounter.ts : Dienstprogramme zum Zählen von Token in einer Eingabeaufforderung

  • src/gemini.ts : Implementierung des Gemini-API-Clients

  • src/openai.ts : OpenAI API-Clientimplementierung für das O3-Modell

  • src/debateOrchestrator.ts : Multi-Modell-Debattenorchestrierung für Sage-Plan

  • src/prompts/debatePrompts.ts : Vorlagen für Debattenanregungen und -anweisungen

  • test/run-test.js : Test für das Sage-Opinion-Tool

  • test/test-expert.js : Test für das Sage-Review-Tool

  • test/run-sage-plan.js : Test für das Sage-Plan-Tool

  • test/test-o3.js : Test für die Modellauswahllogik

Lizenz

ISC

Available Tools

3 tools
sage-opinionA

Send a prompt to sage-like model for its opinion on a matter.

Include the paths to all relevant files and/or directories that are pertinent to the matter.

IMPORTANT: All paths must be absolute paths (e.g., /home/user/project/src), not relative paths.

Do not worry about context limits; feel free to include as much as you think is relevant. If you include too much it will error and tell you, and then you can include less. Err on the side of including more context.
ParametersJSON Schema
NameRequiredDescriptionDefault
pathsYesPaths to include as context. MUST be absolute paths (e.g., /home/user/project/src). Including directories will include all files contained within recursively.
promptYesThe prompt to send to the external model.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses key behavioral traits: the tool sends a prompt to an external model, handles file paths as context, uses absolute paths, and may error if too much context is included. However, it lacks details on rate limits, authentication needs, or what the 'sage-like model' entails (e.g., model type, limitations). The description doesn't contradict annotations since none exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded, with the core purpose stated first. It uses bullet-like formatting for key points (paths, absolute paths, context limits), but includes some redundancy (e.g., repeating absolute path requirement). Most sentences earn their place by clarifying usage, though it could be slightly more streamlined.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (2 parameters, no output schema, no annotations), the description is somewhat complete but has gaps. It covers the basic operation and constraints, but lacks details on the model's behavior, error handling specifics, or output expectations. Without annotations or an output schema, more context on what 'opinion' entails would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters ('paths' and 'prompt') with descriptions. The description adds minimal value beyond the schema: it reiterates the need for absolute paths and context inclusion but doesn't provide additional syntax, format details, or examples. Baseline 3 is appropriate as the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Send a prompt to sage-like model for its opinion on a matter.' It specifies the verb ('send'), resource ('sage-like model'), and action ('for its opinion'). However, it doesn't explicitly differentiate from sibling tools like 'sage-plan' or 'sage-review' beyond the 'opinion' focus, which is implied but not contrasted.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear usage context: 'Include the paths to all relevant files and/or directories that are pertinent to the matter' and advises on absolute paths and context limits. It implicitly suggests using this tool for opinion-seeking tasks, but it doesn't explicitly state when to choose this over siblings like 'sage-plan' or 'sage-review', nor does it list exclusions or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sage-planA

Generate an implementation plan via multi-model debate.

This tool leverages multiple AI models to debate, critique, and refine implementation plans.

Models will generate initial plans, critique each other's work, refine their plans based on critiques,
and finally produce a consensus plan that combines the best ideas.

IMPORTANT: All paths must be absolute paths (e.g., /home/user/project/src), not relative paths.

The process creates detailed, well-thought-out implementation plans that benefit from
diverse model perspectives and iterative refinement.

When the optional outputPath parameter is provided, the final plan will be saved to that file path,
and a complete transcript of the debate will be saved to a companion file with "-full-transcript"
added to the filename. This is strongly recommended for preserving the expensive results of the debate.
ParametersJSON Schema
NameRequiredDescriptionDefault
maxTokensNoMaximum token budget for the debate
outputPathNoMarkdown file path to save the final plan. Will also save a full transcript to a '-full-transcript.md' suffixed file.
pathsYesPaths to include as context. MUST be absolute paths (e.g., /home/user/project/src). Including directories will include all files contained within recursively.
promptYesThe task to create an implementation plan for
roundsNoNumber of debate rounds (default: 3)

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes key behavioral traits: the multi-model debate process (generation, critique, refinement, consensus), the creation of detailed plans, and file-saving behavior when outputPath is provided. It also notes the expense of the debate, which is useful context. However, it lacks details on error handling or performance expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded, starting with the core purpose. Most sentences add value, such as explaining the debate process and file-saving behavior. However, some redundancy exists (e.g., reiterating absolute paths), and the structure could be slightly tighter by integrating the IMPORTANT note more seamlessly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of a 5-parameter tool with no annotations and no output schema, the description does a good job of covering the tool's behavior and key usage aspects. It explains the debate process and file outputs, but it could be more complete by detailing the format of the output (e.g., Markdown structure) or potential limitations, which would help set clearer expectations for the agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds some value by emphasizing the importance of absolute paths for the 'paths' parameter and explaining the file-saving behavior for 'outputPath', but it does not provide additional semantic context beyond what the schema offers, such as typical use cases for parameters like 'maxTokens' or 'rounds'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Generate an implementation plan via multi-model debate.' It specifies the verb ('generate') and resource ('implementation plan'), and distinguishes it from siblings by detailing the unique multi-model debate process, which is not implied by the sibling names 'sage-opinion' and 'sage-review'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use this tool: for creating detailed, well-thought-out implementation plans through iterative debate. However, it does not explicitly state when not to use it or mention alternatives like the sibling tools, which could help differentiate use cases more precisely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sage-reviewA

Send code to the sage model for expert review and get specific edit suggestions as SEARCH/REPLACE blocks.

Use this tool any time the user asks for a "sage review" or "code review" or "expert review".

This tool includes the full content of all files in the specified paths and instructs the model to return edit suggestions in a specific format with search and replace blocks.

IMPORTANT: All paths must be absolute paths (e.g., /home/user/project/src), not relative paths.

If the user hasn't provided specific paths, use as many paths to files or directories as you're aware of that are useful in the context of the prompt.
ParametersJSON Schema
NameRequiredDescriptionDefault
instructionYesThe specific changes or improvements needed.
pathsYesPaths to include as context. MUST be absolute paths (e.g., /home/user/project/src). Including directories will include all files contained within recursively.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It explains that the tool includes 'full content of all files in the specified paths' and returns 'edit suggestions in a specific format with search and replace blocks', which adds useful context beyond basic functionality. However, it doesn't cover potential limitations like rate limits, authentication needs, or error conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and appropriately sized, with key information front-loaded. However, the second paragraph could be more concise, and the 'IMPORTANT' section repeats path information already stated elsewhere, slightly reducing efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (code review with file processing) and lack of annotations/output schema, the description is moderately complete. It explains the core behavior and format of suggestions but doesn't detail what happens with invalid paths, how large files are handled, or the structure of the returned edit blocks, leaving some gaps for an AI agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters thoroughly. The description reinforces that paths 'must be absolute paths' and mentions directory recursion, but this is already covered in the schema. It adds minimal value beyond what the structured schema provides, meeting the baseline for high coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verbs ('send code', 'get specific edit suggestions') and resources ('sage model', 'SEARCH/REPLACE blocks'). It distinguishes from sibling tools by specifying this is for 'expert review' with edit suggestions, unlike 'sage-opinion' or 'sage-plan' which likely serve different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidelines: 'Use this tool any time the user asks for a "sage review" or "code review" or "expert review"'. It also includes alternative handling when paths aren't specified ('use as many paths... as you're aware of'), giving clear context for when and how to invoke the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.0.0
    • First observedsage-opinion
    • First observedsage-plan
    • First observedsage-review

TDQS

A4/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: sage-opinion provides opinions on matters, sage-plan generates implementation plans through debate, and sage-review offers code review with edit suggestions. There is no overlap in functionality, and the descriptions clearly differentiate their roles.

Naming Consistency5/5

All tool names follow a consistent 'sage-' prefix with a descriptive suffix (opinion, plan, review), using kebab-case throughout. This pattern is predictable and enhances readability, making it easy to identify the tool's function at a glance.

Tool Count4/5

With 3 tools, the count is appropriate for a server focused on AI-assisted development tasks, as it covers key areas like opinion generation, planning, and code review. It is slightly lean but reasonable, as each tool serves a distinct and valuable purpose without redundancy.

Completeness4/5

The tool set covers core AI-assisted development workflows: opinion generation, planning, and code review. Minor gaps exist, such as the lack of tools for executing plans or managing project states, but agents can work around these by combining tools or using external methods.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    A
    quality
    D
    maintenance
    An MCP server that connects Gemini 2.5 Pro to Claude Code, enabling users to generate detailed implementation plans based on their codebase and receive feedback on code changes.
    5
    14
    -
  • A
    license
    A
    quality
    D
    maintenance
    An MCP server that gives your IDE or agent access to Google Gemini with autonomous codebase exploration, enabling deep code analysis, architectural reviews, and bug hunting.
    20
    10
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    An MCP server that integrates Google Gemini CLI with Claude Code for AI-powered development assistance, enabling code review, bug analysis, feature planning, and code explanation without requiring an API key.
    8
    MIT