nexus-convergence-mcp
nexus-convergence-mcp
MCP-сервер, обертывающий Balanced Intelligence Convergence Pipeline.
Предоставляет полный конвейер в виде 4 инструментов MCP, которые можно использовать в любом клиенте, совместимом с MCP (Claude Code, Claude Desktop и т. д.).
Инструменты
converge_query
Выполняет веерный запрос к нескольким LLM, достигает консенсуса и возвращает структурированный результат.
{
"query": "Is this treatment effective?",
"models": ["gpt-4o", "claude-3-5-sonnet-20241022", "gemini-1.5-pro"],
"policy_set": "hipaa"
}Возвращает: run_id, agreement_score, stability, agreed_claims, disputed_claims, final_answer.
get_evidence_ladder
Возвращает полную «лестницу доказательств» (Evidence Ladder) для запроса — неизменяемый аудиторский след каждого этапа конвейера.
{ "query_id": "run-uuid" }Возвращает: упорядоченный список записей от QUERY → DECOMPOSE → MODEL_EXECUTE → CONSENSUS → VERIFY → CONCLUDE.
check_compliance
Проверяет контент на соответствие политикам HIPAA/EU_AI_ACT/NIST/CUSTOM.
{
"content": "Patient SSN is 123-45-6789",
"categories": ["HIPAA"]
}Возвращает: passed, blocked, violations, warnings, logs.
list_model_disagreements
Возвращает информацию о том, где модели разошлись во мнениях — инверсии (прямые противоречия), спорные утверждения, пары с низкой степенью сходства.
{ "query_id": "run-uuid" }Возвращает: inversions, disputed_claims, low_similarity_pairs.
Related MCP server: agent-orchestrator
Конфигурация MCP
Добавьте в свою конфигурацию MCP:
{
"mcpServers": {
"nexus-convergence": {
"command": "npx",
"args": ["-y", "@gonzih/nexus-convergence-mcp"],
"env": {
"CONVERGENCE_SERVICE_URL": "http://your-convergence-service:3000",
"CONSENSUS_SERVICE_URL": "http://your-consensus-service:3001",
"EVIDENCE_SERVICE_URL": "http://your-evidence-service:3002",
"COMPLIANCE_SERVICE_URL": "http://your-compliance-service:3003",
"NEXUS_API_KEY": "your-api-key"
}
}
}
}Окружение
CONVERGENCE_SERVICE_URL=http://localhost:3000
CONSENSUS_SERVICE_URL=http://localhost:3001
EVIDENCE_SERVICE_URL=http://localhost:3002
COMPLIANCE_SERVICE_URL=http://localhost:3003
NEXUS_API_KEY=your-api-keyРазработка
npm ci
npm run build
npm start
npm testAvailable Tools
4 toolscheck_complianceA
Check content against compliance policy sets. Evaluates against HIPAA (PHI detection), EU_AI_ACT (prohibited use cases), NIST (PII/secrets), and CUSTOM rules. Returns: passed/blocked status, list of violations (BLOCK), warnings (WARN), and log entries.
| Name | Required | Description | Default |
|---|---|---|---|
| content | Yes | The content to evaluate against compliance policies | |
| categories | No | Filter to specific policy categories. Leave empty to check against all active policies. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and adequately explains the behavior: evaluating content against policies and returning status, violations, warnings, and logs. It does not mention side effects or permissions, but the read-only nature is implied.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loading the main action then listing specifics. No wasted words – every sentence provides essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema, but the description covers the return fields (passed/blocked, violations, warnings, logs). It is sufficiently complete for a compliance checker, though more detail on interpreting results could help.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers both parameters with descriptions (100% coverage). The description adds value by enumerating the policy categories and detailing the return structure, which goes beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Check content against compliance policy sets' and lists specific policy sets (HIPAA, EU_AI_ACT, NIST, CUSTOM), making it highly specific and distinguishable from its unrelated sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for compliance checks but does not provide explicit guidance on when to use this tool versus alternatives or when not to use it. Since siblings are unrelated, the lack of explicit guidance is a minor gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
converge_queryA
Fan out a query to multiple LLMs in parallel (OpenAI, Claude, Gemini, Ollama), run multi-stage consensus validation, and return a structured result with agreement score, truth stability classification, and provenance. NEVER rely on one model answer. Friction between disagreements = intelligence.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | The query to fan out to multiple models | |
| models | No | List of model names to query. E.g. ["gpt-4o", "claude-3-5-sonnet-20241022", "gemini-1.5-pro"]. Leave empty for defaults. | |
| policy_set | No | Optional policy set identifier for compliance filtering (e.g. "hipaa", "eu-ai-act") |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It describes the parallel fan-out and multi-stage consensus validation, but does not disclose potential costs, rate limits, failure handling (e.g., if a model times out), or the fact that it may be slower than single-model queries. The output format is partially described, but behavioral traits like destructive or read-only are not clarified. The description is adequate but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that efficiently conveys the core purpose, process, and output. It uses imperative language ('NEVER rely') to emphasize proper usage. Every phrase serves a purpose, with no redundancy or filler. It is appropriately sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 3 parameters, no output schema, and no annotations, the description covers the main aspects: what it does, how it works (parallel fan-out, consensus validation), and what it returns (agreement score, truth stability, provenance). It does not explain error behavior or performance trade-offs, but for a tool of this complexity, it is largely complete and leaves minimal ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description adds minimal extra value: for 'models' it notes 'Leave empty for defaults' which is partly in schema (default value), and for 'policy_set' it provides examples ('hipaa', 'eu-ai-act'). However, it does not explain the significance of the default models or the behavior when a model is unavailable. Overall, it adds only slight incremental meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the tool fans out a query to multiple LLMs in parallel, runs multi-stage consensus validation, and returns a structured result with agreement score, truth stability, and provenance. It uses a specific verb ('fan out') and identifies the resource (multiple LLMs). It distinguishes from siblings like check_compliance or list_model_disagreements by focusing on consensus across models.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides strong usage guidance: 'NEVER rely on one model answer' and 'Friction between disagreements = intelligence' explicitly indicate when to use this tool (when multi-model consensus is needed). However, it does not mention when not to use it or suggest alternatives (e.g., if speed is critical, use a single model), which would elevate it to a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_evidence_ladderA
Return the full Evidence Ladder for a query — immutable append-only audit trail showing every step: QUERY → DECOMPOSE → MODEL_EXECUTE → CONSENSUS → VERIFY → CONCLUDE. Each entry records: step, actor (model name or "system"), content, confidence, and timestamp.
| Name | Required | Description | Default |
|---|---|---|---|
| query_id | Yes | The run/query ID returned by converge_query |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description explicitly notes the tool is 'immutable append-only audit trail', indicating no side effects or destructive actions. It also details the exact steps recorded, providing full transparency about the content.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences: first introduces purpose with steps, second details fields. No wasted words. Front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description lists all fields (step, actor, content, confidence, timestamp), so return format is clear. Also covers immutability. For a simple single-param tool, this is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds value by linking query_id to 'converge_query', clarifying its origin. This extra context justifies a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns the full Evidence Ladder for a query, listing the specific steps (QUERY → DECOMPOSE → ...). It uses a specific verb 'Return' and resource 'Evidence Ladder', and distinguishes from siblings by describing the audit trail nature.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when one needs the full audit trail for a query, but does not explicitly state when to use this vs alternatives like list_model_disagreements or check_compliance. However, the purpose is clear enough that an agent would infer appropriate context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_model_disagreementsA
Return where models diverged on a query — direct contradictions (inversions), disputed claims, and low-similarity pairs with their similarity scores. Use this to understand the friction space: inversions are not errors, they are intelligence.
| Name | Required | Description | Default |
|---|---|---|---|
| query_id | Yes | The run/query ID returned by converge_query |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description does not mention authentication, rate limits, or side effects. It adds the behavioral nuance that inversions are not errors, but overall minimal disclosure for a tool with no annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no unnecessary words, front-loaded with core function and then usage guidance. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one parameter and no output schema, the description adequately explains what is returned (inversions, disputed claims, low-similarity pairs with scores) and why to use it. No gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers 100% of the single parameter (query_id) with a complete description. The tool description adds no additional parameter meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the tool returns where models diverged, specifying types of disagreements (inversions, disputed claims, low-similarity pairs) and similarity scores. Distinct from siblings (check_compliance, converge_query, get_evidence_ladder).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use it to understand the friction space, including the insight that inversions are not errors. Lacks explicit alternatives or when-not-to-use, but provides clear usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
check_compliance - First observed
converge_query - First observed
get_evidence_ladder - First observed
list_model_disagreements
TDQS
Scored across 4 tools
Each tool has a clearly distinct purpose: compliance checking, multi-model querying, audit trail retrieval, and disagreement analysis. There is no overlap or ambiguity between them.
All tools follow a consistent verb_noun pattern in snake_case (check_compliance, converge_query, get_evidence_ladder, list_model_disagreements), making naming predictable and easy to understand.
With only 4 tools, the server covers a narrow scope. While each tool is essential, the low count may leave agents wanting for configuration or management tools, but it is still reasonable for a focused server.
The domain of multi-model querying with compliance and audit lacks essential management operations such as listing available models, configuring compliance rules, or retrieving past query history. Significant gaps will cause agent failures.
Maintenance
Related MCP Connectors
AI Reasoning Cache & Consensus Layer with 11 MCP tools via Streamable HTTP.
Free OpenAI-compatible inference with signed provenance receipts and 3 focused MCP tools.
MCP-native AI evaluation: rubric audits, eval suites, and proof reports for AI/LLM output.
100+ MCP tools for AI agents: content metadata, trade intelligence, business-expertise analysis.
Related MCP Servers
- AlicenseAqualityCmaintenanceAn MCP server that enables users to query, compare, and synthesize responses from multiple local and cloud LLMs simultaneously using existing subscriptions. It provides tools for parallel model evaluation, consensus polling with an LLM-as-judge, and response synthesis across different model providers.815 npm16MIT
- AlicenseNot gradedqualityCmaintenanceEnables multi-model leader-worker agent orchestration, workflow execution, and deterministic validation via structured MCP tools.9 npmApache 2.0
- AlicenseNot gradedqualityDmaintenanceExposes RAG and document intelligence pipelines as 8 composable tools for MCP-compatible clients, enabling querying, indexing, classifying, extracting, and assessing documents.1MIT

telos-mcpofficial
AlicenseAqualityBmaintenanceExposes TELOS governance primitives—action scoring, receipt verification, Purpose Anchor inspection, audit-chain queries, and CCRS counterfactual replay—as MCP tools, resources, and prompts for any MCP-compatible client.5Apache 2.0