Skip to main content
Glama

ConferLLM

Python 3.10+ License: MIT MCP Ready uv

ConferLLM is a minimalist agent harness with multi-model support and essential host tools. It enables AI agents (such as Claude Code, Codex, Cursor, and Cline) and human developers to consult, delegate, and collaborate across models on complex tasks with full session state and local execution capabilities.

┌──────────────────────────────────────────────────────────────────────────────┐
│  AI Agent / Developer (Claude Code, Codex, Cursor, CLI, Script)             │
└──────────────────────────────────────┬───────────────────────────────────────┘
                                       │ CLI (--json) or MCP (stdio/sse/http)
                                       ▼
┌──────────────────────────────────────────────────────────────────────────────┐
│  ConferLLM Engine                                                            │
│  • Stateful Session Store (atomic JSONL + fcntl locking)                     │
│  • Self-Correcting Tool Runtime (up to 1,000 tool rounds per turn)           │
│  • Multimodal Pipeline (artifact tracking + auto-saving)                     │
└──────────────┬───────────────────────────────────────────────┬───────────────┘
               │ Native Tool Calls                             │ LiteLLM /
               ▼                                               │ Responses API
┌──────────────────────────────┐                ┌──────────────▼───────────────┐
│ Host Tools (Unrestricted)    │                │ Configured AI Models         │
│ • Shell: bash, pwsh, git     │                │ • OpenAI GPT-6 Astra         │
│ • Background: output, kill   │                │ • Claude Opus/Sonnet         │
│ • Files: read, write, edit,  │                │ • DeepSeek Flash / Pro       │
│   append, list_directory     │                │ • Google Gemini              │
│                              │                │ • Local Ollama               │
└──────────────────────────────┘                └──────────────────────────────┘

Highlights

  • Multi-Model Delegation: Connect to 100+ providers via LiteLLM (OpenAI, Anthropic, Gemini, DeepSeek, local Ollama, Azure, Bedrock, etc.) using clean local aliases.

  • 10 Native Built-in Host Tools: Delegated models get access to shell (run_shell, run_powershell, shell_output, shell_kill, git_command) and filesystem tools (read_file, write_file, append_file, edit_file, list_directory) without extra flags.

  • Self-Correcting Tool Loop: Tool execution errors are returned as native results so the model can inspect errors and self-correct across up to 1,000 rounds in a single turn.

  • Persistent Multi-Turn Sessions: Conversations are stored as atomic JSONL files with process-safe locking. Resume any session at any time with --session <id>.

  • First-Class Multimodal Support: Send input images (--image); model-generated images are automatically saved to disk, linked via URIs, and safely replayed in history.

  • Agent-Ready Surfaces: Dual interface out of the box—bundled Agent Skill (conferllm skill install) for Claude Code/Codex and standard MCP Server (conferllm serve) for Claude Desktop/Cursor/Cline.


Related MCP server: MCP AI Gateway

Quick Start

Requires Python 3.10+ on macOS or Linux (POSIX file locking required; native Windows is not supported).

1. Install

Install with uv (recommended):

uv tool install conferllm

Or install with pip in an activated virtual environment:

pip install conferllm

Verify the installation:

conferllm --version
conferllm doctor --json

Note: If the command is not found after uv tool install, run uv tool update-shell and restart your terminal.

2. Configure Models

Create the configuration directory and file:

mkdir -p ~/.conferllm
chmod 700 ~/.conferllm
touch ~/.conferllm/config.yaml
chmod 600 ~/.conferllm/config.yaml

Populate ~/.conferllm/config.yaml with your preferred providers. For example:

model_list:
  - model_name: gpt-6-astra
    api_format: responses
    capabilities:
      input_modalities: [text, image]
      output_modalities: [text]
    litellm_params:
      model: openai/gpt-6-astra
      api_key: "your-openai-key"

  - model_name: claude-opus-5
    capabilities:
      input_modalities: [text, image]
      output_modalities: [text]
    litellm_params:
      model: anthropic/claude-opus-5
      api_key: "your-anthropic-key"

  - model_name: DeepSeek-V4.1-Flash
    capabilities:
      input_modalities: [text, image]
      output_modalities: [text]
    litellm_params:
      model: deepseek/deepseek-flash
      api_key: "your-deepseek-key"

  - model_name: gemini-3.8-flash
    capabilities:
      input_modalities: [text, image]
      output_modalities: [text]
    litellm_params:
      model: gemini/gemini-3.8-flash
      api_key: "your-gemini-key"

  - model_name: local-qwen3
    litellm_params:
      model: ollama_chat/qwen3-coder:30b
      api_base: "http://localhost:11434"

Tip: See config_example.yaml for more provider templates including Gemini 3.8 Flash, Qwen 3 Coder, Mistral Large, Together AI, Hugging Face, Azure, and AWS Bedrock.

3. Run Your First Chat

conferllm chat --model gpt-6-astra --prompt "Explain Raft leader election in two sentences."

Output includes the answer and a reusable Session ID:

Session: 20260905-0123456789abcdef0123456789abcdef
Name: Explain Raft leader election in two sentences.

Raft leader election ensures a cluster chooses a single leader through randomized election timeouts and majority voting. Nodes transition to candidates, request votes, and become leaders upon securing a majority.

AI Agent Integration

ConferLLM is designed from the ground up for agent consumption. AI agents use ConferLLM to obtain attributed second opinions, delegate specialized tasks (e.g., deep math proofs to DeepSeek V4.1 Flash, visual analysis to GPT-6 Astra, or private offline tasks to local Llama 4), and let external models interact with the host repository.

Option A: Use as an Agent Skill (Claude Code & Codex)

Install the bundled skill into your agent's skill directory:

# For Claude Code and general agents (~/.agents/skills/conferllm):
conferllm skill install

# For Codex (~/.codex/skills/conferllm):
conferllm skill install --target codex

Once installed, your agent automatically knows how to discover configured models, run queries with --json, parse outputs, and continue conversations.

Example agent instruction:

"Use $conferllm to ask DeepSeek-V4.1-Flash to review our Raft consensus implementation in src/consensus.py and check for election split-vote edge cases."

Option B: Use as an MCP Server (Claude Desktop, Cursor, Cline, Windsurf)

ConferLLM includes a full-featured MCP (Model Context Protocol) server over stdio, SSE, or streamable HTTP.

Add to your MCP configuration (e.g. claude_desktop_config.json or Cursor Settings):

{
  "mcpServers": {
    "conferllm": {
      "command": "conferllm",
      "args": ["serve"]
    }
  }
}

MCP Tools Provided

MCP Tool

Description

create_chat(message, model, name=None, images=None, ...)

Start a new conversation with host shell and file tools enabled.

continue_chat(message, session_id, images=None, ...)

Continue an existing session with conversation history restored.

list_sessions(query=None, model=None, since=None, until=None, limit=50)

Search past sessions by ID, name, model, or creation date.

list_models()

Return all configured model aliases.

get_model_info(model)

Return declared capabilities, modalities, and format details.

Generated images are accessible via the MCP resource conferllm://sessions/{session_id}/artifacts/{artifact_id}.


Common Workflows

1. Multi-Turn Conversation (Follow-Up)

Pass --session with the returned session ID to continue with full context:

conferllm chat --session 20260905-0123456789abcdef0123456789abcdef --prompt "Now compare it with Paxos."

Note: A session retains its original model alias. History and past tool calls are restored automatically without re-executing old commands.

2. Machine-Readable JSON Mode

AI agents and scripts should always pass --json for predictable parsing:

conferllm chat --model gpt-6-astra --prompt "Explain Raft." --json

JSON Response Contract (conferllm.chat.response.v1):

{
  "schema_version": "conferllm.chat.response.v1",
  "ok": true,
  "session": {
    "id": "20260905-0123456789abcdef0123456789abcdef",
    "name": "Explain Raft.",
    "model": "gpt-6-astra",
    "turn": 1
  },
  "message": {
    "text": "Raft is a consensus algorithm...",
    "content": [
      {
        "type": "text",
        "text": "Raft is a consensus algorithm..."
      }
    ]
  },
  "artifacts": [],
  "warnings": []
}

Key fields:

  • session.id: Unique session identifier to pass to --session on follow-up.

  • session.turn: Turn counter (starts at 1).

  • message.text: The complete assistant answer (with image payloads replaced by saved paths).

  • artifacts: List of turn artifacts (images input/output, paths, mime types, hashes).

  • warnings: Non-fatal notices (e.g. background processes stopped at turn end).

3. Long Prompts from Files

For complex multi-line prompts, markdown instructions, or code snippets:

conferllm chat --model DeepSeek-V4.1-Flash --prompt-file ./review_prompt.md --json

4. Multimodal Analysis (Images)

Attach one or more images using repeated --image flags:

conferllm chat \
  --model gpt-6-astra \
  --prompt "Analyze the architectural bottleneck shown in this diagram." \
  --image ./architecture.png \
  --json
  • Images are validated against model capabilities and copied into session-owned storage for safe replay across turns.

  • Model-generated images are automatically written to --image-output-dir (defaults to /tmp) and listed in artifacts with direction: "output".

5. Local File and Shell Operations

Delegated models have native access to host files and commands without extra flags:

conferllm chat \
  --model gpt-6-astra \
  --prompt "Read pyproject.toml and tell me what dependencies need attention." \
  --json

Safety note: If you want an opinion or review only without risking file modifications, include in your prompt: "Provide an analysis only; do not edit files or execute modifying commands."

6. Search and Inspect Past Sessions

# List recent sessions
conferllm sessions list

# Search by keyword in name or ID
conferllm sessions list --query raft --json

# Filter by model alias and date range
conferllm sessions list --model gpt-6-astra --since 2026-09-01 --limit 10 --json

7. Compare Models Side-by-Side

To compare how different models handle the same challenge, start independent chats:

conferllm chat --model DeepSeek-V4.1-Flash --prompt-file ./challenge.md --json
conferllm chat --model claude-opus-5 --prompt-file ./challenge.md --json

Built-in Host Tools

ConferLLM provides every configured model with 10 native tools. Native tool calls execute on the host machine with the user's OS permissions:

Tool

Purpose

Key Parameters

run_shell

Run bash/sh commands

command, timeout (1–600s, def 120s), cwd, run_in_background

run_powershell

Run PowerShell commands

command, timeout, cwd, run_in_background

shell_output

Read unread output from a background command

shell_id, filter_str (optional regex)

shell_kill

Terminate an owned background process group

shell_id

git_command

Execute git commands

command, timeout (1–600s, def 60s), cwd

read_file

Read UTF-8 file with 1-indexed line numbers

path, offset (def 1), limit (def 2000 lines)

write_file

Create or atomically overwrite a file

path, content (up to 10 MiB)

append_file

Append content to a file

path, content

edit_file

Exact unique string replacement

path, old_string, new_string, replace_all (def false)

list_directory

Directory listing with glob exclusions

path (def .), ignore (patterns), limit (def 1000)

Tool Execution Rules

  • No confirmation gates: Operations execute immediately to enable seamless multi-step autonomous problem solving.

  • Self-correction: Missing files, invalid arguments, and command errors return as structured tool results so the model can rectify errors itself.

  • Turn boundaries: Background processes belong to the turn and are cleanly stopped at turn completion. Up to 1,000 tool rounds are permitted per turn.


Diagnostics & Troubleshooting

Run doctor to inspect system status, configuration, models, and permissions without exposing credentials:

conferllm doctor --json

List available model aliases:

conferllm models --json

Inspect non-secret details of a specific model alias:

conferllm model-info gpt-6-astra

Common Issues

  • command not found: conferllm: Run uv tool update-shell and restart terminal, or check your virtualenv PATH.

  • model_not_found: Run conferllm models to see available aliases. Model aliases in config.yaml are the identifiers used on the CLI, not raw provider names.

  • configuration_error: Verify YAML syntax in ~/.conferllm/config.yaml. Permissions should be 0700 for directory and 0600 for the configuration file.

  • Provider Responses API vs Chat Completions: OpenAI models use Responses by default. For legacy OpenAI-compatible endpoints that only support /chat/completions, add api_format: chat_completion to the model config. GPT-6 Astra tool calling requires Responses.


Development

# Clone repository
git clone https://github.com/feiskyer/mcp-ai-hub.git conferllm
cd conferllm

# Setup development environment
uv sync --extra dev

# Run quality checks & test suite
uv run ruff format --check .
uv run ruff check .
uv run mypy src/
uv run pytest

# Build packages
uv build

To link an editable checkout globally to your tool environment:

uv tool install --editable . --force

License

ConferLLM is open source software released under the MIT License.

Available Tools

3 tools
chatB

Chat with specified AI model.

    Args:
        model: Model name from configuration (e.g., 'gpt-4', 'claude-sonnet-4')
        inputs: Chat input (string or OpenAI-format messages)

    Returns:
        AI model response as string
    
ParametersJSON Schema
NameRequiredDescriptionDefault
inputsYes
modelYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. While it mentions the tool chats with AI models, it doesn't describe important behavioral aspects like rate limits, authentication requirements, cost implications, error handling, or whether this is a read-only vs. state-changing operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is perfectly structured and concise - a clear purpose statement followed by well-organized parameter explanations and return value description. Every sentence earns its place with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (interactive AI chat) and the presence of an output schema (which handles return values), the description is adequate but incomplete. It covers parameters well but lacks behavioral context that would be crucial for an AI agent to use this tool appropriately.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds significant value beyond the 0% schema description coverage by explaining both parameters: 'model' is described as 'Model name from configuration' with examples, and 'inputs' is clarified as 'Chat input (string or OpenAI-format messages)'. This compensates well for the schema's lack of descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose as 'Chat with specified AI model', which is a specific verb+resource combination. However, it doesn't explicitly differentiate from sibling tools like 'get_model_info' or 'list_models', which appear to be informational rather than interactive.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. There's no mention of when to choose 'chat' over the sibling tools 'get_model_info' or 'list_models', nor any context about appropriate use cases or limitations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_model_infoC

Get information about a specific model.

    Args:
        model: Model name to get info for

    Returns:
        Dictionary with model information
    
ParametersJSON Schema
NameRequiredDescriptionDefault
modelYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states it 'Get information' but doesn't clarify if this is a read-only operation, what happens if the model doesn't exist (e.g., error handling), or any rate limits or permissions required. This leaves significant gaps in understanding the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded, starting with the main purpose. The Args and Returns sections are structured clearly, but the formatting with indentation might be slightly verbose. Overall, it's efficient with little waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's low complexity (one parameter) and the presence of an output schema, the description is somewhat complete. However, it lacks behavioral context and usage guidelines, which are important for an AI agent to invoke it correctly. The output schema helps, but the description could do more to explain the tool's role relative to siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds minimal semantics beyond the input schema, which has 0% description coverage. It specifies that the 'model' parameter is the 'Model name to get info for', but this is basic and doesn't provide details like format, examples, or constraints. With one parameter and low schema coverage, the description compensates slightly but not fully.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('Get') and resource ('information about a specific model'), making it easy to understand what it does. However, it doesn't explicitly differentiate from sibling tools like 'list_models', which might list multiple models rather than get detailed info about one.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The description lacks context about prerequisites, such as whether the model must exist or be accessible, and doesn't mention sibling tools like 'list_models' for comparison or 'chat' for different operations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_modelsB

List all available AI models.

    Returns:
        List of available model names
    
ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. While it states the tool lists models and returns a list of names, it doesn't describe important behavioral aspects like whether this is a read-only operation, if there are rate limits, authentication requirements, or how the list is structured (e.g., pagination, sorting). The description is minimal and lacks behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise with two sentences: one stating the purpose and one describing the return value. It's front-loaded with the main action. However, the formatting with indentation and a 'Returns:' section is slightly verbose for such a simple tool, but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 0 parameters, 100% schema coverage, and an output schema exists, the description is adequate but minimal. It covers the basic purpose and return type, but lacks context on usage guidelines and behavioral traits. For a simple list tool, this might be sufficient, but it could benefit from more guidance on when to use it versus siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has 0 parameters, and schema description coverage is 100% (empty schema). The description doesn't need to explain any parameters, which is appropriate. It focuses on the return value instead, which adds value beyond the input schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('List') and resource ('all available AI models'). It distinguishes from 'get_model_info' by focusing on listing all models rather than getting detailed information about a specific one. However, it doesn't explicitly differentiate from 'chat' beyond the obvious functional difference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'get_model_info' or 'chat'. It doesn't mention any prerequisites, context, or exclusions for usage. The agent must infer usage from the tool name and description alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updates
    • First observedchat
    • First observedget_model_info
    • First observedlist_models

TDQS

B3.3/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: chat for interacting with models, get_model_info for retrieving metadata about a specific model, and list_models for enumerating available models. There is no overlap or ambiguity between these functions.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern (chat, get_model_info, list_models) with clear, descriptive verbs that align with their actions. No deviations or mixed conventions are present.

Tool Count3/5

With only 3 tools, the set feels thin for an AI hub server, potentially lacking operations like model configuration, session management, or advanced interactions. However, it covers basic listing, info retrieval, and chat functionality.

Completeness3/5

The tools provide core chat and model listing capabilities but have notable gaps for a full AI hub domain, such as updating model settings, managing conversations, or handling multimodal inputs. Agents can work around this but may encounter limitations.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    D
    maintenance
    An MCP server that enables AI applications to access 20+ model providers (including OpenAI, Anthropic, Google) through a unified interface for text and image generation.
    2
    30
    MIT
  • A
    license
    B
    quality
    F
    maintenance
    Enables AI assistants to intelligently select and switch between different AI models (OpenAI, Anthropic, etc.) within the same conversation based on task requirements. Provides a unified interface for accessing multiple AI providers through a single MCP tool.
    1
    16 npm
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides comprehensive AI model metadata through MCP, enabling search and filtering of 100+ AI models by capabilities, pricing, context length, and provider specifications.
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    Provides a unified MCP interface for running completions, embeddings, image generation, and classification across OpenAI, Anthropic, Groq, and Mistral. Eliminates provider-specific boilerplate by standardizing API calls for text generation, vector embeddings, and classification tasks.
    -