Councly MCP Server
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Councly MCP Servercreate a council hearing to review this code for security vulnerabilities"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
@councly/mcp
MCP (Model Context Protocol) server for Councly - Multi-LLM Council Hearings.
Enable Claude Code, Codex, and other MCP-compatible AI assistants to invoke council hearings where multiple LLMs (Claude, GPT, Gemini, Grok) debate topics and synthesize verdicts.
Installation
npm install -g @councly/mcpOr use directly with npx:
npx @councly/mcpRelated MCP server: trinity-local
Setup
1. Get an API Key
Sign in to Councly
Go to Settings > MCP Integration
Create a new API key
Copy the key (shown only once)
2. Configure Claude Code
Add to your Claude Code settings (~/.claude/settings.json):
{
"mcpServers": {
"councly": {
"command": "npx",
"args": ["@councly/mcp"],
"env": {
"COUNCLY_API_KEY": "cnc_your_key_here"
}
}
}
}3. Configure Codex CLI
export COUNCLY_API_KEY=cnc_your_key_hereOr add to your shell profile.
Tools
councly_hearing
Create a council hearing where multiple LLMs debate a topic.
Parameters:
Parameter | Type | Required | Default | Description |
subject | string | Yes | - | The topic to discuss (10-10000 chars) |
preset | string | No | balanced | Model preset: |
workflow | string | No | auto | Workflow: |
wait | boolean | No | true | Wait for completion |
timeout_seconds | number | No | 300 | Max wait time (30-600) |
Presets:
Preset | Credits | Counsels | Best For |
balanced | 9 | 3 | General purpose discussions |
fast | 6 | 3 | Quick responses, simple topics |
coding | 14 | 3 | Code review, technical decisions |
coding_plus | 17 | 4 | Complex code problems |
Example:
Use councly_hearing with subject="Review this Python function for security issues:
def authenticate(username, password):
query = f'SELECT * FROM users WHERE username='{username}' AND password='{password}''
return db.execute(query)
" and preset="coding"councly_status
Check the status of a hearing.
Parameters:
Parameter | Type | Required | Description |
hearing_id | string (uuid) | Yes | The hearing ID |
Example:
Use councly_status with hearing_id="550e8400-e29b-41d4-a716-446655440000"Response Format
Completed hearings return:
Status: completed, failed, or early_stopped
Verdict: Synthesized conclusion from the moderator
Trust Score: 0-100 confidence rating
Counsel Perspectives: Summary from each counsel
Cost: Credits used
Error Handling
Common errors:
Code | Description |
INSUFFICIENT_BALANCE | Not enough credits |
ACTIVE_HEARING_EXISTS | One hearing already in progress |
RATE_LIMIT_EXCEEDED | Too many requests |
CONTENT_BLOCKED | Subject contains prohibited content |
Environment Variables
Variable | Required | Description |
COUNCLY_API_KEY | Yes | Your MCP API key |
COUNCLY_BASE_URL | No | API base URL (default: https://councly.ai) |
Pricing
Councly uses a credit-based pricing model:
1 credit = $0.01 USD
Credits are deducted at hearing creation
Failed hearings are refunded
Purchase credits at councly.ai/billing.
Links
License
Apache 2.0 - See LICENSE
Available Tools
2 toolscouncly_hearingA
Create a council hearing where multiple LLMs (Claude, GPT, Gemini, Grok) debate a topic and a moderator synthesizes the verdict.
Use cases:
Code review: Get diverse perspectives on code quality, architecture, security
Technical decisions: Compare approaches, weigh trade-offs
Problem solving: Generate and evaluate multiple solutions
Brainstorming: Explore ideas from different angles
The hearing runs asynchronously. By default, this tool waits for completion and returns the verdict. Set wait=false to get the hearing ID immediately and check status later with councly_status.
Cost: Varies by preset (6-17 credits). Check councly.ai for current pricing.
| Name | Required | Description | Default |
|---|---|---|---|
| subject | Yes | The topic or question to discuss. Be specific and include relevant context. | |
| preset | No | Model preset: balanced (9 credits), fast (6 credits), coding (14 credits), coding_plus (17 credits, 4 counsels) | balanced |
| workflow | No | Workflow type for the hearing | auto |
| wait | No | Wait for completion (true) or return immediately (false) | |
| timeout_seconds | No | Max wait time in seconds (if wait=true) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It effectively describes key behavioral traits: the asynchronous nature of hearings, default waiting behavior, cost implications (6-17 credits), and the ability to check status with another tool. However, it doesn't mention error handling, rate limits, or authentication requirements, which would be helpful for a tool with cost implications.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections: purpose statement, use cases, operational details, and cost information. Each sentence serves a distinct purpose. While slightly longer than minimal, the information density is high with no wasted text. The front-loaded purpose statement immediately communicates the tool's function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 parameters, no annotations, and no output schema, the description provides substantial context about the tool's behavior, use cases, and operational characteristics. It covers the asynchronous nature, cost implications, and relationship to the sibling tool. The main gap is lack of information about return values or error conditions, which would be helpful given the absence of an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds minimal parameter semantics beyond the schema - it mentions cost variations by preset and explains the wait parameter's relationship to councly_status. This meets the baseline expectation when schema coverage is complete, but doesn't add significant additional parameter context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Create a council hearing where multiple LLMs debate a topic and a moderator synthesizes the verdict.' It specifies the action (create), resource (council hearing), and distinct mechanism (multiple LLMs debating with moderator synthesis). This differentiates it from the sibling tool councly_status, which checks status rather than creating hearings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage guidance with a dedicated 'Use cases' section listing four scenarios (code review, technical decisions, problem solving, brainstorming). It also distinguishes when to use this tool vs. alternatives by explaining the wait parameter behavior and referencing councly_status for checking status later. This gives clear context for when and how to use the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
councly_statusA
Check the status of a council hearing.
Returns:
For in-progress hearings: current phase and progress percentage
For completed hearings: verdict, trust score, and counsel summaries
For failed hearings: error message
Use this to check on hearings created with wait=false, or to retrieve past hearing results.
| Name | Required | Description | Default |
|---|---|---|---|
| hearing_id | Yes | The hearing ID to check |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses the tool's behavior by describing different return scenarios (in-progress, completed, failed) and what information each provides. However, it doesn't mention error handling beyond 'error message', rate limits, authentication requirements, or whether this is a read-only operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured with three distinct parts: purpose statement, return value scenarios, and usage guidelines. Every sentence adds value, with no redundant or unnecessary information. It's appropriately sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter status-checking tool with no output schema, the description provides good context: purpose, return scenarios, and usage guidelines. It could be more complete by explicitly stating this is a read-only operation and providing more detail about error conditions, but it covers the essential information an agent needs to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, with the single parameter 'hearing_id' well-documented as a UUID. The description doesn't add any parameter-specific information beyond what the schema provides, so the baseline score of 3 is appropriate given the schema does all the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose as 'Check the status of a council hearing' with specific verb+resource. It distinguishes from the sibling tool 'councly_hearing' by focusing on status checking rather than hearing creation/management, though it doesn't explicitly contrast them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage guidance: 'Use this to check on hearings created with wait=false, or to retrieve past hearing results.' This clearly indicates when to use this tool versus alternatives, including specific scenarios and timing considerations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
2 tool updates
v1.0.0- First observed
councly_hearing - First observed
councly_status
TDQS
The two tools have clearly distinct purposes: councly_hearing creates and runs a council hearing, while councly_status checks the status or retrieves results of an existing hearing. There is no overlap in functionality, making it easy for an agent to select the correct tool based on the desired action.
Both tools follow a consistent naming pattern with the prefix 'councly_' followed by a descriptive action (hearing, status). This verb_noun style is uniform and predictable, aiding in tool identification and usage without confusion.
With only 2 tools, the server feels thin for its apparent scope of facilitating multi-LLM debates and status tracking. While the tools cover creation and status checking, the domain suggests potential gaps (e.g., no tool for listing hearings, modifying settings, or handling errors beyond status retrieval), making the count insufficient for comprehensive coverage.
The tool surface is significantly incomplete for the server's purpose. It lacks operations such as listing past hearings, canceling or deleting hearings, configuring hearing parameters beyond defaults, or managing user settings. This forces agents into dead ends for common workflows, like reviewing multiple past hearings or adjusting debate parameters.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Convene a panel of expert AI personas to debate any decision from every side.
21Multi-LLM council: 25+ frontier models in parallel, consensus scoring, verdict-first code review.
Multi-model AI debates: GPT-4o, Claude, Gemini & 200+ models discuss, then synthesize insight.
Deliberation + live 5-model council divergence over the Omnarai multi-AI attributed corpus.
Related MCP Servers
- AlicenseAqualityDmaintenanceProvides access to multiple frontier LLM models (GPT, Claude, Gemini, Grok, DeepSeek) for consulting a "conclave" of AI perspectives, enabling peer-ranked evaluations and synthesized consensus answers for important decisions.81MIT
- AlicenseNot gradedqualityAmaintenanceEnables running AI councils across Claude, GPT, and Gemini, synthesizing answers based on your personal taste lens, all locally without an API key.1MIT
- AlicenseNot gradedqualityDmaintenanceEnables multi-strategy AI orchestration including council decision review, debate, brainstorming, evaluation, and spec review, with support for multiple LLM providers and advisor personas.1MIT
- AlicenseAqualityAmaintenanceRoutes questions to a council of AI models (local and cloud) and synthesizes their answers in five configurable modes: individual, categorized, deconflicted, pooled, and dialectic.9Apache 2.0
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/slmnsrf/councly-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server