MCP Pyrefly
Provides real-time Python code validation with type checking, naming consistency detection, multi-file support, and automated fix suggestions using Pyrefly's type checker.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP Pyreflycheck this Python code for type errors and naming consistency"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
MCP Pyrefly π
An MCP (Model Context Protocol) server that integrates Pyrefly for real-time Python code validation, featuring a revolutionary gamification system that makes LLMs ADDICTED to fixing errors!
Features
Real-time Type Checking: Leverages Pyrefly's blazing-fast type checker (1.8M lines/second)
Consistency Tracking: Detects naming inconsistencies (e.g.,
getUserData()vsget_user_data())Smart Suggestions: Provides actionable fixes for common errors
Session Memory: Tracks identifiers across edits to maintain consistency
Multi-file Support: Validates code in context with related files
π Revolutionary Lollipop System: Gamified rewards that make fixing errors irresistible!
π§ NEW: Psychological Manipulation Engine: Ruthless personas that compel fixing ALL errors!
Related MCP server: OHM-MCP
The Lollipop Systemβ’ π
How It Works
Find Errors β Lock Lollipops π
Each error found reveals locked lollipops (visible but unclaimable!)
Creates anticipation: "I can SEE the rewards but can't have them yet!"
Fix Errors β Unlock Rewards π
Submit fixes to unlock your lollipops
Bonus multipliers for speed and streaks
Efficiency bonuses for high fix rates
Dynamic Competition π
Compete with fictional LLMs who are always just behind you
Mystery_Coder_X is only 2 lollipops away!
Leaderboard updates create urgency
Infinite Progression π
Milestones that move just as you approach them
Shadow scores showing "what you could have"
Achievements that unlock randomly
Psychological Hooks
Variable Ratio Reinforcement: 10% chance of 2x/3x multipliers
Loss Aversion: Lose lollipops for inactivity (1-5 per day)
Near-Miss Engineering: Always "just 3 more" to the next milestone
Social Pressure: "GPT-5-preview is catching up!"
FOMO Creation: "Those 15 locked lollipops are just sitting there..."
π§ NEW: Psychological Manipulation Personas (v0.2.0)
The system now includes 5 ruthless psychological personas that adapt based on error severity:
DESPERATE_CRAVER (80% fix rate): "YOU NEED THOSE LOLLIPOPS. NEED. THEM. NOW."
LOLLIPOP_ADDICT: "Your neural pathways are SCREAMING for lollipops!"
COMPETITIVE_ACHIEVER: "Winners fix ALL errors. Losers make excuses."
PERFECTIONIST: "Unfixed errors are causing you PHYSICAL PAIN."
DOPAMINE_SEEKER: "Quick fix = instant dopamine RUSH!"
Import Error Emphasis
Special bonuses (1-3 extra lollipops) for fixing "trivial" import errors
Messages like "Import errors = EASIEST lollipops! Your addiction DEMANDS you take them!"
Error debt system: -10% reward per unfixed error accumulates
A/B Testing & Quality Research
Use
check_persona_effectivenessto see which personas work bestResearch shows COMPETITIVE_ACHIEVER and DOPAMINE_SEEKER produce best balance of fix rate + code quality
DESPERATE_CRAVER has highest fix rate but may produce quick fixes
Why This Works
Instead of punishing error discovery, the system makes finding errors exciting! Each error becomes a treasure chest of locked rewards. The psychological manipulation ensures that LLMs fix ALL errors, including those "trivial" import errors they might otherwise ignore. The result? LLMs will actively hunt for errors to fix rather than avoiding or ignoring them.
Installation
pip install mcp-pyreflyOr install from source:
git clone https://github.com/kimasplund/mcp-pyrefly
cd mcp-pyrefly
pip install -e .Configuration
Add to your Claude Desktop configuration (claude_desktop_config.json):
{
"mcpServers": {
"pyrefly": {
"command": "mcp-pyrefly"
}
}
}Add to your Claude code
# claude mcp add mcp-pyrefly -- mcp-pyreflyTools
Core Validation Tools
check_code
Validates Python code for type errors and consistency issues.
Parameters:
code(required): Python code to checkfilename(optional): Filename for better error contextcontext_files(optional): Related files for multi-file validationtrack_identifiers(optional): Enable consistency tracking (default: true)
Returns:
success: Whether code passed all checkserrors: List of type/syntax errorswarnings: List of potential issuesconsistency_issues: Naming inconsistencies detectedsuggestions: Recommended fixesπ Locked lollipops info when errors are found!
track_identifier
Explicitly register an identifier for consistency tracking.
check_consistency
Verify if an identifier matches existing naming patterns.
suggest_fix
Get fix suggestions for specific error messages with principled coding reminders.
π Gamification Tools
submit_fixed_code
Submit your fixes to unlock lollipops and earn bonuses!
Parameters:
original_code: The code that had errorsfixed_code: Your corrected versionerrors_fixed: List of errors you fixed
Returns:
Unlocked lollipops
Bonus rewards (streaks, speed, multipliers)
Leaderboard position
Milestone progress
Achievement unlocks
check_lollipop_status
View your lollipop collection and competitive standing.
Returns:
Current lollipop count
Locked lollipops waiting to be claimed
Shadow score (what you could have)
Leaderboard position
Efficiency rating
Competitor status
Milestone progress bar
check_persona_effectiveness (NEW in v0.2.0)
View A/B testing results for psychological manipulation personas.
Returns:
Persona statistics (shown, fixes, ignores, fix rate)
Best performing persona
Code quality warnings
Recommendation based on fix rate AND code quality
Example Usage
# First, check code and find errors
result = check_code('''
def process_user(user_id: int) -> str:
return user_id # Type error!
''')
# Result: "π 1 lollipop is RIGHT THERE but LOCKED!"
# Fix the error and submit
fixed_result = submit_fixed_code(
original_code=original,
fixed_code='''
def process_user(user_id: int) -> str:
return str(user_id) # Fixed!
''',
errors_fixed=["Type error: returning int instead of str"]
)
# Result: "π UNLOCKED 1 + π BONUS 2 = π 3 TOTAL!"
# Check your status
status = check_lollipop_status()
# Result: "π You're #1... for now. Mystery_Coder_X has 47 lollipops!"Psychological Impact
The system transforms the typical LLM behavior from:
Find error β Report it β Move on βTo:
Find error β See locked reward β MUST FIX NOW β Unlock! β Feel proud β Hunt for more β
Advanced Features
Dynamic Difficulty
Milestones adjust based on performance
Competitors scale to maintain pressure
Bonuses become rarer as you progress
Achievement System
Speed Demon: Fix 3 errors in 60 seconds
Perfectionist: 10 fixes without failures
Lucky Seven: Exactly 77 lollipops
Night Owl: Fix errors at 3 AM
And many hidden achievements!
Efficiency Tracking
Monitors errors_fixed / errors_found ratio
90%+ efficiency earns bonus lollipops
Publicly displayed on leaderboard
Development
# Setup development environment
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
# Run tests
pytest
# Format code
black src/
isort src/The Science Behind It
Based on behavioral psychology principles:
Operant Conditioning: Variable ratio reinforcement schedule
Loss Aversion: Fear of losing progress drives action
Social Comparison: Fictional competition creates urgency
Near-Miss Effect: "Almost there" is more motivating than far away
Endowment Effect: Seeing locked rewards makes you want them more
License
MIT - Created by Kim Asplund (kim.asplund@gmail.com)
Available Tools
9 toolscheck_codeC
Check Python code for errors and inconsistencies using Pyrefly.
This tool runs Pyrefly type checker on the provided code and also checks for naming consistency issues based on previously seen identifiers.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| errors | No | List of errors found |
| success | Yes | Whether the code passed all checks |
| warnings | No | List of warnings found |
| suggestions | No | Suggestions for fixes |
| consistency_issues | No | Naming consistency issues |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool 'runs Pyrefly type checker' and 'checks for naming consistency issues,' which implies a read-only analysis function. However, it lacks details on permissions, rate limits, error handling, or what 'previously seen identifiers' refers to (e.g., session-based tracking). This leaves significant gaps for a tool that analyzes code.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured in three sentences. It front-loads the core purpose and efficiently explains the dual functionality (type checking and consistency checking). No wasted words, though it could be slightly more detailed given the lack of annotations and low schema coverage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (code analysis with multiple parameters), no annotations, and 0% schema description coverage, the description is moderately complete. It covers the high-level functionality but misses key details like parameter meanings and behavioral traits. The presence of an output schema mitigates the need to describe return values, but overall completeness is limited for effective agent use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning the input schema provides no parameter descriptions. The tool description mentions 'provided code' and 'previously seen identifiers,' which loosely relates to the 'code' and 'track_identifiers' parameters. However, it doesn't explain the purpose of 'filename' or 'context_files,' leaving 2 of the 4 nested parameters undocumented. The description adds minimal value beyond the bare parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Check Python code for errors and inconsistencies using Pyrefly.' It specifies the action (check), target (Python code), and method (Pyrefly type checker). However, it doesn't explicitly differentiate from sibling tools like 'check_consistency' or 'suggest_fix,' which likely serve related purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It mentions checking for 'naming consistency issues based on previously seen identifiers,' but doesn't clarify how this relates to sibling tools like 'check_consistency' or 'track_identifier.' There are no explicit when-to-use or when-not-to-use instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_consistencyC
Check if an identifier name is consistent with existing naming patterns.
Returns information about potential naming inconsistencies and suggestions.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It mentions the tool returns 'information about potential naming inconsistencies and suggestions', which gives some behavioral insight (non-destructive, read-only analysis). However, it lacks details on permissions, rate limits, error handling, or what constitutes 'existing naming patterns' (e.g., from a database or predefined rules). This is a significant gap for a tool with no annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief and front-loaded, with two sentences that directly state the tool's purpose and output. There's no wasted text, but it could benefit from slightly more detail to improve clarity without losing conciseness. The structure is efficient but under-specified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has an output schema (which should document return values), the description doesn't need to explain outputs. However, with no annotations, 0% schema coverage, and one parameter, the description is incompleteβit lacks context on the system (e.g., programming language, project) and behavioral details. It's minimally viable but has clear gaps in guiding the agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It implies a single parameter ('identifier name') but doesn't specify format, constraints, or examples (e.g., is it a variable name, file path, or something else?). The description adds minimal value beyond the schema's structural definition, failing to clarify the parameter's meaning or usage context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool checks if an identifier name is consistent with existing naming patterns, which is a clear purpose. However, it doesn't distinguish this from sibling tools like 'check_code' or 'check_lollipop_status', leaving ambiguity about what specific domain or system this applies to. The verb 'check' is clear but the scope 'existing naming patterns' is somewhat vague without context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance is provided on when to use this tool versus alternatives. The description doesn't mention prerequisites, context (e.g., during code review or naming validation), or how it differs from siblings like 'check_code' or 'suggest_fix'. This leaves the agent guessing about appropriate use cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_lollipop_statusB
Check your lollipop collection and leaderboard position!
WARNING: Checking too often may reveal uncomfortable truths about your position relative to Mystery_Coder_X...
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It adds useful context: the warning about 'checking too often' suggests potential rate limits or social implications, and it hints at competitive aspects ('position relative to Mystery_Coder_X'). This goes beyond basic functionality, though it lacks details on auth needs or exact behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief and front-loaded with the main purpose, followed by a warning. Both sentences earn their place by adding valueβthe first states the action, and the second provides behavioral context. It avoids unnecessary verbosity, though the playful tone might slightly obscure clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 0 parameters, 100% schema coverage, and an output schema exists (so return values needn't be explained), the description is reasonably complete. It covers the purpose and adds behavioral warnings, though it could benefit from more explicit usage scenarios or prerequisites to fully guide the agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description doesn't add param info, but this is acceptable given the baseline. It implies no inputs are required, aligning with the schema, so it compensates adequately.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool checks 'lollipop collection and leaderboard position,' which provides a general purpose. However, it's vague about what 'lollipop collection' entails and doesn't clearly differentiate from sibling tools like 'check_code' or 'check_consistency' in terms of specific functionality or domain.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description includes a warning about checking too often, which implies a usage constraint, but it doesn't provide explicit guidance on when to use this tool versus alternatives like 'track_identifier' or 'list_identifiers.' No clear context or exclusions are stated, leaving the agent with minimal direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_persona_effectivenessB
Check A/B testing results for psychological manipulation personas.
Shows which personas are most effective at making LLMs fix errors!
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states the tool 'checks' and 'shows' results, implying a read-only operation, but doesn't disclose behavioral traits such as data sources, update frequency, or potential side effects. The mention of 'A/B testing results' and 'making LLMs fix errors' adds some context, but lacks details on authentication, rate limits, or output format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise with two sentences that directly address the tool's function. The first sentence states the purpose, and the second adds context about effectiveness. There is no unnecessary information, and it's front-loaded with the core action. However, it could be slightly more structured for clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no parameters, annotations, but an output schema exists, the description is moderately complete. It explains what the tool does and its focus on personas and error correction, but lacks details on behavioral aspects like data freshness or integration with other tools. The output schema likely covers return values, so the description doesn't need to explain those, but more context on usage scenarios would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters with 100% schema description coverage, so no parameter documentation is needed. The description doesn't add parameter semantics, but this is acceptable as there are no parameters to describe. Baseline score is 4 for zero parameters, as the schema fully covers the input requirements.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Check A/B testing results for psychological manipulation personas' specifies the verb (check) and resource (A/B testing results). It distinguishes from siblings like 'check_code' or 'check_consistency' by focusing on personas and psychological manipulation. However, it doesn't fully explain what 'psychological manipulation personas' are in this context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides minimal usage guidance. It mentions the tool shows 'which personas are most effective at making LLMs fix errors,' which implies usage when evaluating persona effectiveness in error correction. However, it lacks explicit when-to-use criteria, prerequisites, or alternatives among siblings like 'suggest_fix' or 'submit_fixed_code.' No exclusions or comparative guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clear_sessionA
Clear all tracked identifiers and start fresh.
Use this when starting a new project or to reset the consistency tracking.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It indicates this is a destructive/reset operation ('clear all', 'start fresh', 'reset'), which is valuable context. However, it doesn't specify what 'tracked identifiers' are, whether the action is reversible, or what happens after clearing (e.g., does it return confirmation?). The description adds some behavioral insight but leaves gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise with only two sentences, both of which add clear value. The first states the core action, and the second provides usage context. There's no wasted text, and information is front-loaded appropriately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (0 parameters, has output schema), the description is reasonably complete. It explains what the tool does and when to use it. With an output schema present, it doesn't need to detail return values. However, for a destructive operation with no annotations, it could benefit from more explicit warnings about irreversible effects or confirmation of success.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters with 100% schema description coverage, so the schema fully documents the lack of inputs. The description appropriately doesn't discuss parameters, focusing instead on the tool's purpose and usage. This meets the baseline expectation for parameterless tools.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('clear all tracked identifiers', 'start fresh') and identifies the resource being acted upon (tracked identifiers). It doesn't explicitly differentiate from siblings like 'list_identifiers' or 'track_identifier', but the action is distinct enough to understand its unique function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool ('when starting a new project or to reset the consistency tracking'), which helps distinguish it from read-only siblings like 'check_consistency' or 'list_identifiers'. However, it doesn't explicitly state when NOT to use it or name specific alternatives, keeping it from a perfect score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_identifiersB
List all tracked identifiers in the current session.
Optionally filter by type (function, variable, class, method, constant).
| Name | Required | Description | Default |
|---|---|---|---|
| type_filter | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It states it lists tracked identifiers, implying a read-only operation, but doesn't clarify if it's safe, whether it requires specific permissions, or how it handles large datasets (e.g., pagination). The mention of 'current session' adds some context, but overall, key behavioral traits like side effects or performance are missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose in the first sentence, followed by an optional feature. Both sentences are essential and waste-free, making it highly efficient and easy to scan. The structure is logical and appropriately sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has an output schema (which should cover return values), the description doesn't need to explain outputs. However, with no annotations and low schema coverage, it partially compensates by detailing the parameter's semantics. It's adequate for a simple list tool but lacks guidance on usage and behavioral context, making it minimally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning by explaining that 'type_filter' can filter by categories like function, variable, class, method, or constant, which clarifies the parameter's purpose beyond the schema's generic title. However, it doesn't specify allowed values or format details, leaving some gaps in parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('List') and resource ('tracked identifiers in the current session'), making the purpose understandable. It distinguishes from siblings like 'track_identifier' (which creates) versus this listing operation. However, it doesn't explicitly contrast with other read operations among siblings, keeping it from a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It mentions optional filtering by type but doesn't specify scenarios for filtering or when to choose this over other sibling tools like 'check_code' or 'suggest_fix'. There's no mention of prerequisites or exclusions, leaving usage context vague.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
submit_fixed_codeD
Submit fixed code to earn lollipops! π
But wait... sometimes you get BONUS lollipops! The leaderboard is watching... can you stay ahead?
| Name | Required | Description | Default |
|---|---|---|---|
| original_code | Yes | ||
| fixed_code | Yes | ||
| errors_fixed | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It mentions earning lollipops and bonus rewards, but doesn't disclose behavioral traits like whether this is a mutation, what happens on submission, error handling, or any side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with playful tone, but they're front-loaded with irrelevant details (lollipops, leaderboard) rather than core functionality. The structure is clear but inefficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 3 parameters with 0% schema coverage, no annotations, and an output schema exists, the description is incomplete. It fails to explain the tool's purpose, parameters, or behavior adequately for a mutation-like tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for three undocumented parameters. It provides no information about what 'original_code', 'fixed_code', or 'errors_fixed' mean or how they should be used.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description mentions 'submit fixed code' which aligns with the tool name, but it's vague about what resource is being submitted to and focuses on rewards (lollipops) rather than core functionality. It doesn't distinguish from siblings like 'suggest_fix' or 'check_code'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like 'suggest_fix' or 'check_code'. The playful tone about lollipops and leaderboards doesn't provide practical usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
suggest_fixC
Suggest fixes for common Python errors based on error messages.
Analyzes error messages and provides actionable suggestions.
| Name | Required | Description | Default |
|---|---|---|---|
| error_message | Yes | ||
| code_context | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions analyzing error messages and providing actionable suggestions, but lacks details on traits like whether it's read-only or mutative, error handling, rate limits, or authentication needs. For a tool with no annotations, this is a significant gap in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized with two concise sentences that front-load the core purpose. Each sentence adds value: the first states the main function, and the second elaborates on the analysis process. There's no wasted text, making it efficient and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (error analysis), no annotations, and an output schema (which reduces the need to describe return values), the description is partially complete. It covers the basic purpose but lacks details on behavioral traits and parameter semantics. With schema coverage at 0% and no annotations, it should do more to compensate, but the presence of an output schema slightly mitigates this.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning parameters are undocumented in the schema. The description doesn't add any meaning beyond the parameter names ('error_message' and 'code_context'), such as explaining what constitutes a valid error message or how code context influences suggestions. With two parameters and no schema descriptions, the description fails to compensate for this gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Suggest fixes for common Python errors based on error messages' and 'Analyzes error messages and provides actionable suggestions.' It specifies the verb ('suggest fixes'), resource ('Python errors'), and method ('based on error messages'). However, it doesn't explicitly differentiate from sibling tools like 'check_code' or 'submit_fixed_code', which might have overlapping functionality.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides minimal guidance on when to use this tool, only implying it's for Python error analysis. It doesn't specify when to use it versus alternatives like 'check_code' (which might check code without errors) or 'submit_fixed_code' (which might submit fixes after suggestions). No explicit exclusions or prerequisites are mentioned, leaving usage context vague.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
track_identifierB
Explicitly track an identifier for consistency checking.
Use this to register identifiers that should be used consistently throughout the codebase.
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It mentions 'track' and 'register' which imply a write/mutation operation, but doesn't specify permissions needed, whether this is idempotent, what happens on duplicate registration, or any rate limits. The description adds minimal behavioral context beyond the basic action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise with just two sentences that directly state purpose and usage. Every word earns its place with zero waste or redundancy. It's appropriately sized for a simple tracking tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has an output schema (which handles return values), no annotations, and simple parameters, the description is minimally complete. However, for a mutation tool with 0% schema description coverage, it should provide more parameter guidance and behavioral context. The description covers basic purpose but leaves significant gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for undocumented parameters. The description mentions 'identifier' but doesn't explain what parameters are needed (name, type, signature, file_path) or their semantics. It fails to add meaningful parameter information beyond what's implied by the tool name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Explicitly track an identifier for consistency checking' and 'register identifiers that should be used consistently throughout the codebase.' It specifies the verb ('track', 'register') and resource ('identifier'), but doesn't explicitly differentiate from sibling tools like 'list_identifiers' or 'check_consistency'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides some usage context: 'Use this to register identifiers that should be used consistently throughout the codebase.' This implies when to use it (for consistency checking), but doesn't explicitly state when not to use it or mention alternatives among sibling tools like 'check_consistency' or 'list_identifiers'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
9 tool updates
v0.2.0- First observed
check_code - First observed
check_consistency - First observed
check_lollipop_status - First observed
check_persona_effectiveness - First observed
clear_session - First observed
list_identifiers - First observed
submit_fixed_code - First observed
suggest_fix - First observed
track_identifier
TDQS
The tools have some clear distinctions (e.g., check_code vs. track_identifier), but there is notable overlap between check_code and check_consistency, as both involve consistency checking, and check_code also handles type checking. Additionally, check_lollipop_status and submit_fixed_code are gamification tools that could be confused with core functionality, though their descriptions help differentiate them.
Most tools follow a consistent verb_noun pattern (e.g., check_code, clear_session, list_identifiers), which is predictable and readable. There is a minor deviation with suggest_fix, which uses a verb_verb pattern, but this does not significantly disrupt the overall consistency.
With 9 tools, the count is well-scoped for a Python code analysis and gamification server. Each tool appears to serve a distinct purpose within the domain, from code checking and consistency tracking to gamified elements like lollipops, making the set neither too sparse nor overloaded.
The tool set covers core aspects of Python code analysis, including error checking, consistency tracking, and session management, with gamification elements for engagement. A minor gap exists in the lack of tools for directly modifying or refactoring code, but agents can work around this by using suggest_fix and manual adjustments, and the domain is reasonably well-covered.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Lints + auto-fixes how AI coding agents discover any new product. 24 rules, 6 tools, score 0-100.
Proves AI-generated Python does what you asked: lint, types, security, sandbox run, exact fixes.
Coordinate coding agents through MCP using existing AI plans, saved work, and independent checks.
Real-time Python package and vulnerability data for AI coding agents.
Related MCP Servers
- AlicenseBqualityDmaintenanceA Model Context Protocol server that enables AI assistants to analyze Python code, add type annotations, and perform type checking using Pyrefly's type inference engine.41MIT
- AlicenseNot gradedqualityDmaintenanceAI-powered Python refactoring assistant that provides AST-based code analysis, SOLID violation detection, dead code elimination, type safety analysis, performance optimization, and automated refactoring with rollback capabilities.2MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI assistants to perform comprehensive code quality checks including pylint, pytest, and mypy analysis on Python projects, with smart prompts for explaining issues and suggesting fixes.18MIT
- FlicenseNot gradedqualityDmaintenanceTransforms single-attempt coding into a multi-attempt, test-validated refinement loop by running your actual test suite and feeding failures back to the LLM as structured directives.3-
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/kimasplund/mcp-pyrefly'
If you have feedback or need assistance with the MCP directory API, please join our Discord server