refine-backlog-mcp
The refine-backlog-mcp server transforms messy, unstructured backlog items into clean, actionable work items.
Core capabilities:
Refine backlog items: Convert raw descriptions into structured work items with:
Clean title
Problem statement
Testable acceptance criteria
T-shirt size estimate (XS/S/M/L/XL)
Priority with rationale
Relevant tags
Optional assumptions
Format options: Titles as user stories ("As a [user], I want [goal], so that [benefit]") and acceptance criteria in Gherkin format (Given/When/Then)
Project context: Provide optional context (e.g., "B2B SaaS CRM") to improve output relevance
Batch processing: Handle multiple items per request (Free: 5, Pro: 25, Team: 50)
Tiered access: Free tier (5 items/request, 3 requests/day) or higher limits with a license key via
REFINE_BACKLOG_KEYenv variable or inline parameterMCP integration: Works with MCP clients like Claude Desktop and Cursor
Enables automated scoring and linting of GitHub issues to identify vague requirements and missing acceptance criteria. Provides a GitHub Action that can automatically analyze issue content and post quality scores and suggested rewrites as comments directly on the platform.
Speclint
Deterministic spec linter for AI coding agents. Score GitHub issues 0–100 across 5 dimensions before agents touch them — catch vague requirements, missing acceptance criteria, and untestable specs at the source.
What it does
Scores issues like:
"users keep saying login is broken"
"dashboard loads slow"
"need dark mode"
Across 5 dimensions:
Problem clarity — Is the problem statement specific and observable?
Acceptance criteria — Are there testable, concrete pass/fail conditions?
Scope definition — Is the work bounded and decomposable?
Verification steps — Can a CI agent prove it's done?
Testability — Are edge cases and failure modes addressed?
Returns a 0–100 score with per-dimension breakdown, agent-readiness flag, and optional AI-powered rewrite suggestions.
API
Lint a spec
curl -X POST https://speclint.ai/api/lint \
-H "Content-Type: application/json" \
-d '{"items": ["users keep saying login is broken", "dashboard loads slow"]}'With a license key:
curl -X POST https://speclint.ai/api/lint \
-H "Content-Type: application/json" \
-H "x-license-key: SK-YOUR-KEY" \
-d '{"items": ["users keep saying login is broken"]}'Example response:
{
"results": [{
"item": "users keep saying login is broken",
"score": 32,
"agent_ready": false,
"dimensions": {
"problem_clarity": 20,
"acceptance_criteria": 10,
"scope_definition": 45,
"verification_steps": 30,
"testability": 55
},
"rewrite_preview": "**Problem:** Users cannot log in to the application..."
}]
}Rewrite a spec (Lite tier and above)
curl -X POST https://speclint.ai/api/rewrite \
-H "Content-Type: application/json" \
-H "x-license-key: SK-YOUR-KEY" \
-d '{
"item": "dashboard loads slow",
"target_agent": "claude",
"rewrite_mode": "full"
}'Example response:
{
"rewritten": "**Problem:** The dashboard takes >3s to load on standard connections...",
"structured": {
"title": "Optimize dashboard load time to <1s on 4G",
"problem": "The main dashboard takes 3-8s to load, causing 40% of users to abandon before seeing data.",
"acceptance_criteria": [
"Dashboard LCP < 1s on 4G (Lighthouse throttling preset)",
"First contentful paint < 500ms",
"All chart data visible within 2s without skeleton loaders"
],
"verification_steps": [
"Run Lighthouse CI in --preset=perf mode",
"Assert LCP < 1000ms in CI",
"Load dashboard with network throttled to 4G in Playwright test"
]
},
"score_before": 28,
"score_after": 91,
"score_delta": 63
}Full OpenAPI spec: speclint.ai/openapi.yaml
Agent capabilities: speclint.ai/llms.txt
CLI
Install and run from your terminal:
npx @speclint/cli lint "dashboard loads slow"Or install globally:
npm install -g @speclint/cli
speclint lint "dashboard loads slow"
speclint rewrite "dashboard loads slow" --key SK-YOUR-KEY
speclint batch issues.txt --key SK-YOUR-KEYSet your key once via env: export SPECLINT_KEY=SK-YOUR-KEY
MCP Server
Use speclint-mcp directly in Claude Desktop, Cursor, or any MCP-compatible client:
{
"mcpServers": {
"speclint": {
"command": "npx",
"args": ["speclint-mcp"],
"env": { "SPECLINT_KEY": "SK-YOUR-KEY" }
}
}
}This gives your AI assistant a lint_spec tool it can call automatically before writing code.
GitHub Action
Lint specs automatically in CI. Trigger on issue open, manual dispatch, or any GitHub event.
- uses: DavidNielsen1031/speclint-action@v1
with:
items: ${{ github.event.issue.title }}
write-back: "true"
gherkin: "true"
key: ${{ secrets.SPECLINT_KEY }}Posts the score + rewrite suggestions as a comment on the issue.
→ GitHub Marketplace · Full docs + examples
Pricing
Tier | Price | Items/req | Rewrites/day | Keys |
Free | $0 | 5 | 1 preview | 1 |
Lite | $9/mo | 5 | 10 full | 1 |
Solo | $29/mo | 25 | 500 full | 1 |
Team | $79/mo | 50 | 1,000 full | Unlimited |
Free: No signup required. Get a free key to track usage. Rewrite previews are 250 chars.
Lite: Full rewrites (complete rewritten spec + structured fields + score delta). Unlimited lint requests.
Solo: 25 items per batch, 500 rewrites/day,
codebase_contextfield for stack-aware scoring.Team: 50 items per batch, 1,000 rewrites/day, multi-seat. For teams where bad specs cost real money.
Pass your license key via x-license-key header, SPECLINT_KEY env var, or the MCP server config.
Links
Available Tools
1 toolrefine_backlogA
Refine messy backlog items into structured, actionable work items. Returns each item with a clean title, problem statement, acceptance criteria, T-shirt size estimate (XS/S/M/L/XL), priority with rationale, tags, and optional assumptions. Free tier: up to 5 items per request. Pro: 25. Team: 50.
BEFORE calling this tool, ask the user TWO quick questions if they haven't already specified:
Would you like titles formatted as user stories? ("As a [user], I want [goal], so that [benefit]")
Would you like acceptance criteria in Gherkin format? (Given/When/Then) Set useUserStories and useGherkin accordingly based on their answers. Both default to false.
LICENSE KEY: For unlimited requests and higher item limits, set REFINE_BACKLOG_KEY in your MCP server environment config (Claude Desktop → claude_desktop_config.json → env section). Get a key at https://refinebacklog.com/pricing
| Name | Required | Description | Default |
|---|---|---|---|
| items | Yes | Array of raw backlog item strings to refine. Each string is a rough description of work to be done. | |
| context | No | Optional project context to improve relevance. Example: "B2B SaaS CRM for enterprise sales teams" or "Mobile fitness app for casual runners". | |
| licenseKey | No | Optional. Refine Backlog license key for Pro or Team tier. Preferred: set REFINE_BACKLOG_KEY in your MCP server env config instead of passing inline. Get a key at https://refinebacklog.com/pricing. Free tier (5 items, 3 req/day) works without a key. | |
| useUserStories | No | Format titles as user stories: "As a [user], I want [goal], so that [benefit]". Default: false. | |
| useGherkin | No | Format acceptance criteria as Gherkin: Given/When/Then. Default: false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes key behavioral traits: the transformation process (from messy to structured), output format details, tier-based rate limits (items per request), authentication/licensing requirements (license key for higher tiers), and default values for boolean parameters. It doesn't mention error handling or response time, but covers most critical aspects for a tool with no annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately front-loaded with the core purpose, but contains some redundancy (license key information appears twice) and could be more streamlined. The licensing details and URL reference, while important, add length. Most sentences earn their place, but the structure could be tighter with better grouping of related information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (5 parameters, no output schema, no annotations), the description provides substantial context: it explains the transformation process, output structure, tier limits, prerequisites (questions to ask), and licensing. However, without an output schema, it doesn't fully describe the return format (only lists fields without structure details), and some behavioral aspects like error conditions are missing. For a tool with this complexity, it's quite complete but has minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds significant value beyond the schema: it explains the purpose of the 'items' parameter ('raw backlog item strings to refine'), provides concrete examples for 'context', clarifies the relationship between 'licenseKey' and environment configuration, and gives formatting details for 'useUserStories' and 'useGherkin' that go beyond the schema's descriptions. However, it doesn't fully explain the semantics of all parameters (e.g., what 'T-shirt size estimate' means in practice).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Refine messy backlog items into structured, actionable work items' with specific outputs listed (clean title, problem statement, acceptance criteria, T-shirt size estimate, priority with rationale, tags, optional assumptions). It uses specific verbs ('refine', 'returns') and resources ('backlog items', 'work items'), and since there are no sibling tools, it doesn't need to differentiate from them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage guidelines: it instructs the agent to ask two specific questions before calling the tool if the user hasn't already specified them (about user stories and Gherkin format), and explains how to set parameters based on user answers. It also details tier limits (Free: 5 items, Pro: 25, Team: 50) and when to use the licenseKey parameter versus environment configuration.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
With only one tool, there is no possibility of ambiguity or overlap with other tools. The tool has a single, clearly defined purpose: refining backlog items into structured work items.
There is only one tool name, 'refine_backlog', which follows a clear verb_noun pattern. Since there are no other tools to compare against, consistency is inherently perfect.
A single tool is too few for a server that appears to handle backlog refinement, as it lacks complementary operations like listing, updating, or managing refined items. This minimal scope will likely cause agent failures due to incomplete workflows.
The tool surface is severely incomplete for backlog management. While the refine_backlog tool performs a specific transformation, there are no tools for creating, retrieving, updating, or deleting backlog items, leaving significant gaps in the domain coverage.
Related MCP Connectors
Capture feature requests and bug reports from chat into a searchable, AI-categorized backlog.
Opinionated sprint tracker. Read/update tickets, sprints, velocity from Claude/Cursor/Zed.
Track stories, organize sprints, and manage project workflows across your team
Turn raw customer feedback into evidence-cited specs (free, no key) plus 16 PM tools.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/DavidNielsen1031/refine-backlog-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server