Skip to main content
Glama

Speclint

Deterministic spec linter for AI coding agents. Score GitHub issues 0–100 across 5 dimensions before agents touch them — catch vague requirements, missing acceptance criteria, and untestable specs at the source.

What it does

Scores issues like:

  • "users keep saying login is broken"

  • "dashboard loads slow"

  • "need dark mode"

Across 5 dimensions:

  • Problem clarity — Is the problem statement specific and observable?

  • Acceptance criteria — Are there testable, concrete pass/fail conditions?

  • Scope definition — Is the work bounded and decomposable?

  • Verification steps — Can a CI agent prove it's done?

  • Testability — Are edge cases and failure modes addressed?

Returns a 0–100 score with per-dimension breakdown, agent-readiness flag, and optional AI-powered rewrite suggestions.

API

Lint a spec

curl -X POST https://speclint.ai/api/lint \
  -H "Content-Type: application/json" \
  -d '{"items": ["users keep saying login is broken", "dashboard loads slow"]}'

With a license key:

curl -X POST https://speclint.ai/api/lint \
  -H "Content-Type: application/json" \
  -H "x-license-key: SK-YOUR-KEY" \
  -d '{"items": ["users keep saying login is broken"]}'

Example response:

{
  "results": [{
    "item": "users keep saying login is broken",
    "score": 32,
    "agent_ready": false,
    "dimensions": {
      "problem_clarity": 20,
      "acceptance_criteria": 10,
      "scope_definition": 45,
      "verification_steps": 30,
      "testability": 55
    },
    "rewrite_preview": "**Problem:** Users cannot log in to the application..."
  }]
}

Rewrite a spec (Lite tier and above)

curl -X POST https://speclint.ai/api/rewrite \
  -H "Content-Type: application/json" \
  -H "x-license-key: SK-YOUR-KEY" \
  -d '{
    "item": "dashboard loads slow",
    "target_agent": "claude",
    "rewrite_mode": "full"
  }'

Example response:

{
  "rewritten": "**Problem:** The dashboard takes >3s to load on standard connections...",
  "structured": {
    "title": "Optimize dashboard load time to <1s on 4G",
    "problem": "The main dashboard takes 3-8s to load, causing 40% of users to abandon before seeing data.",
    "acceptance_criteria": [
      "Dashboard LCP < 1s on 4G (Lighthouse throttling preset)",
      "First contentful paint < 500ms",
      "All chart data visible within 2s without skeleton loaders"
    ],
    "verification_steps": [
      "Run Lighthouse CI in --preset=perf mode",
      "Assert LCP < 1000ms in CI",
      "Load dashboard with network throttled to 4G in Playwright test"
    ]
  },
  "score_before": 28,
  "score_after": 91,
  "score_delta": 63
}

Full OpenAPI spec: speclint.ai/openapi.yaml

Agent capabilities: speclint.ai/llms.txt

CLI

Install and run from your terminal:

npx @speclint/cli lint "dashboard loads slow"

Or install globally:

npm install -g @speclint/cli
speclint lint "dashboard loads slow"
speclint rewrite "dashboard loads slow" --key SK-YOUR-KEY
speclint batch issues.txt --key SK-YOUR-KEY

Set your key once via env: export SPECLINT_KEY=SK-YOUR-KEY

MCP Server

Use speclint-mcp directly in Claude Desktop, Cursor, or any MCP-compatible client:

{
  "mcpServers": {
    "speclint": {
      "command": "npx",
      "args": ["speclint-mcp"],
      "env": { "SPECLINT_KEY": "SK-YOUR-KEY" }
    }
  }
}

This gives your AI assistant a lint_spec tool it can call automatically before writing code.

GitHub Action

Lint specs automatically in CI. Trigger on issue open, manual dispatch, or any GitHub event.

- uses: DavidNielsen1031/speclint-action@v1
  with:
    items: ${{ github.event.issue.title }}
    write-back: "true"
    gherkin: "true"
    key: ${{ secrets.SPECLINT_KEY }}

Posts the score + rewrite suggestions as a comment on the issue.

GitHub Marketplace · Full docs + examples

Pricing

Tier

Price

Items/req

Rewrites/day

Keys

Free

$0

5

1 preview

1

Lite

$9/mo

5

10 full

1

Solo

$29/mo

25

500 full

1

Team

$79/mo

50

1,000 full

Unlimited

  • Free: No signup required. Get a free key to track usage. Rewrite previews are 250 chars.

  • Lite: Full rewrites (complete rewritten spec + structured fields + score delta). Unlimited lint requests.

  • Solo: 25 items per batch, 500 rewrites/day, codebase_context field for stack-aware scoring.

  • Team: 50 items per batch, 1,000 rewrites/day, multi-seat. For teams where bad specs cost real money.

Pass your license key via x-license-key header, SPECLINT_KEY env var, or the MCP server config.

Available Tools

1 tool
refine_backlogA

Refine messy backlog items into structured, actionable work items. Returns each item with a clean title, problem statement, acceptance criteria, T-shirt size estimate (XS/S/M/L/XL), priority with rationale, tags, and optional assumptions. Free tier: up to 5 items per request. Pro: 25. Team: 50.

BEFORE calling this tool, ask the user TWO quick questions if they haven't already specified:

  1. Would you like titles formatted as user stories? ("As a [user], I want [goal], so that [benefit]")

  2. Would you like acceptance criteria in Gherkin format? (Given/When/Then) Set useUserStories and useGherkin accordingly based on their answers. Both default to false.

LICENSE KEY: For unlimited requests and higher item limits, set REFINE_BACKLOG_KEY in your MCP server environment config (Claude Desktop → claude_desktop_config.json → env section). Get a key at https://refinebacklog.com/pricing

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYesArray of raw backlog item strings to refine. Each string is a rough description of work to be done.
contextNoOptional project context to improve relevance. Example: "B2B SaaS CRM for enterprise sales teams" or "Mobile fitness app for casual runners".
licenseKeyNoOptional. Refine Backlog license key for Pro or Team tier. Preferred: set REFINE_BACKLOG_KEY in your MCP server env config instead of passing inline. Get a key at https://refinebacklog.com/pricing. Free tier (5 items, 3 req/day) works without a key.
useUserStoriesNoFormat titles as user stories: "As a [user], I want [goal], so that [benefit]". Default: false.
useGherkinNoFormat acceptance criteria as Gherkin: Given/When/Then. Default: false.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes key behavioral traits: the transformation process (from messy to structured), output format details, tier-based rate limits (items per request), authentication/licensing requirements (license key for higher tiers), and default values for boolean parameters. It doesn't mention error handling or response time, but covers most critical aspects for a tool with no annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately front-loaded with the core purpose, but contains some redundancy (license key information appears twice) and could be more streamlined. The licensing details and URL reference, while important, add length. Most sentences earn their place, but the structure could be tighter with better grouping of related information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (5 parameters, no output schema, no annotations), the description provides substantial context: it explains the transformation process, output structure, tier limits, prerequisites (questions to ask), and licensing. However, without an output schema, it doesn't fully describe the return format (only lists fields without structure details), and some behavioral aspects like error conditions are missing. For a tool with this complexity, it's quite complete but has minor gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds significant value beyond the schema: it explains the purpose of the 'items' parameter ('raw backlog item strings to refine'), provides concrete examples for 'context', clarifies the relationship between 'licenseKey' and environment configuration, and gives formatting details for 'useUserStories' and 'useGherkin' that go beyond the schema's descriptions. However, it doesn't fully explain the semantics of all parameters (e.g., what 'T-shirt size estimate' means in practice).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Refine messy backlog items into structured, actionable work items' with specific outputs listed (clean title, problem statement, acceptance criteria, T-shirt size estimate, priority with rationale, tags, optional assumptions). It uses specific verbs ('refine', 'returns') and resources ('backlog items', 'work items'), and since there are no sibling tools, it doesn't need to differentiate from them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidelines: it instructs the agent to ask two specific questions before calling the tool if the user hasn't already specified them (about user stories and Gherkin format), and explains how to set parameters based on user answers. It also details tier limits (Free: 5 items, Pro: 25, Team: 50) and when to use the licenseKey parameter versus environment configuration.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A4.1/5.0
Disambiguation5/5

With only one tool, there is no possibility of ambiguity or overlap with other tools. The tool has a single, clearly defined purpose: refining backlog items into structured work items.

Naming Consistency5/5

There is only one tool name, 'refine_backlog', which follows a clear verb_noun pattern. Since there are no other tools to compare against, consistency is inherently perfect.

Tool Count2/5

A single tool is too few for a server that appears to handle backlog refinement, as it lacks complementary operations like listing, updating, or managing refined items. This minimal scope will likely cause agent failures due to incomplete workflows.

Completeness2/5

The tool surface is severely incomplete for backlog management. While the refine_backlog tool performs a specific transformation, there are no tools for creating, retrieving, updating, or deleting backlog items, leaving significant gaps in the domain coverage.

Related MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/DavidNielsen1031/refine-backlog-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server