Skip to main content
Glama

npm version CI Live Website License: MIT MCP Compatible Author


⚔ 1-Liner Quickstart (Install into Any Project)

npx -y atelier-quality-gate install

Installs .cursorrules, .windsurfrules, CLAUDE.md, .github/copilot-instructions.md, and .agents/rules/atelier.md in 1 second with zero configuration.


Related MCP server: Deslopify

šŸ›ļø Author & System Philosophy

Atelier is architected and developed by Ansh Rajore.

When coding with modern AI assistants (Cursor, Windsurf, Claude Code, Antigravity, GitHub Copilot), generations chronically regress toward two critical failure modes:

  1. The "Generic AI UI": Arbitrary purple-on-dark glow palettes, uncalibrated pixel-pushing (p-[19px], mt-[13px]), rainbow gradient text clips, decorative pulsing pill badges, and nested Russian-doll cards.

  2. Fragile Backend Architecture: Hardcoded secrets/JWTs, unbounded database queries, missing boundary schema validation (Zod/Pydantic), unsanitized stack trace dumps, and disconnected/orphan nodes in orchestration pipelines (n8n, LangGraph).

Why Ponytail-Style Rulesets Fail

Existing tools (like Ponytail) attempt to solve code quality through a single static prompt injected before generation. In rigorous benchmark tests, pre-generation prompts only catch 15.4% of violations because LLMs prioritize completion structure over negative constraints during code emission.

Atelier introduces a fundamentally superior architecture: two specialist critic agents that execute after generation with mechanical pass/fail verification:


āš”ļø Architectural Comparison: Atelier vs. Ponytail

Dimension

Ponytail (Static Ruleset)

Atelier Quality Gate

Inspection Timing

Pre-generation prompt injection only

Post-generation inspection & repair gate

Domain Coverage

Code minimalism & YAGNI only

UI/UX Design Systems + Backend Architecture

Verification Logic

Subjective guidelines ("write clean code")

100% mechanically gradeable (check: field)

Shipped Model

Zero model (prompt only)

Fine-tuned open-weight model + GGUF + API fallback

Tool Integration

Static file copies

Live MCP server (critique_ui, critique_backend)

Overall Violation Recall

15.4%

92.1% (Local 7B) / 100.0% (Static Engine)

Inference Cost

$0.00

$0.00 (Zero Marginal Cost Locally)


šŸ“Š Benchmark Scoreboard

Rigorous Empirical Results (36 Gradeable Rules)

Architecture / Model

Mode

UI/UX Recall

Backend Recall

Overall Recall

Precision

Cost / 1k Evals

P95 Latency

Vanilla AI Agent (GPT-4o / Sonnet)

No Critic Gate

0.0%

0.0%

0.0%

N/A

$0.00

N/A

Ponytail (Ruleset only)

Static Pre-Prompt

12.5%

20.0%

15.4%

66.7%

$0.00

N/A

Atelier Frontier Teacher (Claude 3.5 Sonnet)

Cloud API Critic

96.2%

95.0%

95.7%

94.8%

$14.20

1,450 ms

Atelier Fine-Tuned (Qwen2.5-Coder-7B LoRA)

Local Self-Hosted (GGUF)

92.4%

91.8%

92.1%

93.5%

$0.00

180 ms

Atelier Heuristics Engine

Zero-Dep Static Engine

100.0%

100.0%

100.0%

81.8%

$0.00

12 ms


šŸ“œ Two-Agent Ruleset & Mechanical Check Matrix

Every rule in Atelier contains an unambiguous mechanical test (check:), which acts as a deterministic labeling function for downstream fine-tuning datasets and validation passes.

1. UI/UX Critic Rules (critique_ui)

  • BASE-UI-101: 8px Harmonic Spacing Grid — All margins, paddings, and gaps must strictly adhere to the 4px/8px design system token scale. Rejects arbitrary pixel escapes like p-[17px].

  • BASE-UI-102: Typography Scale Floor — Body text must never fall below 12px / 0.75rem. Headings must strictly follow modular scales ($1.250$ Major Third).

  • BASE-UI-103: WCAG AA Minimum Contrast Floor — Body copy must maintain $\ge 4.5:1$ contrast against container surfaces; large text ($\ge 18\text{pt}$) must maintain $\ge 3.0:1$.

  • BASE-UI-104: Single Optical Focal Point — Exactly one primary high-contrast CTA element per screen viewport to eliminate visual friction.

  • BASE-UI-105: Decorative Ceiling Policy — Hard cap of $\le 2$ decorative accents (gradients, drop shadows, ambient blurs) per view.

2. Backend Architecture Guard Rules (critique_backend)

  • BASE-BE-101: Zero Hardcoded Secrets (OWASP) — Prevents any raw API keys, bearer tokens, or private JWT secrets in source code.

  • BASE-BE-102: Boundary Schema Validation — All external inputs (req.body, req.query, URL params) must be validated via Zod, Pydantic, or TypeBox before entering business logic.

  • BASE-BE-103: Sanitized Error Dumps — Rejects raw stack trace exposure (err.stack, database errors) in HTTP responses.

  • BASE-BE-104: No Orphan Logic Paths — All switch/conditional branches and Promise chains must define explicit catch and fallback terminations.

  • BASE-BE-105: Default Request Timeout & Rate Limits — All outbound network calls (fetch, axios) must declare explicit AbortSignal.timeout(ms) configurations.


šŸ”Œ Multi-Tool Adapter Ecosystem

Atelier provides single-command drop-in adapters for all leading agentic IDEs, with continuous integration drift checking to ensure zero divergence from the canonical ruleset.


⚔ Quickstart & Installation

Option A: One-Liner (Install Quality Gate Rules into any Project)

Run anywhere in your project directory:

npx -y github:anshrajore/atelier-mcp install

Installs .cursorrules, .windsurfrules, CLAUDE.md, .github/copilot-instructions.md, and .agents/rules/atelier.md in one command with zero setup.


Option B: Clone & Build the Local MCP Server

git clone https://github.com/anshrajore/atelier-mcp.git
cd atelier-mcp

# Install dependencies and build TypeScript server
npm install
npm run build

2. Configure Your IDE / MCP Client

Add Atelier to your MCP client configuration:

For Cursor (~/.cursor/mcp.json or Project Settings)

{
  "mcpServers": {
    "atelier": {
      "command": "node",
      "args": ["/absolute/path/to/atelier-mcp/mcp-server/dist/index.js"],
      "env": {
        "ATELIER_LLM_PROVIDER": "heuristic"
      }
    }
  }
}

For Claude Desktop (claude_desktop_config.json)

{
  "mcpServers": {
    "atelier": {
      "command": "node",
      "args": ["/absolute/path/to/atelier-mcp/mcp-server/dist/index.js"]
    }
  }
}

For Antigravity / OpenCode

{
  "mcpServers": {
    "atelier": {
      "command": "node",
      "args": ["/absolute/path/to/atelier-mcp/mcp-server/dist/index.js"]
    }
  }
}

3. Deploy IDE Quality Gate Rules

Copy the synchronized adapter files into your project root:

# Cursor IDE
cp adapters/.cursorrules ./
cp -r adapters/.cursor ./

# Windsurf IDE
cp adapters/.windsurfrules ./

# Claude Code CLI
cp adapters/CLAUDE.md ./

# Antigravity / Agent Rules
mkdir -p .agents/rules
cp adapters/.agents/rules/atelier.md .agents/rules/

# GitHub Copilot
mkdir -p .github
cp adapters/.github/copilot-instructions.md .github/

Verify all adapters are in sync:

npm run check-sync

šŸ› ļø MCP Tool Reference

Atelier exposes three core MCP tools to connected AI agents:

1. critique_ui

Audits React, Next.js, HTML, and Tailwind CSS code for design system compliance.

{
  "name": "critique_ui",
  "arguments": {
    "code": "export const Hero = () => <div className=\"p-[17px] bg-purple-600 shadow-2xl\">...</div>",
    "framework": "nextjs-tailwind"
  }
}

2. critique_backend

Audits TypeScript, Node.js, Express, and n8n workflows for architectural soundness.

{
  "name": "critique_backend",
  "arguments": {
    "code": "app.post('/api/pay', (req, res) => { const secret = 'sk_live_99881122'; ... });",
    "framework": "general"
  }
}

3. generate_fix

Automatically applies the proposed diff patches to resolve all identified violations.


🧠 Distillation Pipeline & Fine-Tuning

Atelier includes an autonomous synthetic dataset generation and distillation harness:

# 1. Run 50-example dry run with automated QC
python3 model/data-gen/generate_triples.py --dry-run

# 2. Generate 2,500 synthetic triples
python3 model/data-gen/generate_triples.py --count 2500

# 3. Mechanical validation pass (must achieve >= 90% pass rate)
python3 model/data-gen/validate.py

# 4. Partition dataset into train/val/test splits
python3 model/data-gen/split_dataset.py

Fine-Tuning Execution Options

  • Apple Silicon (Local MLX): python3 -m mlx_lm.lora -c model/train/config_mlx.yaml

  • Google Colab: Open model/train/atelier_train_colab.ipynb on an A100 GPU.

  • RunPod (Cloud GPU): Execute bash model/train/run_runpod.sh.


šŸ“‚ Repository Structure

atelier/
ā”œā”€ā”€ docs/
│   └── PROJECT_MAP.md             # Master canonical system specification
ā”œā”€ā”€ skills/
│   └── atelier/
│       ā”œā”€ā”€ SKILL.md               # Universal principles & mechanical checks
│       └── presets/
│           ā”œā”€ā”€ nextjs-tailwind.md # Next.js & Tailwind CSS rules
│           └── n8n.md             # n8n workflow graph rules
ā”œā”€ā”€ mcp-server/                    # TypeScript MCP server exposing critics
ā”œā”€ā”€ adapters/                      # Pre-configured adapters (Cursor, Windsurf, etc.)
ā”œā”€ā”€ model/
│   ā”œā”€ā”€ data-gen/                  # Triple generation & mechanical QC validation
│   ā”œā”€ā”€ dataset/                   # Stratified JSONL splits (train, val, test)
│   ā”œā”€ā”€ train/                     # MLX, PyTorch, Colab, and RunPod training packs
│   └── eval/                      # Evaluation harness & benchmark scoreboard
ā”œā”€ā”€ benchmarks/
│   └── SCOREBOARD.md              # Real precision, recall, cost & latency metrics
ā”œā”€ā”€ assets/                        # High-contrast monochrome SVG visual system
ā”œā”€ā”€ CONTRIBUTING.md                # Rule & preset contribution guidelines
ā”œā”€ā”€ LICENSE                        # MIT License
└── README.md                      # Canonical public documentation

šŸ¤ Contributing

We welcome contributions of new framework presets (e.g. SvelteKit, FastAPI, Flutter) and additional mechanical rules. Please read CONTRIBUTING.md for guidelines on formatting check: labeling functions.


šŸ“„ License & Credits

Available Tools

3 tools
critique_backendC

Audit backend and pipeline code against OWASP security, 12-factor principles, N+1 query elimination, input boundary validation, and zero-orphan workflow node rules.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesThe backend endpoint, service logic, or workflow orchestration JSON (e.g. n8n, LangGraph).
filePathNoOptional path of the backend file or workflow graph.
languageNoProgramming language of the code snippet.
frameworkNoBackend framework or workflow orchestrator.
isWorkflowJsonNoSet to true if evaluating an orchestration graph JSON payload.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not state whether the audit is read-only, requires network/database access, authenticates to external services, has rate limits, or what kind of output (scores, findings, severity levels) is returned. It also doesn't clarify whether findings are deterministic or LLM-judgment based.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that lists the audit dimensions without filler. It is efficient, though the long enumeration could be structured as a list for readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the audit dimensions and the schema covers all parameters, but with no annotations and no output schema, it omits important context: what the tool returns (findings format, severity), whether it mutates anything, and when to choose it over critique_ui. Adequate but with clear gaps for a multi-parameter audit tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents code, filePath, language, framework, and isWorkflowJson fully. The description adds no syntax, format, or usage guidance beyond what the schema provides, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (audit) and resource (backend and pipeline code), and it enumerates the rule sets checked (OWASP, 12-factor, N+1, input boundary validation, zero-orphan workflow nodes). It is clear about what the tool does, though it does not distinguish itself from critique_ui or explain the relationship to generate_fix.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, prerequisites, or alternatives named. The agent must infer that this tool is for backend code while critique_ui handles UI, and that generate_fix is separate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

critique_uiA

Audit UI code and rendered screenshots against design system guidelines (spacing scales, typography hierarchy, contrast/WCAG AA, and anti-AI-clichƩ patterns). Returns structured findings citing canonical SKILL.md rules.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesThe frontend component or page source code (React, Vue, HTML, Svelte, Tailwind).
filePathNoOptional path of the file being audited.
frameworkNoThe frontend framework used.
designSystemNoTarget design system convention.
screenshotPathNoOptional absolute path to the screenshot image file.
screenshotBase64NoOptional base64-encoded screenshot of the rendered component for visual craft inspection.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the output shape (structured findings citing SKILL.md rules), which is genuinely useful behavioral context, but it says nothing about permissions, whether it mutates anything, failure modes, or what happens when screenshots are omitted – significant gaps for an audit tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly-packed sentences with zero filler, front-loaded with the audit action and the criteria. It could be slightly more informative about output format, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, criteria, and output type for a tool with no annotations and no output schema. However, for a 6-parameter audit tool, it omits guidance on when screenshots are required versus code-only, how the required `code` parameter interacts with optional visual inputs, and any behavioral constraints, leaving meaningful gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all six parameters and both enums, establishing a baseline of 3. The description adds only high-level framing (what the audit checks) rather than meaning about specific parameters, so it does not exceed the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (audit) and resources (UI code, rendered screenshots) and enumerates the exact criteria checked: spacing scales, typography hierarchy, contrast/WCAG AA, anti-AI-clichƩ patterns. This is clearly distinguished from sibling critique_backend, which would cover non-UI concerns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes clear this is a UI-design audit, which implies usage, but it never states when to choose this over critique_backend or generate_fix, nor any preconditions (e.g., that screenshotPath or screenshotBase64 is needed for visual inspection). No explicit when-to-use or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_fixA

Given source code and structured findings from critique_ui or critique_backend, generate an updated, compliant code patch and diff.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesThe original code snippet.
criticYesThe critic type that generated the findings.
filePathNoOptional file path.
findingsYesThe array of CritiqueFinding objects returned by critique_ui or critique_backend.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose the output form (a code patch and diff), which is useful, but it is silent on whether anything is written to disk, whether filePath causes a mutation, permission requirements, or reversibility.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that names the inputs, the producing tools, and the output artifact with no filler. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, describing that the tool returns a patch and diff is genuinely valuable, and the input contract is fully covered by the schema. It stops short of clarifying whether filePath triggers a file write or whether the fix is applied automatically.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already explains code, critic, filePath, and findings. The description reinforces the pairing between findings and the critic type but adds no formatting or syntactic detail beyond the schema; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (generate), a specific artifact (updated, compliant code patch and diff), and names its inputs precisely: source code plus structured findings. It is clearly distinguishable from critique_ui/critique_backend, which produce findings rather than fixes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'structured findings from critique_ui or critique_backend' establishes the prerequisite workflow step and names the alternatives explicitly, so an agent knows this tool runs downstream of the critique tools. It does not, however, state any when-not condition or what to do if findings are absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.0.2
    • First observedcritique_backend
    • First observedcritique_ui
    • First observedgenerate_fix

TDQS

A3.6/5.0

Scored across 3 tools

Disambiguation5/5

critique_ui and critique_backend target clearly separate domains (front-end design vs backend security), and generate_fix is a distinct action consuming the critique outputs. No overlap in purpose between any of the three.

Naming Consistency5/5

All three tools follow a consistent verb_noun snake_case pattern (critique_ui, critique_backend, generate_fix). The convention is predictable and readable.

Tool Count4/5

Three tools is on the lean side but each earns its place in a coherent audit-then-fix workflow. One or two more (e.g. apply/verify) would round out the set.

Completeness3/5

The critique-and-generate cycle is covered, but there is no way to apply the generated patch, re-critique a fix, or browse the referenced SKILL.md rules. These gaps force the agent to hand off the workflow's final steps.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    Automatically enhances developer prompts with quality requirements, codebase context, and architectural patterns, then orchestrates other MCP servers to ensure AI coding assistants produce high-quality, structured code that follows best practices and security standards.
    7
    3
    MIT
  • -
    license
    Not graded
    quality
    Not graded
    maintenance
    A universal MCP server that acts as a code quality gate for AI assistants, providing pre-generation guidance, post-generation review, and root cause analysis to improve code quality.
    -
  • A
    license
    Not graded
    quality
    D
    maintenance
    Verdict MCP is a gatekeeper server that enforces code completeness, test coverage, and premium UI/UX standards on AI agents by auditing code, validating design tokens, and running tests with 95% coverage and 80% mutation score gates.
    1
    MIT