atelier-mcp
Audits Express applications for backend architectural soundness, including schema validation and secure error responses.
Audits LangGraph pipelines for architectural soundness, detecting orphan nodes and ensuring explicit error handling.
Audits n8n workflows for architectural soundness, detecting orphan nodes and ensuring explicit error handling.
Audits Next.js code for UI design system compliance, including spacing, typography, contrast, and decorative element limits.
Audits Node.js code for backend architectural soundness, including secret handling, schema validation, and error handling.
Audits React code for UI design system compliance, including spacing, typography, contrast, and decorative element limits.
Audits Tailwind CSS code for design token adherence and UI quality.
Audits TypeScript code for backend architectural soundness, including secret handling, schema validation, and error handling.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@atelier-mcpCritique the UI of my generated landing page"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ā” 1-Liner Quickstart (Install into Any Project)
npx -y atelier-quality-gate installInstalls .cursorrules, .windsurfrules, CLAUDE.md, .github/copilot-instructions.md, and .agents/rules/atelier.md in 1 second with zero configuration.
Related MCP server: Deslopify
šļø Author & System Philosophy
Atelier is architected and developed by Ansh Rajore.
When coding with modern AI assistants (Cursor, Windsurf, Claude Code, Antigravity, GitHub Copilot), generations chronically regress toward two critical failure modes:
The "Generic AI UI": Arbitrary purple-on-dark glow palettes, uncalibrated pixel-pushing (
p-[19px],mt-[13px]), rainbow gradient text clips, decorative pulsing pill badges, and nested Russian-doll cards.Fragile Backend Architecture: Hardcoded secrets/JWTs, unbounded database queries, missing boundary schema validation (Zod/Pydantic), unsanitized stack trace dumps, and disconnected/orphan nodes in orchestration pipelines (n8n, LangGraph).
Why Ponytail-Style Rulesets Fail
Existing tools (like Ponytail) attempt to solve code quality through a single static prompt injected before generation. In rigorous benchmark tests, pre-generation prompts only catch 15.4% of violations because LLMs prioritize completion structure over negative constraints during code emission.
Atelier introduces a fundamentally superior architecture: two specialist critic agents that execute after generation with mechanical pass/fail verification:
āļø Architectural Comparison: Atelier vs. Ponytail
Dimension | Ponytail (Static Ruleset) | Atelier Quality Gate |
Inspection Timing | Pre-generation prompt injection only | Post-generation inspection & repair gate |
Domain Coverage | Code minimalism & YAGNI only | UI/UX Design Systems + Backend Architecture |
Verification Logic | Subjective guidelines ("write clean code") | 100% mechanically gradeable ( |
Shipped Model | Zero model (prompt only) | Fine-tuned open-weight model + GGUF + API fallback |
Tool Integration | Static file copies | Live MCP server ( |
Overall Violation Recall | 15.4% | 92.1% (Local 7B) / 100.0% (Static Engine) |
Inference Cost | $0.00 | $0.00 (Zero Marginal Cost Locally) |
š Benchmark Scoreboard
Rigorous Empirical Results (36 Gradeable Rules)
Architecture / Model | Mode | UI/UX Recall | Backend Recall | Overall Recall | Precision | Cost / 1k Evals | P95 Latency |
Vanilla AI Agent (GPT-4o / Sonnet) | No Critic Gate | 0.0% | 0.0% | 0.0% | N/A | $0.00 | N/A |
Ponytail (Ruleset only) | Static Pre-Prompt | 12.5% | 20.0% | 15.4% | 66.7% | $0.00 | N/A |
Atelier Frontier Teacher (Claude 3.5 Sonnet) | Cloud API Critic | 96.2% | 95.0% | 95.7% | 94.8% | $14.20 | 1,450 ms |
Atelier Fine-Tuned (Qwen2.5-Coder-7B LoRA) | Local Self-Hosted (GGUF) | 92.4% | 91.8% | 92.1% | 93.5% | $0.00 | 180 ms |
Atelier Heuristics Engine | Zero-Dep Static Engine | 100.0% | 100.0% | 100.0% | 81.8% | $0.00 | 12 ms |
š Two-Agent Ruleset & Mechanical Check Matrix
Every rule in Atelier contains an unambiguous mechanical test (check:), which acts as a deterministic labeling function for downstream fine-tuning datasets and validation passes.
1. UI/UX Critic Rules (critique_ui)
BASE-UI-101: 8px Harmonic Spacing Gridā All margins, paddings, and gaps must strictly adhere to the 4px/8px design system token scale. Rejects arbitrary pixel escapes likep-[17px].BASE-UI-102: Typography Scale Floorā Body text must never fall below 12px / 0.75rem. Headings must strictly follow modular scales ($1.250$ Major Third).BASE-UI-103: WCAG AA Minimum Contrast Floorā Body copy must maintain $\ge 4.5:1$ contrast against container surfaces; large text ($\ge 18\text{pt}$) must maintain $\ge 3.0:1$.BASE-UI-104: Single Optical Focal Pointā Exactly one primary high-contrast CTA element per screen viewport to eliminate visual friction.BASE-UI-105: Decorative Ceiling Policyā Hard cap of $\le 2$ decorative accents (gradients, drop shadows, ambient blurs) per view.
2. Backend Architecture Guard Rules (critique_backend)
BASE-BE-101: Zero Hardcoded Secrets (OWASP)ā Prevents any raw API keys, bearer tokens, or private JWT secrets in source code.BASE-BE-102: Boundary Schema Validationā All external inputs (req.body,req.query, URL params) must be validated via Zod, Pydantic, or TypeBox before entering business logic.BASE-BE-103: Sanitized Error Dumpsā Rejects raw stack trace exposure (err.stack, database errors) in HTTP responses.BASE-BE-104: No Orphan Logic Pathsā All switch/conditional branches and Promise chains must define explicit catch and fallback terminations.BASE-BE-105: Default Request Timeout & Rate Limitsā All outbound network calls (fetch,axios) must declare explicitAbortSignal.timeout(ms)configurations.
š Multi-Tool Adapter Ecosystem
Atelier provides single-command drop-in adapters for all leading agentic IDEs, with continuous integration drift checking to ensure zero divergence from the canonical ruleset.
ā” Quickstart & Installation
Option A: One-Liner (Install Quality Gate Rules into any Project)
Run anywhere in your project directory:
npx -y github:anshrajore/atelier-mcp installInstalls .cursorrules, .windsurfrules, CLAUDE.md, .github/copilot-instructions.md, and .agents/rules/atelier.md in one command with zero setup.
Option B: Clone & Build the Local MCP Server
git clone https://github.com/anshrajore/atelier-mcp.git
cd atelier-mcp
# Install dependencies and build TypeScript server
npm install
npm run build2. Configure Your IDE / MCP Client
Add Atelier to your MCP client configuration:
For Cursor (~/.cursor/mcp.json or Project Settings)
{
"mcpServers": {
"atelier": {
"command": "node",
"args": ["/absolute/path/to/atelier-mcp/mcp-server/dist/index.js"],
"env": {
"ATELIER_LLM_PROVIDER": "heuristic"
}
}
}
}For Claude Desktop (claude_desktop_config.json)
{
"mcpServers": {
"atelier": {
"command": "node",
"args": ["/absolute/path/to/atelier-mcp/mcp-server/dist/index.js"]
}
}
}For Antigravity / OpenCode
{
"mcpServers": {
"atelier": {
"command": "node",
"args": ["/absolute/path/to/atelier-mcp/mcp-server/dist/index.js"]
}
}
}3. Deploy IDE Quality Gate Rules
Copy the synchronized adapter files into your project root:
# Cursor IDE
cp adapters/.cursorrules ./
cp -r adapters/.cursor ./
# Windsurf IDE
cp adapters/.windsurfrules ./
# Claude Code CLI
cp adapters/CLAUDE.md ./
# Antigravity / Agent Rules
mkdir -p .agents/rules
cp adapters/.agents/rules/atelier.md .agents/rules/
# GitHub Copilot
mkdir -p .github
cp adapters/.github/copilot-instructions.md .github/Verify all adapters are in sync:
npm run check-syncš ļø MCP Tool Reference
Atelier exposes three core MCP tools to connected AI agents:
1. critique_ui
Audits React, Next.js, HTML, and Tailwind CSS code for design system compliance.
{
"name": "critique_ui",
"arguments": {
"code": "export const Hero = () => <div className=\"p-[17px] bg-purple-600 shadow-2xl\">...</div>",
"framework": "nextjs-tailwind"
}
}2. critique_backend
Audits TypeScript, Node.js, Express, and n8n workflows for architectural soundness.
{
"name": "critique_backend",
"arguments": {
"code": "app.post('/api/pay', (req, res) => { const secret = 'sk_live_99881122'; ... });",
"framework": "general"
}
}3. generate_fix
Automatically applies the proposed diff patches to resolve all identified violations.
š§ Distillation Pipeline & Fine-Tuning
Atelier includes an autonomous synthetic dataset generation and distillation harness:
# 1. Run 50-example dry run with automated QC
python3 model/data-gen/generate_triples.py --dry-run
# 2. Generate 2,500 synthetic triples
python3 model/data-gen/generate_triples.py --count 2500
# 3. Mechanical validation pass (must achieve >= 90% pass rate)
python3 model/data-gen/validate.py
# 4. Partition dataset into train/val/test splits
python3 model/data-gen/split_dataset.pyFine-Tuning Execution Options
Apple Silicon (Local MLX):
python3 -m mlx_lm.lora -c model/train/config_mlx.yamlGoogle Colab: Open
model/train/atelier_train_colab.ipynbon an A100 GPU.RunPod (Cloud GPU): Execute
bash model/train/run_runpod.sh.
š Repository Structure
atelier/
āāā docs/
ā āāā PROJECT_MAP.md # Master canonical system specification
āāā skills/
ā āāā atelier/
ā āāā SKILL.md # Universal principles & mechanical checks
ā āāā presets/
ā āāā nextjs-tailwind.md # Next.js & Tailwind CSS rules
ā āāā n8n.md # n8n workflow graph rules
āāā mcp-server/ # TypeScript MCP server exposing critics
āāā adapters/ # Pre-configured adapters (Cursor, Windsurf, etc.)
āāā model/
ā āāā data-gen/ # Triple generation & mechanical QC validation
ā āāā dataset/ # Stratified JSONL splits (train, val, test)
ā āāā train/ # MLX, PyTorch, Colab, and RunPod training packs
ā āāā eval/ # Evaluation harness & benchmark scoreboard
āāā benchmarks/
ā āāā SCOREBOARD.md # Real precision, recall, cost & latency metrics
āāā assets/ # High-contrast monochrome SVG visual system
āāā CONTRIBUTING.md # Rule & preset contribution guidelines
āāā LICENSE # MIT License
āāā README.md # Canonical public documentationš¤ Contributing
We welcome contributions of new framework presets (e.g. SvelteKit, FastAPI, Flutter) and additional mechanical rules. Please read CONTRIBUTING.md for guidelines on formatting check: labeling functions.
š License & Credits
License: MIT License ā see LICENSE for details.
Architect & Developer: Ansh Rajore.
Available Tools
3 toolscritique_backendC
Audit backend and pipeline code against OWASP security, 12-factor principles, N+1 query elimination, input boundary validation, and zero-orphan workflow node rules.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | The backend endpoint, service logic, or workflow orchestration JSON (e.g. n8n, LangGraph). | |
| filePath | No | Optional path of the backend file or workflow graph. | |
| language | No | Programming language of the code snippet. | |
| framework | No | Backend framework or workflow orchestrator. | |
| isWorkflowJson | No | Set to true if evaluating an orchestration graph JSON payload. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does not state whether the audit is read-only, requires network/database access, authenticates to external services, has rate limits, or what kind of output (scores, findings, severity levels) is returned. It also doesn't clarify whether findings are deterministic or LLM-judgment based.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that lists the audit dimensions without filler. It is efficient, though the long enumeration could be structured as a list for readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the audit dimensions and the schema covers all parameters, but with no annotations and no output schema, it omits important context: what the tool returns (findings format, severity), whether it mutates anything, and when to choose it over critique_ui. Adequate but with clear gaps for a multi-parameter audit tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents code, filePath, language, framework, and isWorkflowJson fully. The description adds no syntax, format, or usage guidance beyond what the schema provides, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (audit) and resource (backend and pipeline code), and it enumerates the rule sets checked (OWASP, 12-factor, N+1, input boundary validation, zero-orphan workflow nodes). It is clear about what the tool does, though it does not distinguish itself from critique_ui or explain the relationship to generate_fix.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, prerequisites, or alternatives named. The agent must infer that this tool is for backend code while critique_ui handles UI, and that generate_fix is separate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
critique_uiA
Audit UI code and rendered screenshots against design system guidelines (spacing scales, typography hierarchy, contrast/WCAG AA, and anti-AI-clichƩ patterns). Returns structured findings citing canonical SKILL.md rules.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | The frontend component or page source code (React, Vue, HTML, Svelte, Tailwind). | |
| filePath | No | Optional path of the file being audited. | |
| framework | No | The frontend framework used. | |
| designSystem | No | Target design system convention. | |
| screenshotPath | No | Optional absolute path to the screenshot image file. | |
| screenshotBase64 | No | Optional base64-encoded screenshot of the rendered component for visual craft inspection. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the output shape (structured findings citing SKILL.md rules), which is genuinely useful behavioral context, but it says nothing about permissions, whether it mutates anything, failure modes, or what happens when screenshots are omitted ā significant gaps for an audit tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly-packed sentences with zero filler, front-loaded with the audit action and the criteria. It could be slightly more informative about output format, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, criteria, and output type for a tool with no annotations and no output schema. However, for a 6-parameter audit tool, it omits guidance on when screenshots are required versus code-only, how the required `code` parameter interacts with optional visual inputs, and any behavioral constraints, leaving meaningful gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all six parameters and both enums, establishing a baseline of 3. The description adds only high-level framing (what the audit checks) rather than meaning about specific parameters, so it does not exceed the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (audit) and resources (UI code, rendered screenshots) and enumerates the exact criteria checked: spacing scales, typography hierarchy, contrast/WCAG AA, anti-AI-clichƩ patterns. This is clearly distinguished from sibling critique_backend, which would cover non-UI concerns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes clear this is a UI-design audit, which implies usage, but it never states when to choose this over critique_backend or generate_fix, nor any preconditions (e.g., that screenshotPath or screenshotBase64 is needed for visual inspection). No explicit when-to-use or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_fixA
Given source code and structured findings from critique_ui or critique_backend, generate an updated, compliant code patch and diff.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | The original code snippet. | |
| critic | Yes | The critic type that generated the findings. | |
| filePath | No | Optional file path. | |
| findings | Yes | The array of CritiqueFinding objects returned by critique_ui or critique_backend. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does disclose the output form (a code patch and diff), which is useful, but it is silent on whether anything is written to disk, whether filePath causes a mutation, permission requirements, or reversibility.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that names the inputs, the producing tools, and the output artifact with no filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, describing that the tool returns a patch and diff is genuinely valuable, and the input contract is fully covered by the schema. It stops short of clarifying whether filePath triggers a file write or whether the fix is applied automatically.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already explains code, critic, filePath, and findings. The description reinforces the pairing between findings and the critic type but adds no formatting or syntactic detail beyond the schema; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (generate), a specific artifact (updated, compliant code patch and diff), and names its inputs precisely: source code plus structured findings. It is clearly distinguishable from critique_ui/critique_backend, which produce findings rather than fixes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'structured findings from critique_ui or critique_backend' establishes the prerequisite workflow step and names the alternatives explicitly, so an agent knows this tool runs downstream of the critique tools. It does not, however, state any when-not condition or what to do if findings are absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.2- First observed
critique_backend - First observed
critique_ui - First observed
generate_fix
TDQS
Scored across 3 tools
critique_ui and critique_backend target clearly separate domains (front-end design vs backend security), and generate_fix is a distinct action consuming the critique outputs. No overlap in purpose between any of the three.
All three tools follow a consistent verb_noun snake_case pattern (critique_ui, critique_backend, generate_fix). The convention is predictable and readable.
Three tools is on the lean side but each earns its place in a coherent audit-then-fix workflow. One or two more (e.g. apply/verify) would round out the set.
The critique-and-generate cycle is covered, but there is no way to apply the generated patch, re-critique a fix, or browse the referenced SKILL.md rules. These gaps force the agent to hand off the workflow's final steps.
Related MCP Connectors
MCP server for visual regression testing: triage a PR's UI diffs from your coding agent.
Official DevSpeak MCP server ā translate technical text into formal specs from any AI IDE or agent
Website QA for your coding agent: audit SEO, performance, security, accessibility over MCP.
MCP server for Mint ā AI-powered QA that runs your app in a real browser on every PR.
Related MCP Servers
- AlicenseAqualityCmaintenanceAutomatically enhances developer prompts with quality requirements, codebase context, and architectural patterns, then orchestrates other MCP servers to ensure AI coding assistants produce high-quality, structured code that follows best practices and security standards.73MIT
- -licenseNot gradedqualityNot gradedmaintenanceA universal MCP server that acts as a code quality gate for AI assistants, providing pre-generation guidance, post-generation review, and root cause analysis to improve code quality.-
- AlicenseNot gradedqualityDmaintenanceVerdict MCP is a gatekeeper server that enforces code completeness, test coverage, and premium UI/UX standards on AI agents by auditing code, validating design tokens, and running tests with 95% coverage and 80% mutation score gates.1MIT
- FlicenseAqualityDmaintenanceAn MCP server implementing a 7-stage agentic frontend workflowāfrom design audit to PR reviewāincluding AI-driven component generation, browser validation, E2E testing, and CI self-healing.410 npm-