PQS - Prompt Quality Score
OfficialPQS (Prompt Quality Score) is an MCP server that lets you score, optimize, and compare LLM prompts before sending them to any AI model.
Score any prompt for free (
score_prompt): Get a quality grade (A–F), a score out of 40, and a percentile ranking — no API key or payment required. Supports domain-specific scoring across verticals like software, crypto, education, business, content, science, and general.Optimize prompts (
optimize_prompt): For $0.025 USDC, receive a full 8-dimension quality breakdown (based on PEEM, RAGAS, G-Eval, and MT-Bench frameworks) plus an improved version of your prompt. Requires a free PQS API key.Compare Claude vs GPT-4o (
compare_models): For $0.50 USDC, run a head-to-head comparison judged by a third model, returning a winner, scores, and a recommendation on which model suits that prompt type. Requires a free PQS API key.Use as a quality gate: Block low-quality prompts (grade D or below, < 28/40) before they consume inference spend, ensuring only high-quality prompts reach your LLM.
Available as a GitHub Marketplace action (PQS Check) for integrating prompt quality scoring into GitHub workflows and CI/CD pipelines.
PQS MCP Server
Score prompt quality before it reaches any AI model. An MCP server for PQS.
Score and optimize LLM prompts before they hit any model. Built on PEEM, RAGAS, MT-Bench, G-Eval, and ROUGE.
Install
Claude Desktop (stdio)
Add to your config (~/Library/Application Support/Claude/claude_desktop_config.json):
{
"mcpServers": {
"pqs": {
"command": "npx",
"args": ["-y", "pqs-mcp-server"]
}
}
}Remote (HTTP)
Use this when your MCP client supports streamable-HTTP transport (no local npm install required):
{
"mcpServers": {
"pqs": {
"url": "https://promptqualityscore.com/api/mcp"
}
}
}Smithery
smithery mcp add onchaintel/pqsRelated MCP server: aipaygen-mcp
Tools
score_prompt (Free, no API key required)
Returns a 0-80 score, A-F grade, full 8-dimension breakdown (clarity, specificity, context, constraints, output_format, role_definition, examples, cot_structure), and the weakest dimension. Rate-limited per IP: 5/min, 10/day, 100/month.
Low- and mid-band scores also include a structured suggestion field with a message, a next_tool pointer to optimize_prompt, and a subscribe URL the consuming LLM can paraphrase back to the user.
Example output (low-band score, suggestion attached):
{
"pqs_version": "2.0",
"prompt": "analyze this wallet",
"score": 9,
"out_of": 80,
"grade": "F",
"dimensions": {
"clarity": 2,
"specificity": 1,
"context": 1,
"constraints": 1,
"output_format": 1,
"role_definition": 1,
"examples": 1,
"cot_structure": 1
},
"weakest_dimension": "specificity",
"powered_by": "PQS — promptqualityscore.com",
"suggestion": {
"message": "This prompt scored 9/80 (F) — significant room to improve. The optimize_prompt tool rewrites it and shows side-by-side outputs from a frontier model, so you can see the impact. optimize_prompt is part of PQS Pro ($19.99/mo, 1,000 calls/mo). Subscribe at https://promptqualityscore.com/pricing?utm_source=mcp&utm_medium=suggestion_v140&utm_campaign=2026-05-mcp-tools-v140.",
"next_tool": "optimize_prompt",
"subscribe_url": "https://promptqualityscore.com/pricing?utm_source=mcp&utm_medium=suggestion_v140&utm_campaign=2026-05-mcp-tools-v140"
}
}If the per-IP rate limit is hit, the response is a structured rate_limit_exceeded payload with subscribe and account URLs.
optimize_prompt (Pro subscription required)
Rewrites a prompt to score higher and runs both versions through a frontier model so the user can see the before/after output. Returns the optimized prompt, before/after dimension scores (with totals), improvement_pct, and side-by-side sample outputs.
Pro subscription required ($19.99/mo, 1,000 calls/mo, includes batch + A/B comparison). Subscribe at promptqualityscore.com/pricing.
If the API key is missing, invalid, or on the Free tier, the tool returns a structured error pointing the user at the right URL. No silent failures, no inventing keys. Errors emitted:
api_key_required: noapi_keyargument was sentapi_key_invalid: key not recognizedsubscription_required: key is valid but Free tier (subscribe to upgrade)rate_limited: per-minute burst limit reached (Pro is rate-limited per minute, not per month) or temporary upstream capacity issueservice_unavailable: upstream 5xx
Quality Gate Pattern
Use PQS as a pre-inference quality gate:
const score = await fetch("https://promptqualityscore.com/api/score/free", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ prompt: userPrompt })
});
const { score: pqsScore } = await score.json();
if (pqsScore < 56) throw new Error("Prompt quality too low. Improve and retry.");Grade D or below (under 56/80) means the prompt will waste inference spend.
x402 (legacy pay-per-call)
The MCP tools in this package use the SaaS API-key model. A separate x402-native pay-per-call path is available via the canonical PQS HTTP API (no API key, caller settles USDC on Base on-chain). For x402 integration, see the canonical pricing and discovery artifacts at promptqualityscore.com.
Self-hosting
Override the PQS backend URL with the PQS_BASE environment variable:
PQS_BASE=https://your-pqs-host.example.com npx pqs-mcp-serverDefaults to https://promptqualityscore.com.
Built by
OnChainIntel, @OnChainAIIntel promptqualityscore.com
Available Tools
2 toolsoptimize_promptAInspect
Rewrite a prompt to score higher on the PQS rubric, AND show before/after output comparisons so the user can see the impact. Returns the optimized prompt, the original PQS score, the optimized PQS score, and side-by-side sample outputs from a frontier model using both versions.
USE WHEN:
The user got a low score from score_prompt and asks how to improve.
The user explicitly asks to "improve" / "rewrite" / "fix" / "optimize" a prompt they pasted.
The user is dissatisfied with output quality from a previous prompt and asks how to get better results.
score_prompt returned a suggestion to invoke this tool.
DO NOT USE WHEN:
The user just asked for a score (use score_prompt only — don't double up).
The user wants you to write a new prompt from scratch (write it directly).
REQUIRES: A PQS API key from a Pro subscription ($19.99/month, 1,000 calls/mo, includes batch + A/B comparison). If the user has not provided one, the tool returns a clear subscription URL — pass that response to the user verbatim. Do not invent or guess API keys. There is no free trial of this tool; the user must subscribe before the first call.
COST: Counted against your Pro subscription's monthly call quota.
LATENCY: ~6-8 seconds.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | The prompt to optimize. Max 8000 characters. | |
| api_key | No | PQS API key from a Pro subscription. Required. Format: pqs_live_… (32+ characters). Subscribe at https://promptqualityscore.com/pricing?utm_source=mcp&utm_medium=schema_description_v140&utm_campaign=2026-05-mcp-tools-v140 if you don't have one, or look up an existing key at https://promptqualityscore.com/account?utm_source=mcp&utm_medium=schema_description_v140&utm_campaign=2026-05-mcp-tools-v140. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite no annotations, the description thoroughly discloses behavior: it rewrites and shows comparisons, returns scores and outputs, requires an API key with subscription details, cost against quota, and latency of 6-8 seconds. Also explains the fallback when no API key is provided.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with sections and bullet points, and the main purpose is front-loaded. However, it could be slightly more concise; some repetition exists, such as 'score higher on the PQS rubric' appearing multiple times.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Comprehensive: explains return values despite no output schema, covers prerequisites, cost, latency, and when to use sibling. The tool's purpose and constraints are fully captured, leaving no significant gaps for an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with both parameters described. The description adds context: the API key format, subscription URL, and behavior when key is missing. This goes beyond the schema but is not essential since schema already covers basics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool rewrites a prompt to score higher on the PQS rubric and provides before/after comparisons. It distinguishes from the sibling 'score_prompt' by explicitly noting that 'score_prompt only scores' and this tool optimizes and compares.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit 'USE WHEN' and 'DO NOT USE WHEN' sections, listing specific conditions such as when the user asks for improvement or when score_prompt suggests it. Also clarifies when not to use it, e.g., if only a score is requested.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
score_promptAInspect
Score a prompt's quality across 8 dimensions BEFORE sending it to an expensive model. Returns a 0-80 score, an A-F grade, the per-dimension breakdown (clarity, specificity, context, constraints, output_format, role_definition, examples, cot_structure), and the weakest dimension.
USE WHEN:
The user is workshopping a prompt and asks "is this good?" / "will this work?" / "should I add more detail?"
The user is about to send a long or expensive prompt to GPT-4, Claude Opus, or any frontier model, especially in a batch or automation context where rework is costly.
The user mentions iterating on a prompt that produced poor output and wants to diagnose what's missing.
The user pastes a prompt and asks for feedback on it.
DO NOT USE WHEN:
The user is asking you to write a prompt for them (write it yourself first, then optionally call score_prompt to verify).
The prompt is conversational chat (this scores task-shaped prompts).
COST: Free, no API key required. Rate-limited per IP: 5/min, 10/day, 100/month. If the user exceeds the limit, the response will include a structured upgrade path with subscribe and account URLs.
LATENCY: ~2 seconds.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | The prompt text to score. Single prompt, not a conversation. Max 8000 characters. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses cost (free), rate limits (5/min, 10/day, 100/month), latency (~2 seconds), and behavior when limits exceeded (structured upgrade path). No annotations provided, so description carries full burden and does so comprehensively.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Description is concise and well-structured: starts with purpose and return values, then lists usage guidelines, cost, latency. Every sentence adds value with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (1 parameter, no output schema), the description is complete. It explains return values, usage context, limitations, and behavior, fully informing an agent about when and how to use it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a description for the single 'prompt' parameter. The description adds extra context: specifies it must be a single prompt (not conversation) and max 8000 characters, which goes beyond the schema's description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool scores prompt quality across 8 dimensions, returns a score, grade, breakdown, and weakest dimension. It distinguishes from sibling 'optimize_prompt' by focusing on evaluation before sending to expensive models, not optimization.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit 'USE WHEN' and 'DO NOT USE WHEN' sections, listing specific scenarios like workshopping prompts, about to send expensive prompts, or iterating. Excludes conversational chat and prompt writing tasks, which helps an agent decide when to invoke.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
The two tools have clearly distinct purposes: one scores a prompt, the other optimizes it. Their use cases and prerequisites are well-defined, so an agent cannot confuse them.
Both tool names follow a consistent verb_noun pattern: 'score_prompt' and 'optimize_prompt'. The naming is uniform and predictable.
With only two tools, the set is narrowly scoped to prompt scoring and optimization. While this is appropriate for the domain, slightly more tools (e.g., subscription management or rubric access) could enhance completeness without bloat.
The server covers the core workflow of scoring and optimizing prompts. However, it lacks tools for subscription management, retrieving the rubric, or performing batch operations, which are minor gaps given the stated dependencies and cost structure.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
MCP server providing access to the Scorecard API to evaluate and optimize LLM systems.
Hosted MCP server for LLM cost estimation, model comparison, and budget-aware routing.
Paid remote MCP for LLM security scans, jailbreak checks, analytics, checkout, and readiness.
Free OpenAI-compatible inference with signed provenance receipts and 3 focused MCP tools.
Related MCP Servers
- AlicenseBqualityDmaintenanceDynamic MCP server — 30+ tools across fact verification, agent memory, Indian NLP, contract risk, security threat modelling, sales call intelligence and more. x402/USDC micropayments on Base.3316MIT
- AlicenseNot gradedqualityCmaintenance250+ AI-powered MCP tools: research, write, code, translate, scrape, sentiment, vision, RAG, agent memory, marketplace, trading signals, and more. 15 models across 7 providers. Pay-per-use via API key or x402 USDC micropayments.2MIT

Parlayofficial
AlicenseNot gradedqualityCmaintenanceRemote MCP server for prediction markets — search and compare live odds across Polymarket, Kalshi, and Limitless from Claude, ChatGPT, or Gemini. Six read-only tools, free tier available.6MIT- AlicenseAqualityDmaintenanceAn MCP server for deterministic prompt optimization in Claude Code. Score prompts across 7 quality dimensions, auto-select from 11 Anthropic techniques, and return a structural scaffold.1242MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/OnChainAIIntel/pqs-mcp-server'
If you have feedback or need assistance with the MCP directory API, please join our Discord server