SEO Crawler MCP
SEO Crawler MCP - Website Crawler & SEO Analyzer for LLMs
Crawl and analyse your website for errors and issues that probably affect your site's SEO
Quick Navigation
Installation | CLI mode | How to use | What gets detected | Data storage | Performance | Tools reference | Available queries
I wanted to build on my experience working with the MCP protocol SDK to see just how far we can extend an AI assistant's capabilities. I decided that I'd quite like to build a crawler to check my site's "technical SEO" health and came across Crawlee - which seemed like the ideal library to base the crawl component of my MCP.
What's interesting is that MCP usually indicates a server connection of some sort. This is not so with SEO Crawler MCP. The MCP protocol is probably more powerful than I realised - this is a self-contained application wrapped in the MCP SDK that handles everything locally:
Smart request scheduling and queue management
Automatic retry logic and error handling
Respectful crawling with configurable delays
Memory-efficient streaming for large sites
Better-SQLite3 embedded database storing every crawled page's HTML, metadata, headers, link relationships, and site structure
Custom SQL analysis engine with 25+ specialised queries detecting content issues, technical SEO problems, security vulnerabilities, and optimisation opportunities
Claude (or your AI assistant of choice) can orchestrate this entire stack through simple function calls. The crawl runs asynchronously, stores everything in SQLite, and then Claude can query that data through natural language - "analyse this crawl for seo opportunities" or "report on internal broken links" - and the MCP server translates that into sophisticated SQL analysis.
You can also run crawls directly from the terminal - perfect for large sites or background processing. The CLI mode lets you run a crawl, get the output directory, and then hand that over to Claude for AI-powered analysis via the MCP tools.
Credits
The core crawling architecture is inspired by the logic and patterns from the LibreCrawl project. We've adapted their proven crawling methodology for use within the MCP protocol whilst adding comprehensive SEO analysis capabilities.
Installation
For Beginners
If you're new to MCP servers, I'd recommend reading these first:
I'd also suggest installing Desktop Commander first - it's useful for working with the crawl output files. See the Desktop Commander setup guide for details.
Quick Install (NPX)
Add this to your Claude Desktop config file:
Windows: C:\Users\[YourName]\AppData\Roaming\Claude\claude_desktop_config.json
Mac: ~/Library/Application Support/Claude/claude_desktop_config.json
{
"mcpServers": {
"seo-crawler-mcp": {
"command": "npx",
"args": ["-y", "@houtini/seo-crawler-mcp"],
"env": {
"OUTPUT_DIR": "C:\\seo-audits"
}
}
}
}Restart Claude Desktop. Four tools will be available:
seo-crawler-mcp:run_seo_auditseo-crawler-mcp:analyze_seoseo-crawler-mcp:query_seo_dataseo-crawler-mcp:list_seo_queries
Claude Code (CLI)
Claude Code uses a different registration mechanism -- it doesn't read claude_desktop_config.json. Use claude mcp add instead:
claude mcp add -e OUTPUT_DIR=/path/to/seo-audits -s user seo-crawler-mcp -- npx -y @houtini/seo-crawler-mcpVerify with:
claude mcp get seo-crawler-mcpYou should see Status: Connected.
Development Install
cd C:\MCP\seo-crawler-mcp
npm install
npm run buildThen use the local path in your config:
{
"mcpServers": {
"seo-crawler-mcp": {
"command": "node",
"args": ["C:\\MCP\\seo-crawler-mcp\\build\\index.js"],
"env": {
"OUTPUT_DIR": "C:\\seo-audits",
"DEBUG": "false"
}
}
}
}Environment Variables:
OUTPUT_DIR: Directory where crawl results are saved (required)DEBUG: Set to"true"to enable verbose debug logging (optional, default:"false")
CLI Usage for Local Development:
When running the CLI from a local build (not installed via npm), use node directly:
# Run crawl
node C:\MCP\seo-crawler-mcp\build\cli.js crawl https://example.com --max-pages=20
# Analyze results
node C:\MCP\seo-crawler-mcp\build\cli.js analyze C:\seo-audits\example.com_2026-02-02_abc123
# List queries
node C:\MCP\seo-crawler-mcp\build\cli.js queries --category=criticalRelated MCP server: seoagent
CLI Mode (Terminal Usage)
For large crawls or background processing, you can run crawls directly from the terminal.
Note: These examples use npx for globally installed packages. For local development, see the "Development Install" section above.
}
}
}
}
---
## CLI Mode (Terminal Usage)
For large crawls or background processing, you can run crawls directly from the terminal:
### Run a Crawl
```bash
# Basic crawl
npx @houtini/seo-crawler-mcp crawl https://example.com
# Large crawl with custom settings
npx @houtini/seo-crawler-mcp crawl https://example.com --max-pages=5000 --depth=5
# Using Googlebot user agent
npx @houtini/seo-crawler-mcp crawl https://example.com --user-agent=googlebotQuick Analysis
# Show summary statistics
npx @houtini/seo-crawler-mcp analyze C:/seo-audits/example.com_2026-02-01_abc123
# Detailed JSON output
npx @houtini/seo-crawler-mcp analyze C:/seo-audits/example.com_2026-02-01_abc123 --format=detailedList Available Queries
# All queries
npx @houtini/seo-crawler-mcp queries
# Security queries only
npx @houtini/seo-crawler-mcp queries --category=security
# Critical priority queries
npx @houtini/seo-crawler-mcp queries --priority=CRITICALWorkflow: Terminal + Claude
Run large crawl from terminal (runs in background, can close terminal)
npx @houtini/seo-crawler-mcp crawl https://bigsite.com --max-pages=5000Get the output path from the crawl results
Output Path: C:\seo-audits\bigsite.com_2026-02-02T10-15-30_abc123In Claude Desktop, analyze with AI
Analyze the crawl at C:\seo-audits\bigsite.com_2026-02-02T10-15-30_abc123 Show me the critical issues What are the biggest SEO problems? Give me a detailed report on broken internal links
This workflow is perfect for:
Large sites (1000+ pages) where you want the crawl to run overnight
Multiple sites you want to crawl in batch
Automated crawling via cron jobs or scheduled tasks
Keeping terminal-based workflow whilst using Claude for intelligent analysis
How to Use This
Complete SEO Audit
The typical workflow goes like this:
Crawl the website
Use seo-crawler-mcp to crawl https://example.com with maxPages=2000Run the analysis
Analyse the crawl at C:/seo-audits/example.com_2026-02-01_abc123Investigate specific issues
Show me the broken internal links from that crawl
Claude handles the rest - calling the right tools, parsing the results, and presenting everything in readable format.
Security Audit
If you're specifically worried about security headers:
List available security queries
What security checks can you run on an SEO crawl?Run security-focused analysis
Check the security issues in crawl C:/seo-audits/example.com_2026-02-01_abc123Deep dive on specific problems
Show me all pages with unsafe external links
What Gets Detected
The analysis engine includes 25 comprehensive SEO checks across five categories:
Critical Issues (4 checks)
Missing title tags - pages without titles don't rank
Broken internal links - 404/5xx responses that hurt crawlability
Server errors - 5xx responses indicating site problems
404 errors - broken pages that need fixing or redirecting
Content Quality (7 checks)
Duplicate titles across different pages
Duplicate meta descriptions
Duplicate H1 tags
Missing meta descriptions
Missing H1 tags
Multiple H1 tags on single pages
Thin content - pages under 300 words
Technical SEO (5 checks)
Redirect chains and loops
Orphan pages with no internal links
Canonical URL mismatches
Non-HTTPS pages still in use
Heading hierarchy problems (H3 before H2, etc)
Security (6 checks)
Missing Content-Security-Policy headers
Missing HSTS (Strict-Transport-Security)
Missing X-Frame-Options (clickjacking protection)
Missing Referrer-Policy
Unsafe external links (target="_blank" without rel="noopener")
Protocol-relative links (//example.com)
Optimisation (6 checks)
Title tags too long or too short
Meta descriptions length issues
Title matches H1 (opportunity for differentiation)
Pages with no outbound links
Pages with excessive external links
Pages missing images
Data Storage
The crawler stores everything in SQLite databases organised by domain and date:
C:/seo-audits/example.com_2026-02-01_abc123/
├── crawl-data.db # SQLite database
│ ├── pages # Every page crawled
│ ├── links # All link relationships
│ ├── errors # Crawl errors
│ └── crawl_metadata # Statistics
├── config.json # Crawl settings
└── crawl-export.csv # Optional CSV exportPerformance
Typical crawl performance metrics:
Crawl Speed:
Medium site (1,500-2,000 pages): ~15 minutes
300,000+ link relationships tracked
Database size: ~15MB for 2,000 pages
Query Performance:
Simple queries: under 10ms
Complex queries: under 100ms
Join queries: under 200ms
Full analysis: under 600ms
The SQLite approach works well here. Everything stays local, no API rate limits to worry about, and the query performance is more than adequate for SEO analysis.
Limitations
There are 4 additional checks planned for v3.0:
Core Web Vitals - requires Playwright for real browser metrics
Robots.txt validation - needs parser library
Readability scoring - requires text analysis library
Mobile rendering issues - needs device emulation
The current 25 checks cover the most critical aspects of technical SEO that directly impact search engine crawling, indexing, and ranking.
Technical Details
Built with:
TypeScript 5.3
Crawlee 3.7 (HttpCrawler)
better-sqlite3 12.6
Cheerio 1.0 (HTML parsing)
MCP SDK 1.0
The code uses ES modules throughout, with proper Zod validation on inputs and comprehensive error handling. I've kept the architecture clean - separate modules for crawling, analysis, formatting, and tool definitions.
Deployment:
Local MCP server via Node.js
No external dependencies
Configurable output directory
Concurrent crawling (5 workers)
MCP Tools Reference
run_seo_audit
Crawl a website and extract comprehensive SEO data into SQLite.
Parameters:
url(required) - Website URL to crawlmaxPages(optional) - Maximum pages to crawl (default: 1000)depth(optional) - Maximum crawl depth (default: 3)userAgent(optional) - "chrome" or "googlebot" (default: "chrome")
Example:
run_seo_audit({
url: "https://example.com",
maxPages: 2000,
depth: 5,
userAgent: "chrome"
})Returns: Crawl ID and output path
analyze_seo
Run comprehensive SEO analysis on a completed crawl.
Parameters:
crawlPath(required) - Path to crawl output directoryformat(optional) - "structured", "summary", or "detailed" (default: "structured")includeCategories(optional) - Filter by categories: "critical", "content", "technical", "security", "opportunities"maxExamplesPerIssue(optional) - Maximum example URLs per issue (default: 10)
Example:
analyze_seo({
crawlPath: "C:/seo-audits/example.com_2026-02-01_abc123",
format: "structured",
includeCategories: ["critical", "security"],
maxExamplesPerIssue: 5
})Returns: Structured report with issues, affected URLs, and fix recommendations
query_seo_data
Execute specific SEO queries by name.
Parameters:
crawlPath(required) - Path to crawl output directoryquery(required) - Query name (see list_seo_queries)limit(optional) - Maximum results (default: 100)
Example:
query_seo_data({
crawlPath: "C:/seo-audits/example.com_2026-02-01_abc123",
query: "broken-internal-links",
limit: 50
})Returns: Query results with affected URLs and context
list_seo_queries
Discover available SEO analysis queries.
Parameters:
category(optional) - Filter by categorypriority(optional) - Filter by priority level
Example:
list_seo_queries({
category: "security",
priority: "HIGH"
})Returns: List of available queries with descriptions and priorities
Available Queries
The analysis engine includes 28 predefined SQL queries organised by category. Each query includes detailed impact analysis and fix recommendations.
Critical Issues (4 queries)
missing-titles
What it finds: Pages without title tags
Why it matters: Title tags are the most important on-page SEO element. Without them, pages are essentially invisible to search engines.
Fix: Add unique, descriptive title tags (50-60 characters) to all pages immediately.
broken-internal-links
What it finds: Internal links pointing to 404/5xx error pages
Why it matters: Broken links hurt crawlability and waste crawl budget. They create dead ends for users and search engines.
Fix: Update or remove broken links. Add redirects for moved pages.
server-errors
What it finds: Pages returning 5xx status codes
Why it matters: Indicates server problems that prevent search engines from indexing content.
Fix: Investigate server issues, check error logs, ensure adequate resources.
not-found-errors
What it finds: Pages returning 404 status codes
Why it matters: Lost indexing opportunities and poor user experience.
Fix: Add 301 redirects or remove links to non-existent pages.
Content Quality (7 queries)
duplicate-titles
What it finds: Multiple pages sharing identical title tags
Why it matters: Confuses search engines about which page to rank for queries.
Fix: Make each page's title tag unique and descriptive of its specific content.
duplicate-descriptions
What it finds: Multiple pages with identical meta descriptions
Why it matters: Reduces click-through rates as snippets look identical in search results.
Fix: Write unique meta descriptions (150-160 characters) for each page.
duplicate-h1s
What it finds: Multiple pages sharing the same H1 heading
Why it matters: H1 tags signal page topic - duplicates dilute topical clarity.
Fix: Ensure each page has a unique H1 that accurately describes its content.
missing-descriptions
What it finds: Pages without meta description tags
Why it matters: Search engines create their own snippets, often poorly representing content.
Fix: Add compelling meta descriptions (150-160 characters) for all important pages.
missing-h1s
What it finds: Pages without H1 headings
Why it matters: H1 is a primary signal of page topic and structure.
Fix: Add descriptive H1 tags to all content pages.
multiple-h1s
What it finds: Pages with more than one H1 tag
Why it matters: Dilutes topical focus and confuses heading hierarchy.
Fix: Use only one H1 per page. Convert other H1s to H2 or H3.
thin-content
What it finds: Pages with less than 300 words of content
Why it matters: Thin content provides little value and ranks poorly.
Fix: Expand content with valuable information or consolidate into existing pages.
Technical SEO (5 queries)
redirect-pages
What it finds: Pages that redirect to other URLs
Why it matters: Multiple redirects waste crawl budget and slow page loads.
Fix: Update internal links to point directly to final destination.
redirect-chains
What it finds: URLs that redirect multiple times before reaching destination
Why it matters: Each redirect adds latency and risks breaking the chain.
Fix: Implement direct redirects from source to final destination.
orphan-pages
What it finds: Pages with no internal links pointing to them
Why it matters: Search engines may never discover orphan pages.
Fix: Add internal links from relevant pages to connect orphans to site structure.
canonical-issues
What it finds: Pages where canonical URL doesn't match actual URL
Why it matters: Signals duplicate content or indexing preference conflicts.
Fix: Ensure canonical tags point to the correct version of each page.
non-https-pages
What it finds: Pages still using HTTP instead of HTTPS
Why it matters: Security risk, ranking penalty, and browser warnings.
Fix: Implement HTTPS across entire site with proper redirects.
Security (6 queries)
missing-csp
What it finds: Pages without Content-Security-Policy headers
Why it matters: Vulnerability to XSS attacks and code injection.
Fix: Implement CSP headers to control resource loading.
missing-hsts
What it finds: Pages without Strict-Transport-Security headers
Why it matters: Allows protocol downgrade attacks.
Fix: Add HSTS headers to enforce HTTPS connections.
missing-x-frame-options
What it finds: Pages without X-Frame-Options headers
Why it matters: Vulnerability to clickjacking attacks.
Fix: Add X-Frame-Options headers (DENY or SAMEORIGIN).
missing-referrer-policy
What it finds: Pages without Referrer-Policy headers
Why it matters: Potential privacy and security leakage.
Fix: Implement appropriate referrer policy for your use case.
unsafe-external-links
What it finds: Links with target="_blank" but without rel="noopener"
Why it matters: Security vulnerability allowing opened page to control opener window.
Fix: Add rel="noopener noreferrer" to all target="_blank" links.
protocol-relative-links
What it finds: Links using // instead of https://
Why it matters: Can cause mixed content issues and security warnings.
Fix: Use absolute HTTPS URLs for all external resources.
Optimisation Opportunities (6 queries)
title-length-issues
What it finds: Title tags shorter than 30 characters or longer than 60
Why it matters: Too short titles waste opportunity; too long get truncated in search results.
Fix: Aim for 50-60 characters for optimal display in search results.
description-length-issues
What it finds: Meta descriptions shorter than 120 or longer than 160 characters
Why it matters: Poor descriptions reduce click-through rates.
Fix: Write descriptions between 150-160 characters for full display.
title-equals-h1
What it finds: Pages where title tag matches H1 exactly
Why it matters: Missed opportunity to target different keywords or angles.
Fix: Make title and H1 complementary but not identical for broader keyword coverage.
no-outbound-links
What it finds: Pages with zero external links
Why it matters: Can appear spammy or siloed; linking to quality sources builds trust.
Fix: Add relevant external links to authoritative sources where appropriate.
high-external-links
What it finds: Pages with excessive external links (20+)
Why it matters: Can appear spammy and leaks PageRank unnecessarily.
Fix: Reduce external links to most relevant and valuable resources.
missing-images
What it finds: Pages without any images
Why it matters: Images improve engagement and provide additional ranking signals.
Fix: Add relevant, optimized images with proper alt text.
Using Queries
In Claude Desktop:
List all available queries
Show me the critical queries only
Run the missing-titles query on my crawl
What does the orphan-pages query check for?In CLI:
# List all queries
seo-crawler-mcp queries
# Filter by category
seo-crawler-mcp queries --category=security
# Filter by priority
seo-crawler-mcp queries --priority=CRITICALEach query returns:
Affected URLs
Relevant context (word count, status codes, etc.)
Count of affected pages
Organized by severity
Development
# Build
npm run build
# Development mode
npm run dev
# Run tests
npm testVersion History
v2.0.1 (2026-02-02)
Fixed MemoryStorage cleanup bug (added explicit purge in finally block)
Added CLI mode for terminal-based crawling
Removed proprietary tool references from documentation
Ensures guaranteed fresh state between consecutive crawls
v2.0.0 (2026-02-01)
Added comprehensive SQL-based analysis engine
28 SEO queries covering industry-standard audit requirements
Three analysis tools: analyze_seo, query_seo_data, list_seo_queries
86% coverage of standard SEO audit requirements
v1.1.0 (2026-02-01)
Enhanced data collection with security headers
Heading structure validation (H1-H6)
Link security analysis
Response time accuracy improvements
v1.0.0 (2026-01-31)
Initial release with SQLite storage
LibreCrawl pattern implementation
Basic crawl tool (run_seo_audit)
Licence
Apache License 2.0
Copyright 2026 Richard Baxter
This product includes software developed by Apify and the Crawlee project. See NOTICE file for details.
Support
GitHub: https://github.com/houtini-ai/seo-crawler-mcp
Issues: https://github.com/houtini-ai/seo-crawler-mcp/issues
Author: Richard Baxter hello@houtini.com
Tags: seo, crawler, audit, technical-seo, mcp, crawlee, sqlite, web-scraping, site-analysis
Available Tools
4 toolsanalyze_seoB
Analyze SEO data from a completed crawl. Runs 25+ SQL queries to detect critical issues, content problems, technical SEO issues, security vulnerabilities, and optimization opportunities. Returns structured report with affected URLs and fix recommendations.
| Name | Required | Description | Default |
|---|---|---|---|
| crawlPath | Yes | Path to crawl output directory (e.g., C:/seo-audits/example.com_2026-02-01_abc123) | |
| includeCategories | No | Optional: Filter analysis by categories. Default: all categories | |
| maxExamplesPerIssue | No | Maximum example URLs to return per issue. Default: 10 | |
| format | No | Output format: "structured" (organized format, default), "summary" (text overview), "detailed" (full JSON). Default: structured |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It mentions the tool 'runs 25+ SQL queries' and returns a 'structured report', which gives some insight into computational intensity and output format. However, it lacks critical details like execution time, resource requirements, error handling, or whether it modifies data (though 'analyze' suggests read-only).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized with two sentences that efficiently convey core functionality. It's front-loaded with the main purpose, though the second sentence could be slightly more streamlined. Every phrase adds value without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 4 parameters, 100% schema coverage, and no output schema, the description provides adequate context about what the tool does and returns. However, it lacks details about the report structure, error conditions, or performance characteristics that would help an agent use it effectively, especially given the computational intensity implied by '25+ SQL queries'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, providing good documentation for all parameters. The description adds marginal value by mentioning 'critical issues, content problems, technical SEO issues, security vulnerabilities, and optimization opportunities', which loosely maps to the 'includeCategories' enum values. However, it doesn't explain parameter interactions or provide usage examples beyond what the schema offers.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with specific verbs ('analyze SEO data', 'detect issues', 'returns structured report') and resources ('completed crawl', '25+ SQL queries'). It distinguishes from siblings by focusing on post-crawl analysis rather than listing queries, querying data, or running audits.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context ('from a completed crawl') but doesn't explicitly state when to use this tool versus alternatives like 'run_seo_audit' or 'query_seo_data'. No exclusions or prerequisites are mentioned, leaving the agent to infer appropriate usage scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_seo_queriesB
List all available SEO analysis queries with descriptions, priorities, and fix recommendations. Optionally filter by category or priority level.
| Name | Required | Description | Default |
|---|---|---|---|
| category | No | Optional: Filter by category | |
| priority | No | Optional: Filter by priority level |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It describes a read-only listing operation but doesn't mention critical behaviors such as whether results are paginated, if there are rate limits, authentication requirements, or what the output format looks like. For a tool with no annotation coverage, this leaves significant gaps in understanding how it behaves.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured in two sentences: the first states the core purpose and included data, the second adds optional filtering. Every word contributes to understanding, with no redundant or vague language, making it appropriately sized and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description is incomplete. It doesn't explain the return values (e.g., format of 'descriptions, priorities, and fix recommendations'), potential limitations, or error handling. For a tool that lists data with multiple attributes, more context is needed to use it effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both parameters well-documented in the schema (including enums and descriptions). The description adds marginal value by mentioning filtering options but doesn't provide additional semantic context beyond what the schema already states. This meets the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'List all available SEO analysis queries with descriptions, priorities, and fix recommendations.' It specifies the verb ('List') and resource ('SEO analysis queries') with details about what information is included. However, it doesn't explicitly differentiate this from sibling tools like 'query_seo_data' or 'analyze_seo', which prevents a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides implied usage guidance by mentioning optional filtering by category or priority level, suggesting when to use these parameters. However, it lacks explicit guidance on when to choose this tool over alternatives like 'analyze_seo' or 'run_seo_audit', and doesn't specify prerequisites or exclusions, leaving room for ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_seo_dataA
Execute a specific SEO analysis query by name. Use list_seo_queries to see available queries. Returns detailed results with affected URLs and context.
| Name | Required | Description | Default |
|---|---|---|---|
| crawlPath | Yes | Path to crawl output directory | |
| query | Yes | Query name (e.g., "missing-titles", "duplicate-h1", "orphan-pages"). Use list_seo_queries to see all available queries. | |
| limit | No | Optional: Maximum number of results to return. Default: 100 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses that the tool 'Returns detailed results with affected URLs and context', which gives some behavioral insight about output format. However, it doesn't mention important traits like whether this is a read-only operation, potential performance/rate limits, authentication needs, or what 'execute' entails computationally. The description adds basic context but leaves significant gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is perfectly concise with three sentences that each earn their place: first states the core purpose, second provides prerequisite guidance, third describes return format. No wasted words, front-loaded with the main action. Excellent structure for quick comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 3 parameters with 100% schema coverage but no annotations and no output schema, the description provides adequate but incomplete context. It covers purpose, prerequisite, and return format at a high level, but doesn't address behavioral aspects like safety, performance, or error handling. For a query execution tool with no output schema, more detail about result structure would be helpful.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds minimal value beyond the schema: it reinforces that 'query' should be a name and mentions list_seo_queries for discovery (which is also in the schema). It doesn't provide additional semantic context about how parameters interact or usage patterns. Baseline 3 is appropriate when schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Execute a specific SEO analysis query by name' with the resource being 'SEO analysis query'. It distinguishes from siblings by mentioning list_seo_queries for discovery, but doesn't explicitly differentiate from analyze_seo or run_seo_audit. The verb 'execute' is specific, though not as precise as it could be regarding what execution entails.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool: 'Use list_seo_queries to see available queries' establishes a prerequisite. It implies usage for executing named queries rather than other SEO operations, but doesn't explicitly state when NOT to use it or name alternatives among siblings like analyze_seo or run_seo_audit, which could cause confusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_seo_auditB
Crawl a website and extract comprehensive SEO data using Crawlee HttpCrawler. Returns crawl ID and output path.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Starting URL to crawl (must include http:// or https://) | |
| maxPages | No | Maximum number of pages to crawl (1-10000). Default: 1000 | |
| depth | No | Maximum crawl depth (1-10). Default: 3 | |
| userAgent | No | User agent to identify as: "chrome" (default, Chrome browser) or "googlebot" (Googlebot crawler). Default: chrome |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the tool returns a crawl ID and output path, which adds some context, but fails to cover critical aspects such as whether this is a long-running operation, potential rate limits, authentication needs, or what 'comprehensive SEO data' entails. The use of 'Crawlee HttpCrawler' hints at technical implementation but doesn't clarify behavioral traits for the agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded and concise, consisting of two sentences that efficiently convey the core functionality and return values without unnecessary details. Every sentence earns its place by stating the action and output, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a crawling tool with 4 parameters and no output schema, the description is moderately complete. It specifies the action and return values but lacks details on output format, error handling, or operational constraints. Without annotations, it should do more to guide the agent on usage and behavior, but it meets a minimum viable level for understanding the tool's basic function.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all parameters. The description adds no additional meaning beyond the schema, such as explaining interactions between parameters or providing usage examples. The baseline score of 3 reflects adequate coverage by the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Crawl a website and extract comprehensive SEO data') and resource ('website'), specifying the tool's purpose. It distinguishes from siblings by mentioning the crawling aspect, though it doesn't explicitly contrast with tools like 'analyze_seo' or 'list_seo_queries'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'analyze_seo' or 'list_seo_queries'. The description implies usage for crawling and extracting SEO data but lacks explicit context, prerequisites, or exclusions, leaving the agent to infer based on tool names alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v2.1.3- First observed
analyze_seo - First observed
list_seo_queries - First observed
query_seo_data - First observed
run_seo_audit
TDQS
Scored across 4 tools
The tools have mostly distinct purposes, but there is some overlap between analyze_seo and query_seo_data, as both involve executing SEO analysis queries. However, analyze_seo runs a predefined set of queries, while query_seo_data allows for executing specific queries by name, which helps differentiate them. The other tools (list_seo_queries and run_seo_audit) are clearly distinct.
All tool names follow a consistent verb_noun pattern with snake_case (e.g., analyze_seo, list_seo_queries, query_seo_data, run_seo_audit). The naming is predictable and readable, with no deviations in style or convention.
With 4 tools, the count is reasonable for an SEO crawler server, covering key operations like crawling, listing queries, executing queries, and analyzing data. It might be slightly thin for a comprehensive SEO toolset, but it is well-scoped and each tool earns its place.
The toolset covers core SEO analysis workflows, including crawling, query listing, and data analysis. However, there are notable gaps, such as missing update or delete operations for managing queries or crawl data, and no tools for monitoring or reporting beyond the initial analysis. This could limit agent flexibility in handling ongoing SEO tasks.
Maintenance
Related MCP Connectors
Crawls a website for broken links, redirects and SEO issues, and lists changes since the last run
- RampifyOAuthdev.rampify
SEO MCP server: crawl your site, find AI-visibility gaps, and ship the fix from your coding agent.
Crawl a site for broken links, 404s, dead images, redirect chains and slow pages, with sources
Free technical-SEO audit MCP: crawl a site, run checks, return an LLM-ready shareable report.
Related MCP Servers
- AlicenseAqualityFmaintenanceDownloads your entire Search Console dataset into a local SQLite database, then gives your LLM a pre-built SQL query library for every standard SEO analysis type, with context available for your LLM to perform any SQL query to answer your questions and analyse for you.1214 npm16Apache 2.0
- AlicenseNot gradedqualityCmaintenanceAgent-first SEO toolkit with 24 MCP tools for keyword research, rank tracking, site audits up to 50k pages, competitor analysis, content gap detection, domain reputation, backlink intelligence, Google Search Console integration, and AI-powered strategy generation with Claude, GPT, and Ollama. SQLite-backed and bring-your-own-key.MIT
- AlicenseAqualityDmaintenanceEnables SEO auditing and site analysis by crawling websites, identifying issues, and generating reports like sitemaps and markdown exports.510 npm4MIT
- AlicenseNot gradedqualityCmaintenanceOpen-source technical SEO crawler MCP server built on LibreCrawl. Runs full audits inside Claude, Cursor, or Codex — 50+ checks (hreflang, schema.org, security headers, WAF detection on 200-OK pages), chunked-progressive engine for large sites, ephemeral by design (server forgets every audit after download).42MIT