Skip to main content
Glama
paulet4a-commits

WebDataTools Domain & website intelligence MCP server

tech_stack_detector

Detect the CMS, e-commerce platform, analytics, ad pixels, chat widgets, and payment providers a website uses, returning one row per domain.

Instructions

Tech Stack Detector finds the CMS, e-commerce platform, analytics, ad pixels, chat widget and payment provider any website runs — one row per domain. Billed to your own Apify account: ~$0.02 per result (Apify free-plan price, lower on paid plans).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
startUrlsYesWebsites — Enter the websites to fingerprint, one row is returned per website. Use bare domains or full URLs, e.g. example.com or https://example.com. Example: [{"url":"https://apify.com"}].
categoriesNoTechnology categories — Select which technology categories to detect, e.g. cms, ecommerce. Leave empty to detect everything; narrowing the list makes runs cheaper and keeps the output to the columns you care about. Options: cms = CMS; ecommerce = E-commerce platform; analytics = Analytics; advertising = Advertising & retargeting pixels; emailMarketing = E-mail & marketing automation; chat = Chat & support widgets; payments = Payment providers; frameworks = Front-end frameworks; cookieConsent = Cookie consent; cdn = CDN & hosting; server = Web server.
followSubdomainsNoFollow subdomains — Turn this on to also inspect subdomains such as shop.example.com, which often run a different platform than the main site.
maxPagesPerDomainNoMax pages per website — Enter how many pages to fetch per website, e.g. 3. The homepage exposes most of the stack; 2-4 pages also catches checkout and marketing scripts that only load on shop or contact pages.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden. It usefully discloses that runs are billed to the caller's own Apify account at roughly $0.02 per result, which implies an Apify account/token prerequisite and gives cost predictability. It stops short of explaining execution behavior (e.g. asynchronous actor runs, result delivery, rate limits) or whether the scrape is non-destructive, leaving some behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences: the capability statement is front-loaded, and the billing note follows as secondary but relevant operational context. No filler or redundant restatement of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given four simple parameters, full schema coverage and no output schema, the description covers purpose, output shape ('one row per domain') and cost. It would be fully complete if it hinted at the returned columns or the execution model, but nothing essential for calling the tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with each parameter fully documented in the schema including examples, defaults, ranges and enum meanings, so the description is not required to compensate. The description adds no parameter detail beyond the schema — it only echoes the result-row semantics. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: it fingerprints a website's CMS, e-commerce, analytics, ad pixels, chat widget and payment providers, with 'one row per domain' defining the output granularity. This is clearly distinct from all siblings (contact_extractor, dns_email_security_checker, seo_page_audit), none of which detect installed technologies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the tool's narrow purpose, and the description notes that narrowing `categories` makes runs cheaper, which is useful operational guidance. However, it never says when to prefer this over a sibling such as core_web_vitals_audit or seo_page_audit, and gives no exclusions or prerequisites beyond the billing note.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.