Skip to main content
Glama
Crawlora-org

Crawlora MCP

Official

web_techstack

Detect the web technologies a site uses by fetching a public URL. Returns frameworks, CMS, analytics, CDNs, and server details with confidence scores.

Instructions

Tech stack — detect what a website is built with. Fetches a public URL and fingerprints the web technologies it is built with — a BuiltWith / Wappalyzer-style detector. Returns a list of detected technologies, each with its categories, a confidence (high, medium, low), an optional version, and the evidence that matched. Covers JavaScript frameworks and libraries (React, Vue.js, Angular, Svelte, jQuery), web frameworks / static site generators (Next.js, Nuxt.js, Gatsby, Remix, SvelteKit, Astro, Hugo), CMS and website builders (WordPress, Drupal, Joomla, Ghost, Wix, Squarespace, Webflow), e-commerce (Shopify, WooCommerce, Magento, BigCommerce), analytics, ad pixels, and tag managers (Google Analytics, Google Tag Manager, Meta Pixel, LinkedIn, Bing, TikTok/Pinterest/Reddit pixels, Segment, Hotjar, Microsoft Clarity), CDNs, UI frameworks and fonts, payments (Stripe, PayPal, Klarna), live chat, marketing automation, A/B testing, consent management, CAPTCHAs (reCAPTCHA, hCaptcha, Turnstile), video, and search. It also inspects response headers (from a plain HTTP fetch) to identify the web server (nginx, Apache, IIS), the CDN / hosting provider (Cloudflare, CloudFront, Fastly, Vercel, Netlify), and the server-side language / framework (PHP, ASP.NET, Ruby on Rails, Django, Laravel, Express). Results are directional, not exhaustive. The render fetch strategy is one of browser (headless browser that executes JavaScript — the default, so client-injected scripts like analytics, tag managers and pixels are detected), auto (Chrome-impersonated HTTP, escalating to a real browser only when blocked or JS-rendered), or http (HTTP only, no JavaScript — fastest, but sees only the server HTML); defaults to browser. Only public pages are supported; respect each site's terms of use and robots directives. Also returns unmatched_evidence (when present) — third-party script/stylesheet host domains and a <meta generator> value the detector saw on the page but doesn't yet have a named signature for; useful for spotting a vendor worth requesting coverage for. is_infrastructure flags a URL whose host looks like backend CDN/DNS/cloud-vendor infrastructure rather than a real, human-navigable website. reachable is false only when the target could not be fetched at all (even after an automatic www./plain-HTTP retry) — technologies may still be partially populated from DNS-based signals alone in that case, and failure_reason explains what happened.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
requestYesTarget URL (and optional render strategy)

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changedv1.17.5
    • addedInput schema / properties / request / properties
      Added value: +{
      +  "render": {
      +    "enum": [
      +      "browser",
      +      "auto",
      +      "http"
      +    ],
      +    "type": "string"
      +  },
      +  "url": {
      +    "type": "string"
      +  }
      +}
    • addedInput schema / properties / request / required
      Added value: +[
      +  "url"
      +]
  2. Addedv1.5.0

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden, and it exceeds expectations. It discloses that the tool fetches a URL, inspects headers, and returns a list of technologies with confidence, version, and evidence. It explains the render strategies' behavior (browser executes JavaScript, http sees only server HTML), states that results are directional, defines edge cases (unmatched_evidence, is_infrastructure, reachable, failure_reason), and notes partial population from DNS signals when fetch fails. This is comprehensive and leaves no behavioral ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long (over 300 words) but well-structured: it front-loads the primary purpose, then enumerates categories, then explains render strategies, and finishes with return fields and edge cases. Every sentence adds value for a complex tool. It is slightly verbose but appropriately so given the tool's richness; a compact summary would lose critical behavioral detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and a complex return structure, the description explains all key outputs: technologies array with categories, confidence, version, evidence; unmatched_evidence; is_infrastructure; reachable; and failure_reason. It also covers the render parameter semantics, retries, and partial success behavior. Nothing an agent needs to invoke this tool correctly or interpret its results is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema only has a one-line description for the request object ('Target URL (and optional render strategy)'). The description adds substantial meaning: it explains the `render` enum in detail (browser vs auto vs http, what each does, and the default), and clarifies that only public pages are supported. This goes far beyond the schema's minimal annotation, meaningfully aiding parameter selection.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'detect what a website is built with' and 'fingerprints the web technologies it is built with — a BuiltWith / Wappalyzer-style detector.' It clearly distinguishes this tool from all siblings (which are news, product, or market-specific tools). The purpose is unambiguous and immediately understood.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There are no close sibling alternatives, so explicit when-to-use-vs-alternatives is not required. The description provides clear context: it explains the render strategies (browser, auto, http) with their trade-offs and defaults, and states limitations (only public pages, respect terms of use, results are directional). It lacks an explicit 'use this when...' statement but the purpose is self-evident and the usage guidance on parameters is thorough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools