Skip to main content
Glama
Crawlora-org

Crawlora MCP

Official

datasets_techstack_facets

Compute distribution counts for a chosen tech-stack facet, applying search-style filters, to view technology or category market share across websites.

Instructions

Facet the website tech-stack dataset. Returns distribution counts over the website tech-stack index (dataset id enum value techstack), honoring the same filters as search — the technology / category market-share view. Facet enum: technology, category, cms, ecommerce, cdn, web_server, server_language, analytics, tld, render_tier, seed_source.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
qNoSubstring match on the site domain, max 256 characters
cdnNoExact CDN / hosting filter, e.g. Cloudflare, Fastly, Vercel
cmsNoExact CMS filter, e.g. WordPress, Shopify, Webflow
notNoRepeatable exact technology name the site must NOT use (excludes)
tldNoExact top-level-domain filter, e.g. com, org, io
facetYesFacet enum: technology, category, cms, ecommerce, cdn, web_server, server_language, analytics, tld, render_tier, seed_source
any_ofNoRepeatable exact technology name; the site must use at least one (OR)
run_idNoScan run id; defaults to the latest run
categoryNoExact category filter, e.g. Ecommerce, CMS, Analytics
ecommerceNoExact e-commerce platform filter, e.g. Shopify, WooCommerce, Magento
reachableNotrue keeps only sites whose homepage was fetched
technologyNoRepeatable exact technology name the site MUST use (AND)
web_serverNoExact web-server filter, e.g. nginx, Apache, IIS
has_captchaNotrue keeps only sites with a detected CAPTCHA
render_tierNoFetch-tier filter. Enum: http, browser
seed_sourceNoSource filter for where the domain was discovered, e.g. tranco
min_tech_countNoMinimum number of detected technologies, 0 or greater
server_languageNoExact server language / framework filter, e.g. PHP, ASP.NET, Ruby on Rails
is_infrastructureNofalse (the common case) excludes backend CDN/DNS/cloud-vendor hostnames that rank highly but were never meant to serve a public homepage, keeping only real, human-navigable sites; true keeps only those backend hostnames

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed5 schema fields changedv1.17.5
    • addedInput schema / properties / any_of / items
      Added value: +{
      +  "type": "string"
      +}
    • addedInput schema / properties / facet / enum
      Added value: +[
      +  "technology",
      +  "category",
      +  "cms",
      +  "ecommerce",
      +  "cdn",
      +  "web_server",
      +  "server_language",
      +  "analytics",
      +  "tld",
      +  "render_tier",
      +  "seed_source"
      +]
    • addedInput schema / properties / not / items
      Added value: +{
      +  "type": "string"
      +}
    • addedInput schema / properties / render_tier / enum
      Added value: +[
      +  "http",
      +  "browser"
      +]
    • addedInput schema / properties / technology / items
      Added value: +{
      +  "type": "string"
      +}
  2. Changed1 schema field changedv1.16.0
    • addedInput schema / properties / is_infrastructure
      Added value: +{
      +  "description": "false (the common case) excludes backend CDN/DNS/cloud-vendor hostnames that rank highly but were never meant to serve a public homepage, keeping only real, human-navigable sites; true keeps only those backend hostnames",
      +  "type": "boolean"
      +}
  3. Addedv1.5.0

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it does disclose the core output behavior ('Returns distribution counts'). It also conveys that filtering semantics are shared with search. But it leaves gaps typical for an aggregation tool: no mention of result limits, ordering of buckets, whether counts are per-site or per-technology-occurrence, or pagination for high-cardinality facets. Adequate core disclosure without rich behavioral detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the verb and resource, and the second sentence adds the most valuable context (output + filter sharing). However, the third sentence repeats the full 11-value facet enum that already exists in the input schema, which is wasted space. Overall compact and well-ordered, with one redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 19-parameter tool with a single required param and no output schema, the description conveys the essential calling contract: pick a facet enum, optionally pass any search filter, and receive distribution counts. The strongest missing piece is the unit of counts and result shape (bucket ordering/limits), but an agent can invoke the tool correctly based on this description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with examples for every parameter, so the baseline is 3. The description adds genuine cross-tool meaning beyond the schema: 'honoring the same filters as search' tells the agent that all 18 filter parameters are semantically interchangeable with the sibling search tool, enabling reuse of query construction. The facet enum is redundantly listed, but the search-filter equivalence is real added value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action and resource — 'Facet the website tech-stack dataset' — and clarifies what the tool produces ('distribution counts over the website tech-stack index'). It differentiates from sibling tools by framing itself as 'the technology / category market-share view' and by noting it 'honor[s] the same filters as search,' which separates it from datasets_techstack_item and datasets_techstack_search. The dataset id enum value `techstack` is also named, removing ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use this tool — when a distribution / market-share breakdown over the tech-stack index is wanted — and explicitly ties its filter behavior to the search counterpart ('honoring the same filters as search'). However, it never names the alternative tool (datasets_techstack_search) or states the flip-side condition ('use search when you need individual records'), leaving the choice partly implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools