Skip to main content
Glama

awwwards-mcp

Free, open-source MCP server that gives AI agents design inspiration from Awwwards — the Mobbin-style visual reference loop, sourced from the web's best award-winning websites.

Your agent searches in natural language ("dark 3D portfolio sites", "soft pastel e-commerce"), sees real screenshots inline, and can pull the design DNA of any site: color palette, tech stack, design elements, award history. Free-text queries run on a porter-stemmed, prefix-matching FTS5 index with BM25 ranking — "magazines" now finds Magazine-tagged sites (68 on the live index), best matches first, where the old substring path returned zero. Multi-word queries keep AND semantics: every token must hit the same site.

Tools

Tool

What it does

search_sites

Search by color, tags, technology, award type or free-text query. Multi-word queries match against the local FTS5 index and rank BM25 (title hits lead); zero results come with loose-match and taxonomy-tag hints. Default full results include each site's screenshot; use responseMode: "compact" for concise cards and inline previews of only the first two results on the requested page.

get_site_details

Full design DNA for one site: palette, technologies, elements, awards, description.

compare_sites

Compare 2–3 sites' design DNA and jury scores as text-only JSON. Uses cached details or fetches missing detail pages.

get_index_status

Read local index count, crawl progress, last success/error, lock state and freshness without network requests.

get_site_elements

Component-level visuals for one site: each element's poster image inline (3D models, video content, mobile layouts, microcopy…) + video URLs.

list_categories

Every filter the agent can search by (200+ tags, 27 colors).

capture_live_site

Optional: fresh full-page screenshot of any live URL. Waits for load + a settle window with a bounded pre-scroll, so heavy sites work (waitStrategy: "networkidle" available). Pass viewport: "mobile" for the 390×844 iPhone-class render ("desktop" 1440×900 default). (needs playwright).

analyze_page_structure

Section band map of any page (live URL or local file:// build): tag, background, offset, height per band. Compare a reference site's structure against your build. Same heavy-site-friendly wait (waitStrategy: "networkidle" available); viewport: "mobile" analyzes the phone-class layout ("desktop" default). (needs playwright).

record_site_motion

Optional: short motion-through video of a live URL — preloader, scroll-triggered and hover/cursor animations. Returns an inline filmstrip JPEG plus the saved .webm path. viewport: "mobile" records at phone size — the filmstrip renders at the selected viewport, no pillarboxing ("desktop" default). (needs playwright + ffmpeg-static).

search_elements

Search the inspiration-elements gallery (footer, hero, pricing, 404…) by free text; ranks BM25 over title/author/category.

get_element

One element record: title, category, author, built-with stack, related elements and its media URL (image or video) pointing at awwwards' CDN.

get_motion_dna

Runtime motion fingerprint of a live URL: animation libraries, render engines, ScrollTrigger stats (trigger count, scrub ratio), tween easing/duration vocab and the scroll model. Fresh capture or cached capture with timestamp.

search_motion

Search previously captured motion-DNA scans by library, scroll model or easing vocabulary — find references by how a site moves.

new_winners

Poll today's freshly-crowned winners (SOTD / Developer Award / Honorable Mention) against a persisted baseline. First call seeds and dumps the listing; later calls report the delta. Each first-seen winner's Elements section is backfilled into the searchable element corpus, so new winners are element-searchable immediately.

watch_site

Persistent watches over studios, tags or specific sites (add/list/remove). list matches each watch against the freshest cached listing and reports per-watch NEW since last check deltas — it never fetches; new_winners/search_sites keep the pool fresh.

Data posture: element records store metadata + media URLs pointing at awwwards' own CDN — nothing is mirrored. Motion DNA records are local captures, each stamped with the time it was taken.

Related MCP server: A1 Gallery MCP Server

Setup

v1.0.0 — the first stable release. Any MCP-compatible coding agent can use awwwards-mcp — no API key, no account. Requires Node ≥ 22.13 (node -v to check). Pick your agent:

Updates: the server checks the npm registry once a day and prints an stderr notice when a newer awwwards-mcp exists (stdout stays clean for the JSON-RPC channel — your agent sees the notice as a log line). Set AWWWARDS_AUTO_UPDATE=1 in the server's env to opt into background self-update; restart your agent afterwards to load it. Nothing is fetched more than once a day and serving never waits on the check.

Claude Code

claude mcp add awwwards -- npx -y awwwards-mcp

Codex CLI (ChatGPT desktop app and the IDE extension share this config)

codex mcp add awwwards -- npx -y awwwards-mcp

or in ~/.codex/config.toml (project-scoped: .codex/config.toml):

[mcp_servers.awwwards]
command = "npx"
args = ["-y", "awwwards-mcp"]

OpenCode (opencode.json — note the command is an array)

{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "awwwards": {
      "type": "local",
      "command": ["npx", "-y", "awwwards-mcp"]
    }
  }
}

ZCode (~/.zcode/cli/config.json — note servers nest under "mcp": { "servers": ... })

{
  "mcp": {
    "servers": {
      "awwwards": { "command": "npx", "args": ["-y", "awwwards-mcp"], "env": {} }
    }
  }
}

Claude Desktop / Cursor / Windsurf / Gemini CLI / Cline / Continue — anything reading the common mcpServers JSON shape (e.g. ~/.claude/claude_desktop_config.json or ~/.gemini/settings.json):

{
  "mcpServers": {
    "awwwards": { "command": "npx", "args": ["-y", "awwwards-mcp"] }
  }
}

Anything else — awwwards-mcp is a plain stdio MCP server: point your client at npx -y awwwards-mcp and it works. To pin a version, use npx -y awwwards-mcp@1.0.0.

pi coding agent has no built-in MCP by design — it uses skills and extensions instead. Two options:

  1. Install the awwwards-inspiration skill (below). pi reads skills from ~/.pi/agent/skills/ or ~/.agents/skills/ (the latter is shared across agents following the Agent Skills standard). The skill teaches the workflow; for it to reach the live data, add an MCP-supporting pi extension, or run the queries in another agent and paste results.

  2. Skip MCP entirely: ask pi to build you a small CLI wrapper around awwwards.com, or use a shared skills directory (~/.agents/skills/) so the same skill file serves pi and every other agent.

Optional full-page captures (needed by capture_live_site, analyze_page_structure, record_site_motion):

npm install -g playwright && npx playwright install chromium

record_site_motion additionally uses ffmpeg; it resolves the ffmpeg-static package automatically if present.

Skills

This package ships four agent skills. Any agent that follows the Agent Skills standard can load them; copy them into your agent's skills directory:

npm install awwwards-mcp
mkdir -p ~/.agents/skills && cp -r node_modules/awwwards-mcp/skills/awwwards-inspiration node_modules/awwwards-mcp/skills/awwwards-setup node_modules/awwwards-mcp/skills/awwwards-doctor node_modules/awwwards-mcp/skills/awwwards-motion-study ~/.agents/skills/

Skill

What it teaches

awwwards-setup

First-time onboarding: asks the user's preferences (result density, viewport, captures, local index, winner watches), persists them to ~/.awwwards-mcp/preferences.json, and runs any one-time installs they opt into.

awwwards-inspiration

The inspiration loop: search, judge from screenshots, pull design DNA, state a design direction, capture/motion-first builds; staying current with new_winners and watch_site.

awwwards-motion-study

The full video chain: what to record from a live site (and what to skip), frame-by-frame review (video input or tile-per-element), the motion inventory, and build verification by re-recording.

awwwards-doctor

Repair: run npm run doctor, apply its fixes, re-anchor parsers after real awwwards.com drift, recover the in-flight task that surfaced the failure.

Agent

Skills directory

Claude Code

~/.claude/skills/

pi

~/.pi/agent/skills/ (also reads ~/.agents/skills/)

ZCode

~/.zcode/skills/

Agent Skills-standard agents

~/.agents/skills/

Windows: run this from Git Bash, or copy node_modules\awwwards-mcp\skills\awwwards-setup manually.

search_sites works out of the box, but its depth is limited by polite live scraping (~31 sites per filter page). Build a local index once and searches draw from thousands of award-winning sites instantly:

npx -y -p awwwards-mcp awwwards-index      # once published
# or, from a local checkout of this repo:
npm run index
  • Crawls all ~200 tag pages at 1 request/second (~4 minutes) into the local SQLite cache at ~/.awwwards-mcp/.

  • Resumable: interrupt it and re-run — completed pages are skipped.

  • The MCP server re-indexes automatically in the background whenever the index is older than 7 days (never blocking your session). Completed crawl checkpoints are cleared so each refresh actually revisits the tag pages. Run get_index_status to inspect progress or the last crawl error.

Elements index (optional)

The search_elements tool auto-indexes the first gallery page (~48 items) on first use. To build a full corpus (~1,500+ items, ~15 pages at 48/page):

npm run index -- --elements all     # follow pagination until exhausted
npm run index -- --elements 10      # first 10 listing pages

Any crawl beyond the first page also runs the taxonomy pass: every facet page (/elements/footer/, /elements/cta/, … 46 categories) is fetched once and each element under it is stamped with that category, so search_elements' category filter works across the corpus. Element pages carry no breadcrumb, so this listing-side pass is the only category source; elements seen on no facet page stay unsorted.

Element rows are searched by title, author, category and slug tokens (slug is an FTS5-indexed column; a cache opened from an older schema version rebuilds its search index automatically on first open). Each element also carries the slug of the award-winning site it came from (siteSlug/siteUrl in search_elements/get_element results). Two record sources share this corpus: gallery records (indexed from the public elements listing) and source:"site" records backfilled from each new SOTD winner's own Elements section by new_winners — their slugs are namespaced site-<siteslug>-<title> so the two never collide. The elements index has the same 7-day freshness gate as the sites index — a re-run inside the window skips itself.

Site details (palettes, tech stacks) are still fetched on demand and cached for 7 days. Awwwards page and CDN requests have a 10-second deadline per attempt, including response-body reading; transient page failures are retried once, while blocks (403/429) and CDN failures are not retried.

How it works

  • Live, polite scraping of awwwards.com public pages (max 1 request/second, robots.txt-compliant paths only, cached 7 days in SQLite at ~/.awwwards-mcp/).

  • Screenshots are served from Awwwards' own CDN (880×660), cached on disk.

  • No API key, no account, no cost.

Ethics & terms

This tool fetches publicly available pages for personal design-inspiration use, at human-ish request rates, honoring robots.txt. Awwwards' screenshots and content remain the property of Awwwards and the credited creators — don't bulk-scrape, redistribute, or republish them. If you use this commercially, review awwwards.com's terms yourself.

Built with awwwards-mcp: four real sites

Four complete sites were built through the full inspiration loop this MCP enables, using nothing but the server's tools plus the shipped awwwards-inspiration skill. Each one exercised a different corner of the loop — and every correction the loop caught on the way became doctrine in the skill.

Built in one shot, by a model that can't watch video. All three sites were built in a single prompt run on GLM 5.3-flash — which does not support video input. The loop's motion study worked entirely from frame-tiled filmstrips (ffmpeg, 1–2 fps per element) instead of watching the recordings. With a video-native model, those same get_site_elements videos and record_site_motion .webm files could be watched directly — timing, easing and overlap read at full fidelity — and the motion-true results would be better still. The skill's frame-tile doctrine is what closes that gap today.

Watch the whole loop run (1:50):

Screen recording of the agent running the awwwards-inspiration loop end to end with the awwwards MCP tools — searching SOTD references with inline screenshots, pulling design DNA, frame-studying element videos, building, and verifying with band maps + motion recording. If your client doesn't render the player, watch the file directly.

1. Ridge (source) — a Swiss-minimal single-page showcase for a fictional engineering-talent studio, direction Aspen Search (SOTD + Developer Award, jury 7.48): monochrome #FAFAF8/#1A1A1A + mint, giant grotesque section markers, halftone grain, asymmetric panel grid, dark discipline panels in an interior horizontal pin passage, count-up stats, client rows, theme toggle, cursor-follower. Built with the v1.4.0 toolkit: FTS5-ranked direction search, both-viewports reference captures and QA (desktop 7,849px + mobile 390×844), overflow audit (0px both), pin-center shots, film verification — and the skill-memory flywheel recorded the findings. QA evidence: showcase/ridge/_qa/.

The grain-panel hover, studied from aspensearch.com's recording and rebuilt as a canvas dither-dissolve — dots flip to mint around the mouse, the trail elongates, the boundary dissolves:

Ridge grain-panel hover — canvas dither-dissolve following the mouse

Panel grid (desktop)

Horizontal discipline passage

Mobile 390×844

Ridge desktop — Swiss panel grid with mint and grain

Ridge disciplines — pinned horizontal passage mid-slide

Ridge mobile — stacked grid, zero overflow

2. Fallow Press (source) — a flat-2D editorial journal, direction Emergence Magazine (SOTD): pink #FF9398 on cream and black, torn-paper masthead (pure CSS clip-path, zero WebGL), giant grotesque display over grayscale photography, serif-italic brand, three pages with separate horizontal projects/about pages (GSAP ScrollTrigger pin + containerAnimation).

Torn-paper masthead (home)

Horizontal gallery (Fields)

Horizontal chapters (Practices)

Fallow Press home — torn-paper masthead over grayscale photography

Fallow Press Fields — pinned horizontal gallery panel

Fallow Press Practices — pink quote chapter

The loop as it ran:

  1. search_sites (magazine filters) → shortlist judged from inline screenshots → get_site_details on Emergence Magazine.

  2. Capture before building: capture_live_site + record_site_motion on the live site first; full-page PNG and motion .webm kept in fallow-press/ref-motion/ as the evidence trail.

  3. Build, then verify: full-page capture plus panel-center pin shots of both horizontal pages (13 stops each, in fallow-press/_qa/ — capture-qa.mjs is reusable).

  4. The pin shots caught a real bug: horizontal-panel entrances used toggleActions: "play none none reverse", and 100vw panels hide content at midpoints on the way back — copy disappeared mid-view. Fix (one-shot play entrances) is now doctrine: full-viewport panels get one-shot entrances; QA pin shots land at panel centers, not uniform fractions, or you photograph empty transition zones.

3. Cerebrium recreation (C:/Users/Afjal/cerebrium-recreation/) — a fidelity-first recreation of cerebrium.ai, pixel-checked against the live reference: full-page captures of both sides, analyze_page_structure band compare, and SVG icon/legend fixes until the build matched the reference to within 1px of total page height (10,871px vs 10,870px). This is the structure-before-pixels doctrine at its strictest — band maps compared, never just totals.

Cerebrium recreation — full-page build capture

4. The Meridian (C:/Users/Afjal/editorial-site/) — an editorial journal built from ORDR/Hearst references: the first build to run the whole loop end-to-end. analyze_page_structure caught a masthead band bug by comparing the build's band map against the reference's; the reference captures, motion film, and the reusable pre-scroll capture script live in editorial-site/_qa/.

The Meridian editorial journal — full-page build capture

The Meridian — motion filmstrip from record_site_motion

Skills used to build these

Skill

Role in the builds

awwwards-inspiration

The 8-step loop itself (ships with this package): search → judge from screenshots → design DNA → capture/motion study → state direction → build → band-map verify.

gsap-scrolltrigger

The horizontal pin + containerAnimation pattern (ease "none", one-shot entrances) driving both Fallow Press horizontal pages.

gsap-core / gsap-timeline

Tween composition and sequenced hero entrances (torn-paper drop, panel copy rises).

frontend-design

Typography, palette and layout judgment applied when translating reference DNA into original pages.

lenis (library, via skill guidance)

smooth scrolling synced to ScrollTrigger on the Fallow Press home page.

tailwindcss / plain CSS

All builds are plain hand-rolled CSS — flat 2D, no frameworks needed.

The skills self-improve: every loop pass records what verification caught (scripts/skill-memory.mjs record), and a deterministic distiller folds rules seen 2+ times into your installed skill copy — while the shipped copies only change via human PR. A techniques registry (skills/_memory/techniques.json) catalogs researched how-tos per domain (video understanding, motion detection, UI structure, micro-interactions, images).

Reduced-motion, JS-less visits, and capture tools all get graceful fallbacks (vertical stacks; progressive-enhancement reveals).

What the verification loop caught — proof the structure-before-pixels doctrine is load-bearing:

  • Element posters lie: the first showcase build was designed from poster frames alone and rendered a spinning 3D ring as floating static cards. Downloading the element videos (get_site_elements) and frame-tiling them revealed the motion truth — now the skill mandates studying motion before animating.

  • Full-page captures of reveal-on-scroll builds showed blank sections: .reveal animation state vs capture's no-scroll reality. Builds ship content-visible-without-JS progressive enhancement.

  • Horizontal-panel copy vanished mid-view on the Fallow Press pages: toggleActions reverse reverts entrances while a 100vw panel is still holding the viewport (see above).

  • Band-map compare kept the references' rhythm instead of drifting on section heights (Cerebrium, The Meridian).

Prompt counts: 3 for the original showcase build (the build ask, the motion correction that exposed the poster-lie, the structure pass) and 1 for Fallow Press ("create a new website using our MCP and skills… no 3D websites") — its two follow-ups were caught by the QA loop, not by the user. Each correction became doctrine in the shipped awwwards-inspiration skill: frame-study element videos before animating; judge page architecture from the studied passages; tile per element, not one giant filmstrip; capture live sites and animation before building.

Can awwwards-mcp crawl the sitemap? (robots.txt notes)

The awwwards.com robots.txt advertises Sitemap: https://www.awwwards.com/sitemap.xml and — verified live 2026-09-18 — that sitemap URL returns a soft-404 HTML page (as do common child names like /sitemap-websites.xml). So sitemap discovery isn't currently a path to more data; the polite crawl surface is exactly what the indexer uses:

  • Allowed and used: /websites/, /websites/<filter>/, /sites/<slug> (one filter per URL; deep pagination stays un-crawled).

  • Disallowed and never fetched: /tag/, /search-websites, /websites/? (query-string pagination), /elements/*, /vote/, favourites/likes/follows, and the rest of the 33 rules.

  • Our client (src/awwwards.ts buildFilterUrl) constructs only /websites/… paths at 1 request/second — the loop stays inside the published rules by construction, not by convention.

Contributing

PRs welcome! The project especially needs parser-drift fixes — when live awwwards.com markup changes, a fresh HTML snapshot attached to an issue often becomes the new test fixture and the fastest merged PR. See CONTRIBUTING.md for the full guide:

Parser-drift is monitored automatically. A probe script (scripts/parser-drift-probe.mjs, npm run drift) checks every markup anchor the parsers depend on — the split/indexOf/regex literals in src/parsers.ts — against the live listing and detail pages (2 fetches, 1 request/second, same politeness as the client). A daily GitHub Action (.github/workflows/parser-drift.yml) runs it and, on drift, opens/updates a single tracking issue with the exact anchors that changed (and auto-closes it when a later run is green). To run it yourself: npm run drift (live, exit code 0/1/2) or npm run drift -- --fixture (offline, checks the committed fixtures still feed every anchor). Raw HTML is never diffed or stored — anchors only fire when the parsers actually break, so there are no false alarms from cosmetic tweaks.

  • Development setup & project layout (offline fixture-tested, no network in tests)

  • How to create a PR: fork → fix//feat//docs/ branch → typecheck + tests → PR template

  • The politeness constraints new code must keep (1 req/s, robots.txt paths, light runtime deps)

Bugs and feature ideas start as issues with templates. Security problems go privately — see SECURITY.md. By participating you agree to the Code of Conduct.

Development

npm install
npm test        # offline unit tests against committed HTML fixtures
npm run smoke   # manual live smoke test against awwwards.com
npm run build   # compile to dist/

MIT — see LICENSE.

Available Tools

15 tools
analyze_page_structureA
Read-onlyIdempotent

Extract a page's section band map (tag, label, background color, offset, height per band) via a headless browser. Works on live URLs and file:// paths — use it to compare a reference site's structure against your local build.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesAbsolute URL (https:// or file://) of the page to analyze
maxBandsNoCap on returned bands
viewportNodesktop = 1440x900 (default); mobile = 390x844 iPhone-class with deviceScaleFactor 3, isMobile + hasTouchdesktop
waitStrategyNo'load' + settle works on heavy sites; 'networkidle' waits for total quietload

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, open-world). The description adds non-obvious behavioral context: it runs via a headless browser and accepts local file:// paths, which tells the agent this fetches and renders remote/local content. It stops short of noting timing/cost or pagination behavior, but that gap is minor given the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the action and output shape, followed by scope and intended use. Every clause earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully compensates by enumerating the per-band fields returned. Combined with the schema's full parameter coverage, an agent has enough to call it correctly, though return-format details (ordering, error shape) are not spelled out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents url, maxBands, viewport, and waitStrategy in detail. The description's field list describes the OUTPUT shape rather than adding meaning to the parameters, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (extract) and resource (a page's section band map) and enumerates the returned fields (tag, label, background color, offset, height per band). This clearly distinguishes it from siblings like capture_live_site, get_site_elements, and compare_sites, which operate on different data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names a concrete use case — comparing a reference site's structure against a local build — and clarifies it accepts both live URLs and file:// paths. However, it does not explicitly state when NOT to use it or name the sibling alternative for the comparison workflow (e.g. compare_sites).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

capture_live_siteA
Read-onlyIdempotent

Take a fresh full-page screenshot of a live website URL using a headless browser. Requires the optional playwright dependency.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesAbsolute URL of the site to capture
viewportNodesktop = 1440x900 (default); mobile = 390x844 iPhone-class with deviceScaleFactor 3, isMobile + hasTouchdesktop
waitStrategyNo'load' + settle works on heavy sites; 'networkidle' waits for total quietload

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description's only added behavioral fact is the headless-browser mechanism and the playwright dependency; it says nothing about timeouts, rate limits, or failure behavior on unreachable URLs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the action, zero filler. The dependency caveat is the only secondary detail and it is stated tersely.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does not indicate what a successful capture returns (image bytes, file path, URL). For a network-facing capture tool that omission is a real gap, though annotations and the rich schema cover the input side adequately.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so url, viewport and waitStrategy are fully documented in the schema, including enum meanings and defaults. The description adds nothing beyond that, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Take a fresh full-page screenshot of a live website URL') plus the mechanism ('headless browser'). It is clearly distinguishable from siblings like analyze_page_structure or record_site_motion, though it never names a sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied by the name and the verb. The one piece of guidance given ('Requires the optional playwright dependency') is a prerequisite rather than a when-to-use rule, and no alternative tool is mentioned for overlapping tasks like structure analysis.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compare_sitesA
Read-onlyIdempotent

Compare the design DNA of 2–3 Awwwards sites: titles, live URLs, palettes, technologies, elements, awards and jury scores. Text only; missing details are fetched and cached.

ParametersJSON Schema
NameRequiredDescriptionDefault
slugsYesTwo or three distinct site slugs from search_sites

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/openWorld/non-destructive, so the safety profile is covered. The description adds genuinely new context: output is text only (no imagery), and missing details are fetched and cached — implying network calls and possible latency. It stops short of stating caching lifetime or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the comparison scope, then a second sentence covering output mode and data-fetch behavior. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so enumerating the returned facets in the description is exactly what's needed; combined with the text-only and fetch/cache notes, an agent has enough to invoke and interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single param (slugs) is fully documented in the schema, including the 2–3 distinct-slug constraint that mirrors min/maxItems. The description restates the count but adds no format guidance beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Compare) and resource (design DNA of Awwwards sites) with explicit scope (2–3 sites) and enumerates the compared facets (titles, URLs, palettes, technologies, elements, awards, jury scores). This clearly distinguishes it from the single-site sibling get_site_details.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 2–3 site scope implies this is the multi-site comparison tool versus get_site_details for one site, but it never states when to prefer it or what to do for a single site. Usage is inferable rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_elementA
Read-onlyIdempotent

Get one Awwwards element in full: title, author, built-with stack, media URLs, related elements and the parent award-winning project. Use search_elements first for ids.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesElement slug from search_elements

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, non-destructive, and closed-world, so the safety profile is covered. The description adds real value by disclosing the shape of the returned content (related elements, parent project), which is otherwise unknown since there is no output schema. Error behavior and any lookup limits are not mentioned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler; the payload enumeration comes first and the prerequisite is placed last as an actionable instruction. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-param getter with no output schema, the description supplies the return-field inventory and the id-acquisition path, which is most of what an agent needs. It omits failure behavior for unknown slugs, a minor gap given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter and schema description coverage is 100% ('Element slug from search_elements'), so the schema already carries the semantics. The description reinforces the id's origin but adds no format or syntax detail beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (get) plus resource (one Awwwards element) with an explicit enumeration of the returned payload: title, author, built-with stack, media URLs, related elements, and parent project. It separates itself from the sibling search_elements by being the single-item 'full detail' endpoint.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent to search_elements first to obtain ids, which is the correct prerequisite for this tool. It stops short of stating when NOT to use it or what happens with an invalid/unknown slug, so it is clear context rather than full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_index_statusA
Read-onlyIdempotent

Read offline index status: cached site count, progress, last successful finish and error, lock state and freshness. No live requests.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive and closed-world, so the safety profile is covered. The description adds the meaningful operational detail that this is an offline cache read with no live requests, plus a lock state indicator — useful context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that lists the return contents. Every clause earns its place; no padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the return-value burden — and it does, enumerating cached site count, progress, last successful finish, error, lock state and freshness. Nothing an agent needs to call this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4; there is nothing to mis-document and the description correctly implies a parameterless read.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read offline index status') and enumerates exactly what the status comprises: cached site count, progress, last finish/error, lock state, freshness. No sibling tool overlaps with index status, so the agent can route here unambiguously.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'No live requests' signals this reads cached/offline state rather than triggering network activity, which implicitly contrasts with the capture_live_site / watch_site siblings, but the description never says when to prefer this tool over others or any prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_motion_dnaA
Read-onlyIdempotent

Get the Motion DNA of a website: animation stack, ScrollTrigger/pin/scrub counts, easing vocabulary and duration distribution. Serves the corpus first; live-captures unseen URLs (headless browser required).

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesAbsolute site URL
recaptureNoForce a fresh live capture even if a recent record exists

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint. The description adds genuinely useful operational context beyond them: it serves a cached corpus first and only then performs a live headless-browser capture, implying network access and variable latency for unseen URLs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler; the output inventory comes first and the corpus/live-capture behavior follows. Every clause conveys usable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but the description enumerates the returned fields, which compensates. Combined with annotations covering safety and idempotency and the stated headless-browser fallback, an agent has enough to call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (url and recapture are both documented in the schema), so the schema carries parameter meaning. The description does not add syntax or format detail for either parameter, nor does it cross-reference the recapture flag against the corpus-first behavior it describes. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get the Motion DNA of a website') and enumerates the returned content (animation stack, ScrollTrigger/pin/scrub counts, easing vocabulary, duration distribution). This clearly separates it from search_motion or record_site_motion, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: 'Serves the corpus first; live-captures unseen URLs' tells the agent when a live capture happens, but there is no explicit when-to-use guidance versus siblings like search_motion or capture_live_site, and no stated exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_site_detailsA
Read-onlyIdempotent

Get the design DNA of one Awwwards site: color palette, technologies, design elements, awards, description and inline screenshot.

ParametersJSON Schema
NameRequiredDescriptionDefault
slugYesSite slug from search_sites, e.g. 'l-i-s-a'

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, open-world behavior, so the safety profile is covered. The description adds value beyond that by disclosing the actual returned content (palette, tech stack, elements, awards, inline screenshot), which matters since there is no output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence with the verb and resource front-loaded and zero filler; every clause (palette, technologies, elements, awards, description, screenshot) earns its place by telling the agent what comes back.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read tool with full annotation coverage, the definition is nearly complete: it states the payload and the schema supplies slug format and origin. The one gap is the absence of any usage/routing guidance relative to siblings such as search_sites.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the slug property already carries a pattern constraint, a concrete example ('l-i-s-a') and its origin tool. The description adds nothing further about the parameter, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (Get) plus resource (one Awwwards site) and an explicit enumeration of the payload — color palette, technologies, design elements, awards, description, screenshot. That enumeration usefully separates it from siblings like get_site_elements or get_motion_dna, though no sibling is named outright.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use, no prerequisites, and no alternatives. The only workflow cue ('slug from search_sites') lives in the input schema, not in the description, so an agent gets no routing guidance from the text itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_site_elementsA
Read-onlyIdempotent

Get the design-element highlights of one Awwwards site: component-level visuals (3D models, video content, mobile layouts, microcopy) with poster images inline and video URLs.

ParametersJSON Schema
NameRequiredDescriptionDefault
slugYesSite slug from search_sites, e.g. 'l-i-s-a'

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, open-world), so the bar is lower. The description adds genuine context the annotations cannot: posters are returned inline while video is delivered as URLs, which shapes how the agent handles the payload. It omits pagination or result-limit behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the leading clause carries the action and the trailing parenthetical enumerates content types. Slightly dense with the nested list, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully characterizes the return payload (inline posters, video URLs). It stops short of covering volume, ordering, or empty-result behavior for a site with no elements, which is the main remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and even supplies a format hint ('l-i-s-a'), so the parameter is fully documented structurally. The description adds no syntax or sourcing detail beyond implying a single-site identifier, making the baseline 3 correct.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get the design-element highlights of one Awwwards site') and enumerates the component types returned (3D models, video content, mobile layouts, microcopy). The single-site scope implicitly separates it from search_elements/get_element, though no sibling is named outright.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: 'of one Awwwards site' signals per-site retrieval, but there is no explicit when-to-use, when-not, or pointer to alternatives such as search_elements for cross-site queries. An agent must infer the routing itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_categoriesA
Read-onlyIdempotent

List the filter taxonomy available on Awwwards: color hexes and tag/technology slugs usable with search_sites.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive). The description adds value beyond them by disclosing what the return content actually is (color hexes, tag/technology slugs) and how it feeds into search_sites, which is the real behavioral insight for a zero-arg lookup tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler. The resource is named first, the contents second, and the downstream use last—no sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters, no output schema, and annotations covering safety, the description is nearly self-sufficient for such a simple tool. It could optionally note the output shape (e.g., grouped lists) but nothing essential to correct invocation is absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes no parameters, so there is nothing for the description to disambiguate; the baseline for a zero-param tool applies. No parameter meaning is missing.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (filter taxonomy on Awwwards), then enumerates the concrete contents (color hexes and tag/technology slugs). An agent can immediately distinguish this discovery tool from sibling search tools like search_sites.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase "usable with search_sites" makes the purpose context explicit: this is the tool to call first to learn valid filter values before searching. It does not spell out a when-not condition, but the intended workflow is clear from the linkage to a named sibling.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

new_winnersA
Read-only

Poll today's new Awwwards winners as a delta against the previous poll: new-entrant cards since the last call (first call seeds the baseline and reports no delta). Poll daily for a winners feed; pairs with watch_site for structured monitoring.

ParametersJSON Schema
NameRequiredDescriptionDefault
awardNoWhich winners feed to pollsotd

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, destructiveHint=false, and the key idempotentHint=false. The description earns credit by explaining that non-idempotence: the first call seeds the baseline and returns no delta, and later calls return only new entrants. That stateful, call-order-dependent behavior is exactly what an agent needs and is not derivable from the annotations alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences, front-loaded with the core action and delta semantics before the polling cadence and sibling pairing. Every sentence carries information, though the two polling-related clauses overlap slightly in intent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description must convey the return shape, and it does so adequately by defining the result as new-entrant cards since the last call, including the baseline-seeding edge case. It omits error/rate-limit behavior, which is a minor gap for a polling tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single 'award' enum is fully documented in the schema, so the baseline is 3. The description never mentions the award parameter or how it affects the delta, adding nothing beyond the structured schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Poll'), a precise resource ('today's new Awwwards winners'), and the exact mechanism ('as a delta against the previous poll'). It is immediately distinguishable from the read-oriented siblings like search_sites or get_site_details, and it explicitly names its closest relative, watch_site.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear usage context ('Poll daily for a winners feed') and routes the agent to the sibling tool for a related mode ('pairs with watch_site for structured monitoring'). It stops short of stating an explicit when-not condition or full alternatives, so it is strong but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

record_site_motionA
Read-onlyIdempotent

Record a short motion-through video of a live website — preloader, scroll-triggered and hover/cursor animations — and return an inline filmstrip JPEG plus the .webm path. Requires the optional playwright and ffmpeg-static dependencies.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesAbsolute URL of the site to record
framesNoFilmstrip tile count (default 16 → a 4x4 grid)
viewportNodesktop = 1440x900 (default); mobile = 390x844 iPhone-class with deviceScaleFactor 3, isMobile + hasTouchdesktop
waitStrategyNo'load' + settle works on heavy sites; 'networkidle' waits for total quietload

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare this a safe, idempotent, open-world read. The description adds genuinely useful behavioral context beyond them: what the capture covers, that it returns an inline filmstrip JPEG plus a .webm path, and that it requires the optional playwright and ffmpeg-static dependencies. It does not state duration limits, timeouts, or failure behavior for heavy sites, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences: the first front-loads purpose and scope, the second covers output and dependency requirements. No filler, and the most important information leads.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly carries the return format (filmstrip JPEG plus .webm path) and runtime prerequisites. With 100% schema coverage the parameters are handled elsewhere. The only real gap is sibling/usage discrimination, which keeps it from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already explains url, frames, viewport, and waitStrategy fully. The description adds no parameter-level detail (e.g., how long a 'short' recording is or how frame count affects runtime), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource — recording a short motion-through video of a live website — and enumerates the exact animation types captured (preloader, scroll-triggered, hover/cursor). It is unambiguous on its own, but it never names or contrasts with close siblings like capture_live_site or watch_site, so differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, no exclusions, and no routing to alternatives such as capture_live_site (static capture) or watch_site (monitoring). An agent must guess which recording/capture sibling fits a given request.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_elementsA
Read-onlyIdempotent

Search Awwwards' curated element gallery — award-grade UI components (footers, menus, loaders, transitions, micro-interactions…). Usage: pass query text and at most one of category/stack; every result links back to its source element page and parent project.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryNoFree text: what the component is or does
stackNoBuilt-with tokens, e.g. ['gsap', 'webgl']
categoryNoCategory id from the gallery taxonomy, e.g. 'footer', 'menu', 'cta', 'loading', 'mouse_interaction', '404_page'

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, open-world, non-destructive behavior, so the safety profile is covered. The description adds two things annotations cannot: the mutual-exclusivity rule for category/stack and the fact that every result links back to its source element page and parent project, which tells the agent what it gets back.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences: the first names the resource, the second front-loads the usage rule. No filler, no restatement of the name or title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description responsibly notes the shape of results (links to source element page and parent project). Combined with the input constraint, an agent has enough to call correctly; only the limit parameter's tuning guidance is left to the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, so the schema already documents query, stack, and category. The description goes beyond it by stating that at most one of category/stack may be supplied, a constraint not encoded in the schema, which materially affects how the agent builds the call.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (search) and resource (Awwwards' curated element gallery) and clarifies the resource with concrete examples (footers, menus, loaders, transitions). An agent can see this is the component-level search, distinct from search_sites and search_motion, though it never explicitly names those siblings to sharpen the contrast.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Usage:' clause gives a real input constraint (query text plus at most one of category/stack), which is actionable. However, it says nothing about when to prefer this tool over get_site_elements, get_element, or search_sites, so the when-to-use decision is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_motionA
Read-onlyIdempotent

Search captured Motion DNA records: find sites that pin sections, scrub scroll-driven animation, or run a given animation library. Corpus only — live-capture happens via get_motion_dna.

ParametersJSON Schema
NameRequiredDescriptionDefault
libNoLibrary token, e.g. 'gsap', 'lenis', 'webgl'
limitNo
hasPinsNoOnly sites that pin sections
scrubOnlyNoOnly sites with substantial scroll-scrubbed animation

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive and closed-world, covering the safety profile. The description's 'Corpus only' line reinforces the closed-corpus behavior and adds the live-capture contrast, which is useful but largely consistent with openWorldHint=false. Nothing is said about result shape, ranking, or how limit interacts with the corpus.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the verb and resource, then a single routing clause. No filler and nothing repeated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description should convey what comes back; 'find sites that...' implies site-level results, which is adequate for a four-param, all-optional search tool. Missing details are minor — result ordering and pagination behavior relative to limit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75% (lib, hasPins, scrubOnly documented), and the description largely restates those same semantics rather than adding format, token, or matching detail. The undocumented 'limit' is self-explanatory via min/max. Baseline 3 fits when the schema already carries the parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Search) and resource (captured Motion DNA records) and then concretizes the searchable properties: pinned sections, scroll-scrubbed animation, animation libraries. It also contrasts itself with the sibling get_motion_dna, so an agent can route without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly scopes usage to the existing corpus and names the alternative for the live path ('live-capture happens via get_motion_dna'), which is a real when-to-use/when-not split. It does not, however, distinguish itself from search_sites or search_elements, so the routing guidance is clear but incomplete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_sitesB
Read-only

Search award-winning websites on Awwwards. Returns site cards with inline screenshots, live URLs, awards and tags.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNo
tagsNoTag slugs, e.g. ['3d', 'portfolio']
awardNo
colorNoDominant color hex, e.g. '#404040'
countNo
queryNoFree text matched against site titles and tags
sortByNoSort results: by Awwwards jury score (details previously fetched) or newest firstnewest
technologyNoTechnology slug, e.g. 'webgl', 'gsap', 'astro'
responseModeNofull (default) returns all inline previews; compact returns concise cards and at most two page previewsfull

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare the read-only, non-destructive safety profile, so the bar is lower. The description adds useful behavioral value by describing the return shape (site cards with inline screenshots, live URLs, awards, tags), but says nothing about pagination, result limits, or the effect of responseMode on payload size.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler, with the primary action stated first and the return contents immediately after. Nothing in it can be cut without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, describing the card contents is necessary and present, which helps. However, a 9-parameter search tool with no usage scoping and undocumented pagination/count behavior is only minimally covered.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Across 9 parameters the description adds no filtering semantics at all: it only mentions tags and awards as things returned in cards, not as query filters. With 67% schema coverage, page, award, and count have no documented meaning in either place, so the description fails to compensate for the gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource ('Search award-winning websites on Awwwards') with the return payload named, so the agent can tell it is a query tool rather than a detail/fetch tool. It does not explicitly contrast itself with siblings like search_elements, new_winners, or get_site_details.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no named alternatives. The agent must infer from the name that this is the entry point before drilling into get_site_details or compare_sites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

watch_siteA

Longitudinal watchlist over Awwwards listings: track a studio (e.g. 'ToyFight'), a tag/technology slug (e.g. 'webgl'), or a site URL-slug for changes and new entries. add/list/remove; the delta is reported by watch_site list after new_winners or search_sites refreshes the listing.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNostudio = creator name from listing cards; tag = tag/technology slug; url = a /sites/<slug> site-slug
noteNoWhy you're watching (free text)
awardNoRestrict a watch to one award feed
actionYes
patternNoWhat to watch (name/slug per kind)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false and openWorldHint=true, so the safety profile is covered. The description adds non-obvious behavioral context beyond them: the mutation actions (add/list/remove) and the dependency that deltas only surface after new_winners or search_sites refreshes the underlying listing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly packed sentences front-load the resource and the three target kinds before the action list and refresh dependency. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and five parameters, the description adequately covers the targets, the three actions, and how deltas are produced. It leaves some ambiguity about exactly what 'list' returns and how the award filter interacts, though the schema covers those fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 80%, so the schema already documents kind, note, award and pattern. The description reinforces the kind semantics with concrete examples ('ToyFight', 'webgl', URL-slug), adding modest value, but introduces no syntax or format detail the schema lacks. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — a longitudinal watchlist over Awwwards listings — and enumerates the three watchable targets (studio, tag, url) with concrete examples. It is clearly distinguishable from sibling tools like search_sites or new_winners without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear workflow context: watches are managed via add/list/remove, and 'the delta is reported by watch_site list after new_winners or search_sites refreshes the listing,' explicitly naming the sibling calls that feed it. It stops short of stating when NOT to use the tool, so it lands at 4 rather than 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 9 tool updatesv1.7.2
    • Addedcompare_sites
    • Addedget_element
    • Addedget_index_status
    • Addedget_motion_dna
    • Addednew_winners
    • Addedsearch_elements
    • Addedsearch_motion
    • Changedsearch_sites1 field changed
      • addedInput schema / properties / responseMode
        Added value: +{
        +  "default": "full",
        +  "description": "full (default) returns all inline previews; compact returns concise cards and at most two page previews",
        +  "enum": [
        +    "full",
        +    "compact"
        +  ],
        +  "type": "string"
        +}
    • Addedwatch_site
  2. 3 tool updatesv1.6.0
    • Changedanalyze_page_structure1 field changed
      • addedInput schema / properties / viewport
        Added value: +{
        +  "default": "desktop",
        +  "description": "desktop = 1440x900 (default); mobile = 390x844 iPhone-class with deviceScaleFactor 3, isMobile + hasTouch",
        +  "enum": [
        +    "desktop",
        +    "mobile"
        +  ],
        +  "type": "string"
        +}
    • Changedcapture_live_site1 field changed
      • addedInput schema / properties / viewport
        Added value: +{
        +  "default": "desktop",
        +  "description": "desktop = 1440x900 (default); mobile = 390x844 iPhone-class with deviceScaleFactor 3, isMobile + hasTouch",
        +  "enum": [
        +    "desktop",
        +    "mobile"
        +  ],
        +  "type": "string"
        +}
    • Changedrecord_site_motion1 field changed
      • addedInput schema / properties / viewport
        Added value: +{
        +  "default": "desktop",
        +  "description": "desktop = 1440x900 (default); mobile = 390x844 iPhone-class with deviceScaleFactor 3, isMobile + hasTouch",
        +  "enum": [
        +    "desktop",
        +    "mobile"
        +  ],
        +  "type": "string"
        +}
  3. 7 tool updatesv1.0.0
    • First observedanalyze_page_structure
    • First observedcapture_live_site
    • First observedget_site_details
    • First observedget_site_elements
    • First observedlist_categories
    • First observedrecord_site_motion
    • First observedsearch_sites

TDQS

A3.9/5.0

Scored across 15 tools

Disambiguation5/5

Each tool targets a distinct resource/action: site search/details/compare, element gallery, motion DNA/search, live capture/structure/motion recording, and monitoring. Overlaps such as the search_* tools are separated by domain, and descriptions clarify outputs. No pair appears interchangeable.

Naming Consistency4/5

Names are consistently snake_case and mostly verb_noun, making actions predictable. The one minor deviation is new_winners, which is a noun phrase rather than verb_noun, but the overall convention is stable.

Tool Count5/5

15 tools is within a reasonable upper bound for a multi-capability design-research server. Each tool maps to a distinct capability such as search, detail, compare, live capture, motion, or monitoring, so none feels redundant.

Completeness4/5

Core lifecycle for design research is well covered: discovery, detail, comparison, elements, categories, live capture, motion analysis, and watchlist add/list/remove. Minor gaps remain, such as no dedicated new-elements feed or explicit element-category listing, but agents can work around them.

Maintenance

ActivityMaintained
ResponsivenessWithin a week

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI agents to see, analyze, and visually verify web page changes through pixel-perfect diffing, theme extraction, layout analysis, and interactive element detection.
    2 npm
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides curated real website design references with structured JSON data on type, spacing, palette, and layout. Enables AI agents to search, browse, and analyze over 1,000 sites and their sections.
    1
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Enables coding agents to capture and inspect running projects, resolve real assets from free sources, and check source for design tells, providing an evidence loop for improving websites, games, and motion.
    10
    1
    MIT