awwwards-mcp
awwwards-mcp lets AI agents search Awwwards for design inspiration and pull deep design, motion, and structure data from award-winning websites.
Search award-winning sites by free text, tags, color, technology, or award type, with inline screenshots and BM25-ranked local FTS5 results.
Get a site’s design DNA: palette, technologies, design elements, awards, and description.
Compare 2–3 sites’ design DNA and jury scores.
List all available search filters: 200+ tags, 27 colors, and technology slugs.
Fetch component-level visuals and videos for a site (3D models, mobile layouts, microcopy, etc.).
Capture fresh full-page screenshots of any live URL in desktop or mobile viewport.
Analyze a page’s section band map to compare reference sites against local builds.
Record short motion-through videos with inline filmstrip JPEGs.
Search and inspect the inspiration-elements gallery by title, author, category, or slug.
Analyze a live site’s runtime motion DNA and search previously captured motion scans.
Track new Awwwards winners and maintain persistent watches over studios, tags, or sites.
Check local index status and build/refresh a resumable local index for deeper searches.
Provides tools for searching Awwwards award-winning websites by color, tags, technology, award type, or free-text query; viewing site screenshots; and extracting design DNA such as color palettes, technologies, design elements, and award history. Can also capture, analyze, and record motion of live sites for design inspiration.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@awwwards-mcpFind dark 3D portfolio sites and summarize their design DNA."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
awwwards-mcp
Free, open-source MCP server that gives AI agents design inspiration from Awwwards — the Mobbin-style visual reference loop, sourced from the web's best award-winning websites.
Your agent searches in natural language ("dark 3D portfolio sites", "soft pastel e-commerce"), sees real screenshots inline, and can pull the design DNA of any site: color palette, tech stack, design elements, award history. Free-text queries run on a porter-stemmed, prefix-matching FTS5 index with BM25 ranking — "magazines" now finds Magazine-tagged sites (68 on the live index), best matches first, where the old substring path returned zero. Multi-word queries keep AND semantics: every token must hit the same site.
Tools
Tool | What it does |
| Search by color, tags, technology, award type or free-text query. Multi-word queries match against the local FTS5 index and rank BM25 (title hits lead); zero results come with loose-match and taxonomy-tag hints. Default full results include each site's screenshot; use |
| Full design DNA for one site: palette, technologies, elements, awards, description. |
| Compare 2–3 sites' design DNA and jury scores as text-only JSON. Uses cached details or fetches missing detail pages. |
| Read local index count, crawl progress, last success/error, lock state and freshness without network requests. |
| Component-level visuals for one site: each element's poster image inline (3D models, video content, mobile layouts, microcopy…) + video URLs. |
| Every filter the agent can search by (200+ tags, 27 colors). |
| Optional: fresh full-page screenshot of any live URL. Waits for |
| Section band map of any page (live URL or local file:// build): tag, background, offset, height per band. Compare a reference site's structure against your build. Same heavy-site-friendly wait ( |
| Optional: short motion-through video of a live URL — preloader, scroll-triggered and hover/cursor animations. Returns an inline filmstrip JPEG plus the saved .webm path. |
| Search the inspiration-elements gallery (footer, hero, pricing, 404…) by free text; ranks BM25 over title/author/category. |
| One element record: title, category, author, built-with stack, related elements and its media URL (image or video) pointing at awwwards' CDN. |
| Runtime motion fingerprint of a live URL: animation libraries, render engines, ScrollTrigger stats (trigger count, scrub ratio), tween easing/duration vocab and the scroll model. Fresh capture or cached capture with timestamp. |
| Search previously captured motion-DNA scans by library, scroll model or easing vocabulary — find references by how a site moves. |
| Poll today's freshly-crowned winners (SOTD / Developer Award / Honorable Mention) against a persisted baseline. First call seeds and dumps the listing; later calls report the delta. Each first-seen winner's Elements section is backfilled into the searchable element corpus, so new winners are element-searchable immediately. |
| Persistent watches over studios, tags or specific sites (add/list/remove). |
Data posture: element records store metadata + media URLs pointing at awwwards' own CDN — nothing is mirrored. Motion DNA records are local captures, each stamped with the time it was taken.
Related MCP server: A1 Gallery MCP Server
Setup
v1.0.0 — the first stable release. Any MCP-compatible coding agent can use awwwards-mcp — no API key, no account.
Requires Node ≥ 22.13 (node -v to check). Pick your agent:
Updates: the server checks the npm registry once a day and prints an
stderr notice when a newer awwwards-mcp exists (stdout stays clean for the
JSON-RPC channel — your agent sees the notice as a log line). Set
AWWWARDS_AUTO_UPDATE=1 in the server's env to opt into background
self-update; restart your agent afterwards to load it. Nothing is fetched
more than once a day and serving never waits on the check.
Claude Code
claude mcp add awwwards -- npx -y awwwards-mcpCodex CLI (ChatGPT desktop app and the IDE extension share this config)
codex mcp add awwwards -- npx -y awwwards-mcpor in ~/.codex/config.toml (project-scoped: .codex/config.toml):
[mcp_servers.awwwards]
command = "npx"
args = ["-y", "awwwards-mcp"]OpenCode (opencode.json — note the command is an array)
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"awwwards": {
"type": "local",
"command": ["npx", "-y", "awwwards-mcp"]
}
}
}ZCode (~/.zcode/cli/config.json — note servers nest under "mcp": { "servers": ... })
{
"mcp": {
"servers": {
"awwwards": { "command": "npx", "args": ["-y", "awwwards-mcp"], "env": {} }
}
}
}Claude Desktop / Cursor / Windsurf / Gemini CLI / Cline / Continue — anything
reading the common mcpServers JSON shape (e.g. ~/.claude/claude_desktop_config.json
or ~/.gemini/settings.json):
{
"mcpServers": {
"awwwards": { "command": "npx", "args": ["-y", "awwwards-mcp"] }
}
}Anything else — awwwards-mcp is a plain stdio MCP server: point your client
at npx -y awwwards-mcp and it works. To pin a version, use
npx -y awwwards-mcp@1.0.0.
pi coding agent has no built-in MCP by design — it uses skills and extensions instead. Two options:
Install the awwwards-inspiration skill (below). pi reads skills from
~/.pi/agent/skills/or~/.agents/skills/(the latter is shared across agents following the Agent Skills standard). The skill teaches the workflow; for it to reach the live data, add an MCP-supporting pi extension, or run the queries in another agent and paste results.Skip MCP entirely: ask pi to build you a small CLI wrapper around awwwards.com, or use a shared skills directory (
~/.agents/skills/) so the same skill file serves pi and every other agent.
Optional full-page captures (needed by capture_live_site,
analyze_page_structure, record_site_motion):
npm install -g playwright && npx playwright install chromiumrecord_site_motion additionally uses ffmpeg; it resolves the ffmpeg-static
package automatically if present.
Skills
This package ships four agent skills. Any agent that follows the Agent Skills standard can load them; copy them into your agent's skills directory:
npm install awwwards-mcp
mkdir -p ~/.agents/skills && cp -r node_modules/awwwards-mcp/skills/awwwards-inspiration node_modules/awwwards-mcp/skills/awwwards-setup node_modules/awwwards-mcp/skills/awwwards-doctor node_modules/awwwards-mcp/skills/awwwards-motion-study ~/.agents/skills/Skill | What it teaches |
| First-time onboarding: asks the user's preferences (result density, viewport, captures, local index, winner watches), persists them to |
| The inspiration loop: search, judge from screenshots, pull design DNA, state a design direction, capture/motion-first builds; staying current with |
| The full video chain: what to record from a live site (and what to skip), frame-by-frame review (video input or tile-per-element), the motion inventory, and build verification by re-recording. |
| Repair: run |
Agent | Skills directory |
Claude Code |
|
pi |
|
ZCode |
|
Agent Skills-standard agents |
|
Windows: run this from Git Bash, or copy
node_modules\awwwards-mcp\skills\awwwards-setup manually.
Indexing (recommended)
search_sites works out of the box, but its depth is limited by polite live
scraping (~31 sites per filter page). Build a local index once and searches
draw from thousands of award-winning sites instantly:
npx -y -p awwwards-mcp awwwards-index # once published
# or, from a local checkout of this repo:
npm run indexCrawls all ~200 tag pages at 1 request/second (~4 minutes) into the local SQLite cache at
~/.awwwards-mcp/.Resumable: interrupt it and re-run — completed pages are skipped.
The MCP server re-indexes automatically in the background whenever the index is older than 7 days (never blocking your session). Completed crawl checkpoints are cleared so each refresh actually revisits the tag pages. Run
get_index_statusto inspect progress or the last crawl error.
Elements index (optional)
The search_elements tool auto-indexes the first gallery page (~48 items)
on first use. To build a full corpus (~1,500+ items, ~15 pages at 48/page):
npm run index -- --elements all # follow pagination until exhausted
npm run index -- --elements 10 # first 10 listing pagesAny crawl beyond the first page also runs the taxonomy pass: every
facet page (/elements/footer/, /elements/cta/, … 46 categories) is
fetched once and each element under it is stamped with that category, so
search_elements' category filter works across the corpus. Element
pages carry no breadcrumb, so this listing-side pass is the only category
source; elements seen on no facet page stay unsorted.
Element rows are searched by title, author, category and slug tokens
(slug is an FTS5-indexed column; a cache opened from an older schema
version rebuilds its search index automatically on first open). Each
element also carries the slug of the award-winning site it came from
(siteSlug/siteUrl in search_elements/get_element results). Two
record sources share this corpus: gallery records (indexed from the public
elements listing) and source:"site" records backfilled from each new
SOTD winner's own Elements section by new_winners — their slugs are
namespaced site-<siteslug>-<title> so the two never collide. The
elements index has the same 7-day freshness gate as the sites index —
a re-run inside the window skips itself.
Site details (palettes, tech stacks) are still fetched on demand and cached for 7 days. Awwwards page and CDN requests have a 10-second deadline per attempt, including response-body reading; transient page failures are retried once, while blocks (403/429) and CDN failures are not retried.
How it works
Live, polite scraping of awwwards.com public pages (max 1 request/second, robots.txt-compliant paths only, cached 7 days in SQLite at
~/.awwwards-mcp/).Screenshots are served from Awwwards' own CDN (880×660), cached on disk.
No API key, no account, no cost.
Ethics & terms
This tool fetches publicly available pages for personal design-inspiration use, at human-ish request rates, honoring robots.txt. Awwwards' screenshots and content remain the property of Awwwards and the credited creators — don't bulk-scrape, redistribute, or republish them. If you use this commercially, review awwwards.com's terms yourself.
Built with awwwards-mcp: four real sites
Four complete sites were built through the full inspiration loop this MCP
enables, using nothing but the server's tools plus the shipped
awwwards-inspiration skill. Each one exercised a different corner of the
loop — and every correction the loop caught on the way became doctrine in the
skill.
Built in one shot, by a model that can't watch video. All three sites were built in a single prompt run on GLM 5.3-flash — which does not support video input. The loop's motion study worked entirely from frame-tiled filmstrips (ffmpeg, 1–2 fps per element) instead of watching the recordings. With a video-native model, those same
get_site_elementsvideos andrecord_site_motion.webm files could be watched directly — timing, easing and overlap read at full fidelity — and the motion-true results would be better still. The skill's frame-tile doctrine is what closes that gap today.
Watch the whole loop run (1:50):
Screen recording of the agent running the awwwards-inspiration loop end to
end with the awwwards MCP tools — searching SOTD references with inline
screenshots, pulling design DNA, frame-studying element videos, building, and
verifying with band maps + motion recording. If your client doesn't render
the player, watch the file directly.
1. Ridge
(source) — a Swiss-minimal single-page showcase for a fictional engineering-talent studio, direction Aspen Search (SOTD + Developer Award, jury 7.48): monochrome #FAFAF8/#1A1A1A + mint, giant grotesque section markers, halftone grain, asymmetric panel grid, dark discipline panels in an interior horizontal pin passage, count-up stats, client rows, theme toggle, cursor-follower. Built with the v1.4.0 toolkit: FTS5-ranked direction search, both-viewports reference captures and QA (desktop 7,849px + mobile 390×844), overflow audit (0px both), pin-center shots, film verification — and the skill-memory flywheel recorded the findings. QA evidence: showcase/ridge/_qa/.
The grain-panel hover, studied from aspensearch.com's recording and rebuilt as a canvas dither-dissolve — dots flip to mint around the mouse, the trail elongates, the boundary dissolves:

Panel grid (desktop) | Horizontal discipline passage | Mobile 390×844 |
|
|
|
2. Fallow Press
(source) — a flat-2D editorial journal, direction
Emergence Magazine (SOTD): pink #FF9398 on cream and black, torn-paper
masthead (pure CSS clip-path, zero WebGL), giant grotesque display over
grayscale photography, serif-italic brand, three pages with separate
horizontal projects/about pages (GSAP ScrollTrigger pin +
containerAnimation).
Torn-paper masthead (home) | Horizontal gallery (Fields) | Horizontal chapters (Practices) |
|
|
|
The loop as it ran:
search_sites(magazine filters) → shortlist judged from inline screenshots →get_site_detailson Emergence Magazine.Capture before building:
capture_live_site+record_site_motionon the live site first; full-page PNG and motion .webm kept in fallow-press/ref-motion/ as the evidence trail.Build, then verify: full-page capture plus panel-center pin shots of both horizontal pages (13 stops each, in fallow-press/_qa/ —
capture-qa.mjsis reusable).The pin shots caught a real bug: horizontal-panel entrances used
toggleActions: "play none none reverse", and 100vw panels hide content at midpoints on the way back — copy disappeared mid-view. Fix (one-shot play entrances) is now doctrine: full-viewport panels get one-shot entrances; QA pin shots land at panel centers, not uniform fractions, or you photograph empty transition zones.
3. Cerebrium recreation (C:/Users/Afjal/cerebrium-recreation/) — a
fidelity-first recreation of cerebrium.ai, pixel-checked against the live
reference: full-page captures of both sides, analyze_page_structure band
compare, and SVG icon/legend fixes until the build matched the reference to
within 1px of total page height (10,871px vs 10,870px). This is the
structure-before-pixels doctrine at its strictest — band maps compared,
never just totals.

4. The Meridian (C:/Users/Afjal/editorial-site/) — an editorial journal
built from ORDR/Hearst references: the first build to run the whole loop
end-to-end. analyze_page_structure caught a masthead band bug by comparing
the build's band map against the reference's; the reference captures,
motion film, and the reusable pre-scroll capture script live in
editorial-site/_qa/.


Skills used to build these
Skill | Role in the builds |
| The 8-step loop itself (ships with this package): search → judge from screenshots → design DNA → capture/motion study → state direction → build → band-map verify. |
| The horizontal pin + |
| Tween composition and sequenced hero entrances (torn-paper drop, panel copy rises). |
| Typography, palette and layout judgment applied when translating reference DNA into original pages. |
| smooth scrolling synced to ScrollTrigger on the Fallow Press home page. |
| All builds are plain hand-rolled CSS — flat 2D, no frameworks needed. |
The skills self-improve: every loop pass records what verification caught (scripts/skill-memory.mjs record), and a deterministic distiller folds rules seen 2+ times into your installed skill copy — while the shipped copies only change via human PR. A techniques registry (skills/_memory/techniques.json) catalogs researched how-tos per domain (video understanding, motion detection, UI structure, micro-interactions, images).
Reduced-motion, JS-less visits, and capture tools all get graceful fallbacks (vertical stacks; progressive-enhancement reveals).
What the verification loop caught — proof the structure-before-pixels doctrine is load-bearing:
Element posters lie: the first showcase build was designed from poster frames alone and rendered a spinning 3D ring as floating static cards. Downloading the element videos (
get_site_elements) and frame-tiling them revealed the motion truth — now the skill mandates studying motion before animating.Full-page captures of reveal-on-scroll builds showed blank sections:
.revealanimation state vs capture's no-scroll reality. Builds ship content-visible-without-JS progressive enhancement.Horizontal-panel copy vanished mid-view on the Fallow Press pages:
toggleActionsreverse reverts entrances while a 100vw panel is still holding the viewport (see above).Band-map compare kept the references' rhythm instead of drifting on section heights (Cerebrium, The Meridian).
Prompt counts: 3 for the original showcase build (the build ask, the
motion correction that exposed the poster-lie, the structure pass) and
1 for Fallow Press ("create a new website using our MCP and skills… no
3D websites") — its two follow-ups were caught by the QA loop, not by the
user. Each correction became doctrine in the shipped awwwards-inspiration
skill: frame-study element videos before animating; judge page architecture
from the studied passages; tile per element, not one giant filmstrip;
capture live sites and animation before building.
Can awwwards-mcp crawl the sitemap? (robots.txt notes)
The awwwards.com robots.txt advertises
Sitemap: https://www.awwwards.com/sitemap.xml and — verified live
2026-09-18 — that sitemap URL returns a soft-404 HTML page (as do common
child names like /sitemap-websites.xml). So sitemap discovery isn't
currently a path to more data; the polite crawl surface is exactly what the
indexer uses:
Allowed and used:
/websites/,/websites/<filter>/,/sites/<slug>(one filter per URL; deep pagination stays un-crawled).Disallowed and never fetched:
/tag/,/search-websites,/websites/?(query-string pagination),/elements/*,/vote/, favourites/likes/follows, and the rest of the 33 rules.Our client (
src/awwwards.tsbuildFilterUrl) constructs only/websites/…paths at 1 request/second — the loop stays inside the published rules by construction, not by convention.
Contributing
PRs welcome! The project especially needs parser-drift fixes — when live awwwards.com markup changes, a fresh HTML snapshot attached to an issue often becomes the new test fixture and the fastest merged PR. See CONTRIBUTING.md for the full guide:
Parser-drift is monitored automatically. A probe script
(scripts/parser-drift-probe.mjs,
npm run drift) checks every markup anchor the parsers depend on — the
split/indexOf/regex literals in src/parsers.ts — against the live
listing and detail pages (2 fetches, 1 request/second, same politeness as
the client). A daily GitHub Action (.github/workflows/parser-drift.yml)
runs it and, on drift, opens/updates a single tracking issue with the exact
anchors that changed (and auto-closes it when a later run is green). To run
it yourself: npm run drift (live, exit code 0/1/2) or npm run drift -- --fixture
(offline, checks the committed fixtures still feed every anchor). Raw HTML
is never diffed or stored — anchors only fire when the parsers actually
break, so there are no false alarms from cosmetic tweaks.
Development setup & project layout (offline fixture-tested, no network in tests)
How to create a PR: fork →
fix//feat//docs/branch → typecheck + tests → PR templateThe politeness constraints new code must keep (1 req/s, robots.txt paths, light runtime deps)
Bugs and feature ideas start as issues with templates. Security problems go privately — see SECURITY.md. By participating you agree to the Code of Conduct.
Development
npm install
npm test # offline unit tests against committed HTML fixtures
npm run smoke # manual live smoke test against awwwards.com
npm run build # compile to dist/MIT — see LICENSE.
Available Tools
15 toolsanalyze_page_structureARead-onlyIdempotent
Extract a page's section band map (tag, label, background color, offset, height per band) via a headless browser. Works on live URLs and file:// paths — use it to compare a reference site's structure against your local build.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Absolute URL (https:// or file://) of the page to analyze | |
| maxBands | No | Cap on returned bands | |
| viewport | No | desktop = 1440x900 (default); mobile = 390x844 iPhone-class with deviceScaleFactor 3, isMobile + hasTouch | desktop |
| waitStrategy | No | 'load' + settle works on heavy sites; 'networkidle' waits for total quiet | load |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, open-world). The description adds non-obvious behavioral context: it runs via a headless browser and accepts local file:// paths, which tells the agent this fetches and renders remote/local content. It stops short of noting timing/cost or pagination behavior, but that gap is minor given the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the action and output shape, followed by scope and intended use. Every clause earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully compensates by enumerating the per-band fields returned. Combined with the schema's full parameter coverage, an agent has enough to call it correctly, though return-format details (ordering, error shape) are not spelled out.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents url, maxBands, viewport, and waitStrategy in detail. The description's field list describes the OUTPUT shape rather than adding meaning to the parameters, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (extract) and resource (a page's section band map) and enumerates the returned fields (tag, label, background color, offset, height per band). This clearly distinguishes it from siblings like capture_live_site, get_site_elements, and compare_sites, which operate on different data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names a concrete use case — comparing a reference site's structure against a local build — and clarifies it accepts both live URLs and file:// paths. However, it does not explicitly state when NOT to use it or name the sibling alternative for the comparison workflow (e.g. compare_sites).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capture_live_siteARead-onlyIdempotent
Take a fresh full-page screenshot of a live website URL using a headless browser. Requires the optional playwright dependency.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Absolute URL of the site to capture | |
| viewport | No | desktop = 1440x900 (default); mobile = 390x844 iPhone-class with deviceScaleFactor 3, isMobile + hasTouch | desktop |
| waitStrategy | No | 'load' + settle works on heavy sites; 'networkidle' waits for total quiet | load |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description's only added behavioral fact is the headless-browser mechanism and the playwright dependency; it says nothing about timeouts, rate limits, or failure behavior on unreachable URLs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action, zero filler. The dependency caveat is the only secondary detail and it is stated tersely.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does not indicate what a successful capture returns (image bytes, file path, URL). For a network-facing capture tool that omission is a real gap, though annotations and the rich schema cover the input side adequately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so url, viewport and waitStrategy are fully documented in the schema, including enum meanings and defaults. The description adds nothing beyond that, which is the expected baseline when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Take a fresh full-page screenshot of a live website URL') plus the mechanism ('headless browser'). It is clearly distinguishable from siblings like analyze_page_structure or record_site_motion, though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the name and the verb. The one piece of guidance given ('Requires the optional playwright dependency') is a prerequisite rather than a when-to-use rule, and no alternative tool is mentioned for overlapping tasks like structure analysis.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
compare_sitesARead-onlyIdempotent
Compare the design DNA of 2–3 Awwwards sites: titles, live URLs, palettes, technologies, elements, awards and jury scores. Text only; missing details are fetched and cached.
| Name | Required | Description | Default |
|---|---|---|---|
| slugs | Yes | Two or three distinct site slugs from search_sites |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/openWorld/non-destructive, so the safety profile is covered. The description adds genuinely new context: output is text only (no imagery), and missing details are fetched and cached — implying network calls and possible latency. It stops short of stating caching lifetime or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the comparison scope, then a second sentence covering output mode and data-fetch behavior. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so enumerating the returned facets in the description is exactly what's needed; combined with the text-only and fetch/cache notes, an agent has enough to invoke and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single param (slugs) is fully documented in the schema, including the 2–3 distinct-slug constraint that mirrors min/maxItems. The description restates the count but adds no format guidance beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Compare) and resource (design DNA of Awwwards sites) with explicit scope (2–3 sites) and enumerates the compared facets (titles, URLs, palettes, technologies, elements, awards, jury scores). This clearly distinguishes it from the single-site sibling get_site_details.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 2–3 site scope implies this is the multi-site comparison tool versus get_site_details for one site, but it never states when to prefer it or what to do for a single site. Usage is inferable rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_elementARead-onlyIdempotent
Get one Awwwards element in full: title, author, built-with stack, media URLs, related elements and the parent award-winning project. Use search_elements first for ids.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Element slug from search_elements |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, non-destructive, and closed-world, so the safety profile is covered. The description adds real value by disclosing the shape of the returned content (related elements, parent project), which is otherwise unknown since there is no output schema. Error behavior and any lookup limits are not mentioned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler; the payload enumeration comes first and the prerequisite is placed last as an actionable instruction. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-param getter with no output schema, the description supplies the return-field inventory and the id-acquisition path, which is most of what an agent needs. It omits failure behavior for unknown slugs, a minor gap given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter and schema description coverage is 100% ('Element slug from search_elements'), so the schema already carries the semantics. The description reinforces the id's origin but adds no format or syntax detail beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (get) plus resource (one Awwwards element) with an explicit enumeration of the returned payload: title, author, built-with stack, media URLs, related elements, and parent project. It separates itself from the sibling search_elements by being the single-item 'full detail' endpoint.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the agent to search_elements first to obtain ids, which is the correct prerequisite for this tool. It stops short of stating when NOT to use it or what happens with an invalid/unknown slug, so it is clear context rather than full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_index_statusARead-onlyIdempotent
Read offline index status: cached site count, progress, last successful finish and error, lock state and freshness. No live requests.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and closed-world, so the safety profile is covered. The description adds the meaningful operational detail that this is an offline cache read with no live requests, plus a lock state indicator — useful context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that lists the return contents. Every clause earns its place; no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the return-value burden — and it does, enumerating cached site count, progress, last successful finish, error, lock state and freshness. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4; there is nothing to mis-document and the description correctly implies a parameterless read.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read offline index status') and enumerates exactly what the status comprises: cached site count, progress, last finish/error, lock state, freshness. No sibling tool overlaps with index status, so the agent can route here unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'No live requests' signals this reads cached/offline state rather than triggering network activity, which implicitly contrasts with the capture_live_site / watch_site siblings, but the description never says when to prefer this tool over others or any prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_motion_dnaARead-onlyIdempotent
Get the Motion DNA of a website: animation stack, ScrollTrigger/pin/scrub counts, easing vocabulary and duration distribution. Serves the corpus first; live-captures unseen URLs (headless browser required).
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Absolute site URL | |
| recapture | No | Force a fresh live capture even if a recent record exists |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint. The description adds genuinely useful operational context beyond them: it serves a cached corpus first and only then performs a live headless-browser capture, implying network access and variable latency for unseen URLs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler; the output inventory comes first and the corpus/live-capture behavior follows. Every clause conveys usable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description enumerates the returned fields, which compensates. Combined with annotations covering safety and idempotency and the stated headless-browser fallback, an agent has enough to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (url and recapture are both documented in the schema), so the schema carries parameter meaning. The description does not add syntax or format detail for either parameter, nor does it cross-reference the recapture flag against the corpus-first behavior it describes. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get the Motion DNA of a website') and enumerates the returned content (animation stack, ScrollTrigger/pin/scrub counts, easing vocabulary, duration distribution). This clearly separates it from search_motion or record_site_motion, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: 'Serves the corpus first; live-captures unseen URLs' tells the agent when a live capture happens, but there is no explicit when-to-use guidance versus siblings like search_motion or capture_live_site, and no stated exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_site_detailsARead-onlyIdempotent
Get the design DNA of one Awwwards site: color palette, technologies, design elements, awards, description and inline screenshot.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | Yes | Site slug from search_sites, e.g. 'l-i-s-a' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, open-world behavior, so the safety profile is covered. The description adds value beyond that by disclosing the actual returned content (palette, tech stack, elements, awards, inline screenshot), which matters since there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the verb and resource front-loaded and zero filler; every clause (palette, technologies, elements, awards, description, screenshot) earns its place by telling the agent what comes back.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read tool with full annotation coverage, the definition is nearly complete: it states the payload and the schema supplies slug format and origin. The one gap is the absence of any usage/routing guidance relative to siblings such as search_sites.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the slug property already carries a pattern constraint, a concrete example ('l-i-s-a') and its origin tool. The description adds nothing further about the parameter, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (Get) plus resource (one Awwwards site) and an explicit enumeration of the payload — color palette, technologies, design elements, awards, description, screenshot. That enumeration usefully separates it from siblings like get_site_elements or get_motion_dna, though no sibling is named outright.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use, no prerequisites, and no alternatives. The only workflow cue ('slug from search_sites') lives in the input schema, not in the description, so an agent gets no routing guidance from the text itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_site_elementsARead-onlyIdempotent
Get the design-element highlights of one Awwwards site: component-level visuals (3D models, video content, mobile layouts, microcopy) with poster images inline and video URLs.
| Name | Required | Description | Default |
|---|---|---|---|
| slug | Yes | Site slug from search_sites, e.g. 'l-i-s-a' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, open-world), so the bar is lower. The description adds genuine context the annotations cannot: posters are returned inline while video is delivered as URLs, which shapes how the agent handles the payload. It omits pagination or result-limit behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the leading clause carries the action and the trailing parenthetical enumerates content types. Slightly dense with the nested list, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully characterizes the return payload (inline posters, video URLs). It stops short of covering volume, ordering, or empty-result behavior for a site with no elements, which is the main remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and even supplies a format hint ('l-i-s-a'), so the parameter is fully documented structurally. The description adds no syntax or sourcing detail beyond implying a single-site identifier, making the baseline 3 correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get the design-element highlights of one Awwwards site') and enumerates the component types returned (3D models, video content, mobile layouts, microcopy). The single-site scope implicitly separates it from search_elements/get_element, though no sibling is named outright.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: 'of one Awwwards site' signals per-site retrieval, but there is no explicit when-to-use, when-not, or pointer to alternatives such as search_elements for cross-site queries. An agent must infer the routing itself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_categoriesARead-onlyIdempotent
List the filter taxonomy available on Awwwards: color hexes and tag/technology slugs usable with search_sites.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive). The description adds value beyond them by disclosing what the return content actually is (color hexes, tag/technology slugs) and how it feeds into search_sites, which is the real behavioral insight for a zero-arg lookup tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. The resource is named first, the contents second, and the downstream use last—no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters, no output schema, and annotations covering safety, the description is nearly self-sufficient for such a simple tool. It could optionally note the output shape (e.g., grouped lists) but nothing essential to correct invocation is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so there is nothing for the description to disambiguate; the baseline for a zero-param tool applies. No parameter meaning is missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (filter taxonomy on Awwwards), then enumerates the concrete contents (color hexes and tag/technology slugs). An agent can immediately distinguish this discovery tool from sibling search tools like search_sites.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase "usable with search_sites" makes the purpose context explicit: this is the tool to call first to learn valid filter values before searching. It does not spell out a when-not condition, but the intended workflow is clear from the linkage to a named sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
new_winnersARead-only
Poll today's new Awwwards winners as a delta against the previous poll: new-entrant cards since the last call (first call seeds the baseline and reports no delta). Poll daily for a winners feed; pairs with watch_site for structured monitoring.
| Name | Required | Description | Default |
|---|---|---|---|
| award | No | Which winners feed to poll | sotd |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, openWorldHint, destructiveHint=false, and the key idempotentHint=false. The description earns credit by explaining that non-idempotence: the first call seeds the baseline and returns no delta, and later calls return only new entrants. That stateful, call-order-dependent behavior is exactly what an agent needs and is not derivable from the annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the core action and delta semantics before the polling cadence and sibling pairing. Every sentence carries information, though the two polling-related clauses overlap slightly in intent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must convey the return shape, and it does so adequately by defining the result as new-entrant cards since the last call, including the baseline-seeding edge case. It omits error/rate-limit behavior, which is a minor gap for a polling tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'award' enum is fully documented in the schema, so the baseline is 3. The description never mentions the award parameter or how it affects the delta, adding nothing beyond the structured schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Poll'), a precise resource ('today's new Awwwards winners'), and the exact mechanism ('as a delta against the previous poll'). It is immediately distinguishable from the read-oriented siblings like search_sites or get_site_details, and it explicitly names its closest relative, watch_site.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear usage context ('Poll daily for a winners feed') and routes the agent to the sibling tool for a related mode ('pairs with watch_site for structured monitoring'). It stops short of stating an explicit when-not condition or full alternatives, so it is strong but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_site_motionARead-onlyIdempotent
Record a short motion-through video of a live website — preloader, scroll-triggered and hover/cursor animations — and return an inline filmstrip JPEG plus the .webm path. Requires the optional playwright and ffmpeg-static dependencies.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Absolute URL of the site to record | |
| frames | No | Filmstrip tile count (default 16 → a 4x4 grid) | |
| viewport | No | desktop = 1440x900 (default); mobile = 390x844 iPhone-class with deviceScaleFactor 3, isMobile + hasTouch | desktop |
| waitStrategy | No | 'load' + settle works on heavy sites; 'networkidle' waits for total quiet | load |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this a safe, idempotent, open-world read. The description adds genuinely useful behavioral context beyond them: what the capture covers, that it returns an inline filmstrip JPEG plus a .webm path, and that it requires the optional playwright and ffmpeg-static dependencies. It does not state duration limits, timeouts, or failure behavior for heavy sites, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences: the first front-loads purpose and scope, the second covers output and dependency requirements. No filler, and the most important information leads.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly carries the return format (filmstrip JPEG plus .webm path) and runtime prerequisites. With 100% schema coverage the parameters are handled elsewhere. The only real gap is sibling/usage discrimination, which keeps it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already explains url, frames, viewport, and waitStrategy fully. The description adds no parameter-level detail (e.g., how long a 'short' recording is or how frame count affects runtime), so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource — recording a short motion-through video of a live website — and enumerates the exact animation types captured (preloader, scroll-triggered, hover/cursor). It is unambiguous on its own, but it never names or contrasts with close siblings like capture_live_site or watch_site, so differentiation is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no exclusions, and no routing to alternatives such as capture_live_site (static capture) or watch_site (monitoring). An agent must guess which recording/capture sibling fits a given request.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_elementsARead-onlyIdempotent
Search Awwwards' curated element gallery — award-grade UI components (footers, menus, loaders, transitions, micro-interactions…). Usage: pass query text and at most one of category/stack; every result links back to its source element page and parent project.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | No | Free text: what the component is or does | |
| stack | No | Built-with tokens, e.g. ['gsap', 'webgl'] | |
| category | No | Category id from the gallery taxonomy, e.g. 'footer', 'menu', 'cta', 'loading', 'mouse_interaction', '404_page' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, open-world, non-destructive behavior, so the safety profile is covered. The description adds two things annotations cannot: the mutual-exclusivity rule for category/stack and the fact that every result links back to its source element page and parent project, which tells the agent what it gets back.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences: the first names the resource, the second front-loads the usage rule. No filler, no restatement of the name or title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description responsibly notes the shape of results (links to source element page and parent project). Combined with the input constraint, an agent has enough to call correctly; only the limit parameter's tuning guidance is left to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, so the schema already documents query, stack, and category. The description goes beyond it by stating that at most one of category/stack may be supplied, a constraint not encoded in the schema, which materially affects how the agent builds the call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (Awwwards' curated element gallery) and clarifies the resource with concrete examples (footers, menus, loaders, transitions). An agent can see this is the component-level search, distinct from search_sites and search_motion, though it never explicitly names those siblings to sharpen the contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Usage:' clause gives a real input constraint (query text plus at most one of category/stack), which is actionable. However, it says nothing about when to prefer this tool over get_site_elements, get_element, or search_sites, so the when-to-use decision is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_motionARead-onlyIdempotent
Search captured Motion DNA records: find sites that pin sections, scrub scroll-driven animation, or run a given animation library. Corpus only — live-capture happens via get_motion_dna.
| Name | Required | Description | Default |
|---|---|---|---|
| lib | No | Library token, e.g. 'gsap', 'lenis', 'webgl' | |
| limit | No | ||
| hasPins | No | Only sites that pin sections | |
| scrubOnly | No | Only sites with substantial scroll-scrubbed animation |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and closed-world, covering the safety profile. The description's 'Corpus only' line reinforces the closed-corpus behavior and adds the live-capture contrast, which is useful but largely consistent with openWorldHint=false. Nothing is said about result shape, ranking, or how limit interacts with the corpus.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the verb and resource, then a single routing clause. No filler and nothing repeated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should convey what comes back; 'find sites that...' implies site-level results, which is adequate for a four-param, all-optional search tool. Missing details are minor — result ordering and pagination behavior relative to limit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75% (lib, hasPins, scrubOnly documented), and the description largely restates those same semantics rather than adding format, token, or matching detail. The undocumented 'limit' is self-explanatory via min/max. Baseline 3 fits when the schema already carries the parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Search) and resource (captured Motion DNA records) and then concretizes the searchable properties: pinned sections, scroll-scrubbed animation, animation libraries. It also contrasts itself with the sibling get_motion_dna, so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly scopes usage to the existing corpus and names the alternative for the live path ('live-capture happens via get_motion_dna'), which is a real when-to-use/when-not split. It does not, however, distinguish itself from search_sites or search_elements, so the routing guidance is clear but incomplete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_sitesBRead-only
Search award-winning websites on Awwwards. Returns site cards with inline screenshots, live URLs, awards and tags.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | ||
| tags | No | Tag slugs, e.g. ['3d', 'portfolio'] | |
| award | No | ||
| color | No | Dominant color hex, e.g. '#404040' | |
| count | No | ||
| query | No | Free text matched against site titles and tags | |
| sortBy | No | Sort results: by Awwwards jury score (details previously fetched) or newest first | newest |
| technology | No | Technology slug, e.g. 'webgl', 'gsap', 'astro' | |
| responseMode | No | full (default) returns all inline previews; compact returns concise cards and at most two page previews | full |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the read-only, non-destructive safety profile, so the bar is lower. The description adds useful behavioral value by describing the return shape (site cards with inline screenshots, live URLs, awards, tags), but says nothing about pagination, result limits, or the effect of responseMode on payload size.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, with the primary action stated first and the return contents immediately after. Nothing in it can be cut without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, describing the card contents is necessary and present, which helps. However, a 9-parameter search tool with no usage scoping and undocumented pagination/count behavior is only minimally covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Across 9 parameters the description adds no filtering semantics at all: it only mentions tags and awards as things returned in cards, not as query filters. With 67% schema coverage, page, award, and count have no documented meaning in either place, so the description fails to compensate for the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ('Search award-winning websites on Awwwards') with the return payload named, so the agent can tell it is a query tool rather than a detail/fetch tool. It does not explicitly contrast itself with siblings like search_elements, new_winners, or get_site_details.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no named alternatives. The agent must infer from the name that this is the entry point before drilling into get_site_details or compare_sites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
watch_siteA
Longitudinal watchlist over Awwwards listings: track a studio (e.g. 'ToyFight'), a tag/technology slug (e.g. 'webgl'), or a site URL-slug for changes and new entries. add/list/remove; the delta is reported by watch_site list after new_winners or search_sites refreshes the listing.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | studio = creator name from listing cards; tag = tag/technology slug; url = a /sites/<slug> site-slug | |
| note | No | Why you're watching (free text) | |
| award | No | Restrict a watch to one award feed | |
| action | Yes | ||
| pattern | No | What to watch (name/slug per kind) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false and openWorldHint=true, so the safety profile is covered. The description adds non-obvious behavioral context beyond them: the mutation actions (add/list/remove) and the dependency that deltas only surface after new_winners or search_sites refreshes the underlying listing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly packed sentences front-load the resource and the three target kinds before the action list and refresh dependency. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and five parameters, the description adequately covers the targets, the three actions, and how deltas are produced. It leaves some ambiguity about exactly what 'list' returns and how the award filter interacts, though the schema covers those fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 80%, so the schema already documents kind, note, award and pattern. The description reinforces the kind semantics with concrete examples ('ToyFight', 'webgl', URL-slug), adding modest value, but introduces no syntax or format detail the schema lacks. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — a longitudinal watchlist over Awwwards listings — and enumerates the three watchable targets (studio, tag, url) with concrete examples. It is clearly distinguishable from sibling tools like search_sites or new_winners without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear workflow context: watches are managed via add/list/remove, and 'the delta is reported by watch_site list after new_winners or search_sites refreshes the listing,' explicitly naming the sibling calls that feed it. It stops short of stating when NOT to use the tool, so it lands at 4 rather than 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v1.7.2- Added
compare_sites - Added
get_element - Added
get_index_status - Added
get_motion_dna - Added
new_winners - Added
search_elements - Added
search_motion - Changed
search_sites1 field changed- added
Input schema / properties / responseModeAdded value: +{ + "default": "full", + "description": "full (default) returns all inline previews; compact returns concise cards and at most two page previews", + "enum": [ + "full", + "compact" + ], + "type": "string" +}
- Added
watch_site
3 tool updates
v1.6.0- Changed
analyze_page_structure1 field changed- added
Input schema / properties / viewportAdded value: +{ + "default": "desktop", + "description": "desktop = 1440x900 (default); mobile = 390x844 iPhone-class with deviceScaleFactor 3, isMobile + hasTouch", + "enum": [ + "desktop", + "mobile" + ], + "type": "string" +}
- Changed
capture_live_site1 field changed- added
Input schema / properties / viewportAdded value: +{ + "default": "desktop", + "description": "desktop = 1440x900 (default); mobile = 390x844 iPhone-class with deviceScaleFactor 3, isMobile + hasTouch", + "enum": [ + "desktop", + "mobile" + ], + "type": "string" +}
- Changed
record_site_motion1 field changed- added
Input schema / properties / viewportAdded value: +{ + "default": "desktop", + "description": "desktop = 1440x900 (default); mobile = 390x844 iPhone-class with deviceScaleFactor 3, isMobile + hasTouch", + "enum": [ + "desktop", + "mobile" + ], + "type": "string" +}
7 tool updates
v1.0.0- First observed
analyze_page_structure - First observed
capture_live_site - First observed
get_site_details - First observed
get_site_elements - First observed
list_categories - First observed
record_site_motion - First observed
search_sites
TDQS
Scored across 15 tools
Each tool targets a distinct resource/action: site search/details/compare, element gallery, motion DNA/search, live capture/structure/motion recording, and monitoring. Overlaps such as the search_* tools are separated by domain, and descriptions clarify outputs. No pair appears interchangeable.
Names are consistently snake_case and mostly verb_noun, making actions predictable. The one minor deviation is new_winners, which is a noun phrase rather than verb_noun, but the overall convention is stable.
15 tools is within a reasonable upper bound for a multi-capability design-research server. Each tool maps to a distinct capability such as search, detail, compare, live capture, motion, or monitoring, so none feels redundant.
Core lifecycle for design research is well covered: discovery, detail, comparison, elements, categories, live capture, motion analysis, and watchlist add/list/remove. Minor gaps remain, such as no dedicated new-elements feed or explicit element-category listing, but agents can work around them.
Maintenance
Related MCP Connectors
A design-style library for AI agents: search real styles, fetch a ready-to-apply design spec.
UI design from prompts, screenshots, and URLs for AI coding agents and theme tokens.
Give your AI agents a design superpower. Generate, edit, and publish publication-grade decks, reports, landing pages, resumes, and marketing visuals directly within your agent workflow. Delivering frontier-level design quality at 3× the speed and 53× lower cost -from conversational prompt to live link or vector PDF in minutes.
- miromiroOAuthapp.miromiro
Turn any live website into brand colors, fonts, design tokens, SVGs, Lottie and paste-ready code.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to see, analyze, and visually verify web page changes through pixel-perfect diffing, theme extraction, layout analysis, and interactive element detection.2 npmMIT
- AlicenseNot gradedqualityCmaintenanceProvides curated real website design references with structured JSON data on type, spacing, palette, and layout. Enables AI agents to search, browse, and analyze over 1,000 sites and their sections.1MIT
- AlicenseNot gradedqualityCmaintenanceEnables coding agents to turn CollectUI links or any image/GIF/video URL into a detailed design and motion brief with measured timings, easing curves, displacement, palette, and opt-in keyframes.7 npm1MIT
- AlicenseAqualityAmaintenanceEnables coding agents to capture and inspect running projects, resolve real assets from free sources, and check source for design tells, providing an evidence loop for improving websites, games, and motion.101MIT





