Ship Check
Server Details
Scan a deployed app URL for exposed keys, open Supabase tables and missing security headers.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 11 tools
Most tools have clearly distinct actions and resources, especially the booking, messaging, scan, and estimate workflows. However, leak_check is an explicit duplicate of ship_check, and the three cost-estimate tools plus request_estimate/list_services have related pricing purposes that could cause occasional misselection.
All names use snake_case consistently, which is the most important factor. The pattern is mostly verb_noun for action tools, with a few noun-style names for scan and cost-estimate tools, but the deviations are minor and readable.
With 11 tools, the server is within a reasonable range and covers its scan, consulting, booking, and estimate workflows without obvious bloat. One tool, leak_check, is a legacy alias rather than a new capability, so it is not fully earning its place.
The surface covers the core lifecycle: scan an app, request a human review, book or message Matt, list services, and get estimates. Minor gaps exist around retrieving or managing prior bookings directly, but the returned reschedule/cancel links provide a workaround.
Available Tools
11 toolsagent_cost_estimateAgent cost per shipped taskBRead-onlyIdempotentInspect
Estimate blended cost per shipped task for a month of agent runs, given a retry rate.
| Name | Required | Description | Default |
|---|---|---|---|
| modelId | No | Model id to scale avgCostPerTask by, defaults to sonnet-5 | |
| retryRatePct | Yes | ||
| runsPerMonth | Yes | ||
| avgCostPerTask | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered without the description. The description adds that the estimate is "blended" and conditioned on a retry rate, which is modestly useful context, but it says nothing about how the retry rate affects the result or what the returned figure represents.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler, front-loaded with the verb and the output quantity, with the conditioning input stated at the end. Nothing can be trimmed without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 4 parameters, 25% schema coverage, and no output schema, the description should carry more weight than it does. It never states the units for retryRatePct or avgCostPerTask, nor what the returned estimate looks like (single figure, breakdown, currency), leaving real ambiguity for a numeric-estimation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25% – modelId is documented, but runsPerMonth, avgCostPerTask, and retryRatePct have no descriptions. The description names "retry rate" but does not resolve the critical unit ambiguity (percent vs fraction) or clarify whether avgCostPerTask is per attempt or per shipped task.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Estimate") and a precise resource ("blended cost per shipped task for a month of agent runs"), which distinguishes it in substance from siblings like cursor_auto_cost_estimate and swarm_run_cost_estimate. It does not name those siblings or explain how the estimates differ, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of when another estimator (cursor_auto_cost_estimate, swarm_run_cost_estimate) would be the better choice. The only signal is the implied context of "agent runs," which the agent must infer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
book_callBook a call with MattAInspect
Book a free 30-minute intro call with Matt Turley (Continuum) directly on his calendar. The booking is confirmed immediately and Cal.com emails the invite to the user and to Matt. First call get_availability and pass the exact start of an open slot. Only book when the user has asked for the call and agreed to the time; pass their real name, email and time zone. notes: a short, factual summary of what the user wants to discuss (shown to Matt). Returns the booking id, meeting link and reschedule/cancel links. Limit 2 upcoming calls per email.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The user's name. | |
| Yes | The user's own email; the invite goes there. | ||
| event | No | intro-30 | |
| notes | No | What the user wants to discuss, up to 1000 characters. | |
| start | Yes | Slot start from get_availability, ISO 8601 with Z or an offset, e.g. 2026-10-05T15:00:00Z. | |
| timezone | Yes | The user's IANA time zone, e.g. America/New_York. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial context beyond annotations: immediate confirmation, Cal.com emailing both parties, the returned booking id/meeting link/reschedule-cancel links, and a concrete limit of 2 upcoming calls per email. Only minor gaps remain, such as what happens if the slot is taken between availability and booking.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense and front-loaded: purpose, mechanics, prerequisite, constraints and return values in five sentences with no filler. Every sentence carries operational information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although there is no output schema, the description enumerates the returned booking id, meeting link and reschedule/cancel links. Combined with the prerequisite, constraints and rate limit, an agent has everything needed to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 83% schema coverage the baseline is 3, but the description adds real meaning: notes is 'a short, factual summary of what the user wants to discuss (shown to Matt)' and the caller must pass the user's real name, email and time zone. It reinforces the ISO slot-start requirement already in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource with scope: 'Book a free 30-minute intro call with Matt Turley (Continuum) directly on his calendar.' It is clearly distinguishable from siblings like get_availability, which it names as the required predecessor.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use and prerequisites: call get_availability first and pass the exact slot start, and only book after the user has asked for the call and agreed to a time. It also names the sibling to consult beforehand.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cursor_auto_cost_estimateCursor Auto vs pinned model costCRead-onlyIdempotentInspect
Estimate Cursor Auto blended cost versus pinning a single model.
| Name | Required | Description | Default |
|---|---|---|---|
| autoModelMix | Yes | Map of model id to weight, must sum to 1 | |
| pinnedModelId | Yes | ||
| avgInputTokens | Yes | ||
| requestsPerDay | Yes | ||
| avgOutputTokens | Yes | ||
| workDaysPerMonth | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false, so the safety profile is covered structurally. The description adds nothing beyond that: it does not say the tool is a pure offline calculation, what assumptions/units it uses, or whether the result is a single figure or a breakdown.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; the comparison at stake is stated up front. It is efficient, though arguably under-sized for a six-parameter calculation tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With six required parameters, 17% schema coverage, a nested object, and no output schema, the description should carry far more explanatory weight. One sentence leaves the agent without enough information to populate or interpret this tool confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 17% — just one of six required parameters (autoModelMix) is documented, and its 'must sum to 1' constraint lives only in the schema. The description does not clarify units (tokens vs. thousands), the valid form of pinnedModelId, or the meaning of workDaysPerMonth, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Estimate') and a concrete comparison ('Cursor Auto blended cost versus pinning a single model'), which is more informative than the title alone. It does not, however, distinguish itself from the adjacent cost tools (agent_cost_estimate, swarm_run_cost_estimate) that occupy the same sibling namespace.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no statement of prerequisites, and no routing away from the two sibling cost estimators. An agent has to guess which of the three cost tools is appropriate for a given request.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_availabilityGet open call times with MattARead-onlyIdempotentInspect
List open slots for a free 30-minute intro call with Matt Turley (Continuum), read live from his Cal.com calendar. Use it when the user wants to talk to Matt, get help with a project, or book a call, before calling book_call. Window: date_from to date_to (YYYY-MM-DD), at most 14 days; defaults to the next 7 days. Times are returned in the requested timezone (IANA name, default UTC). Read-only: it books nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| event | No | Which call. intro-30 = free 30-minute intro call. | intro-30 |
| date_to | No | Last day, YYYY-MM-DD, at most 14 days after date_from. Default date_from + 6 days. | |
| timezone | No | IANA time zone for the returned times, e.g. America/New_York. Default UTC. | |
| date_from | No | First day, YYYY-MM-DD. Default today. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, destructiveHint=false, idempotentHint), so 'Read-only: it books nothing' largely reinforces them. The description still adds useful behavioral context beyond annotations: the 14-day window cap, the 7-day default, and that times come back in the requested timezone. It does not describe pagination or result volume, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose and usage condition, with window/behavioral details following. Slightly redundant in stating read-only twice ('Read-only: it books nothing' against readOnlyHint=true) and does not state which fields drive selection before the trailing behavior notes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description does step in to note that times are returned in the requested timezone, which is what an agent needs to interpret results. For a zero-required-parameter, read-only listing tool this is near-complete, though return shape (slot objects, count) is unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all four parameters including the date format, defaults, and timezone. The description restates the window and timezone behavior but adds no syntax or constraints beyond what the schema provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (open slots for a 30-minute intro call with Matt Turley), and distinguishes the source (live Cal.com calendar). An agent can immediately separate this from the sibling book_call.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names when to use it ('when the user wants to talk to Matt, get help with a project, or book a call') and routes to the alternative by specifying 'before calling book_call'. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
leak_checkLeak Check (older name for ship_check)ARead-onlyIdempotentInspect
Older name for ship_check, kept for existing callers. Same scan, same input, same result. Prefer ship_check.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The live, publicly reachable URL of the deployed app, e.g. https://my-app.lovable.app. A bare domain gets https://. | |
| fresh | No | Skip the 24h cache and scan again (use after deploying a fix). |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | Yes | |
| score | Yes | 0-100. Each red finding costs 30 points, each yellow 10. |
| cached | Yes | |
| counts | Yes | |
| scan_id | Yes | Pass to request_human_review. Null if the result could not be stored. |
| summary | Yes | |
| verdict | Yes | |
| findings | Yes | |
| report_url | Yes | |
| scanned_at | Yes | |
| limitations | Yes | |
| human_review | Yes | |
| overall_status | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds useful equivalence context ('same scan, same input, same result'), but says nothing about auth needs, rate limits, or deprecation/removal timing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero waste, front-loaded with identity, then equivalence, then the routing directive. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a deprecated alias with full schema coverage and an output schema that carries return values, the description supplies everything an agent needs to select correctly. It could be marginally stronger by noting whether the alias will be removed, but that is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters (url, fresh) are fully documented in the schema, so the baseline of 3 applies. The description's 'same input' adds no syntax or format detail beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states exactly what this tool is – an older alias of ship_check with identical behavior – and thereby distinguishes it from its sibling without opening either schema. It does not directly describe the underlying operation (scanning a deployed URL for leaks), but 'Same scan, same input, same result' conveys it by reference to a named sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit routing directive: 'Prefer ship_check.' Combined with 'kept for existing callers,' the agent knows both when this tool is appropriate (legacy callers only) and when to choose the alternative instead. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_servicesContinuum services and pricesARead-onlyIdempotentInspect
List Continuum service offerings and published prices (the same list as uxcontinuum.com/pricing), plus how to reach Matt: get_availability and book_call to book a free 30-minute call, send_message to send him a note, request_estimate for a ballpark from published pricing.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and closed-world, so the safety profile is fully covered. The description adds only that the output mirrors the public pricing page (static, unpaginated, non-personalized), a modest amount of context beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first clause, and the remaining clauses are a compact routing block. The parenthetical URL reference and the multi-tool tail make it slightly dense, but no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing returns, and it does so at a high level (offerings plus published prices, identical to the pricing page). For a zero-parameter informational read tool this is adequate, though it says nothing about output format or how many services are returned.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema baseline is 4. The description correctly implies a no-input, whole-catalog read rather than hinting at filtering options that do not exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ('List Continuum service offerings and published prices') and anchors the content to a known artifact ('the same list as uxcontinuum.com/pricing'). An agent can immediately distinguish it from cost-estimate and booking siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the follow-on tools and their purposes (get_availability/book_call to book a free call, send_message for a note, request_estimate for a ballpark), which effectively routes the agent. It stops short of an explicit 'use this when you need pricing, not X' exclusion, so 4 rather than 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
request_estimateRequest a project estimateAInspect
Get a ballpark price for a software project from Continuum's published pricing, and send the request to Matt Turley, who replies personally with a real quote. Use it when the user asks what a build, fix, review, ongoing support or AI-search visibility work would cost. The ballpark is the matching published offer(s) and price range from uxcontinuum.com/pricing, not a quote. Show it to the user labeled that way. Then offer a call: get_availability and book_call.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The user's name. | |
| Yes | The user's own email, for Matt's quote. | ||
| budget | No | ||
| timeline | No | ||
| project_type | No | launch-review (check an app before launch), fix-existing-app (stabilize or rescue an app), new-build (MVP from scratch), ongoing-support (maintenance or a product team on retainer), ai-visibility (get recommended by ChatGPT and AI search), other. Inferred from the description if omitted. | |
| project_description | Yes | What the user wants built or fixed, 10 to 4000 characters. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial context beyond the annotations (readOnlyHint=false, openWorldHint=true, idempotentHint=false): the request is delivered to a specific person who replies personally, the returned ballpark is explicitly NOT a quote and must be shown to the user labeled as such, and the workflow continues by offering a call. That is real behavioral disclosure of a non-idempotent, human-routed write.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and the pricing source are front-loaded, and the multi-step flow (show ballpark → offer call) is ordered logically. It is dense but each sentence carries information; only the repeated framing of 'ballpark vs. quote' is slightly redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description usefully explains the return semantics (matching published offer(s) plus price range, labeled as a ballpark) and the post-call handoff. The two undocumented optional parameters (budget, timeline) and the absence of any confirmation/permission step before sending user contact details are the remaining gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67%: name, email, project_type and project_description are documented, while budget and timeline are not. The description echoes the project_type categories in prose ('build, fix, review, ongoing support or AI-search visibility'), adding slight value, but contributes nothing about budget/timeline. Baseline 3 for mid coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource: it fetches a ballpark price from Continuum's published pricing AND sends a request to a named human (Matt Turley). This is clearly separable from the sibling cost-estimator tools (agent_cost_estimate, cursor_auto_cost_estimate, swarm_run_cost_estimate), which estimate compute/agent costs rather than project pricing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger — 'Use it when the user asks what a build, fix, review, ongoing support or AI-search visibility work would cost' — and names the follow-on tools (get_availability, book_call). It does not state a when-not condition or contrast against the similarly-named cost-estimate siblings, so it stops short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
request_human_reviewRequest a human review ($299)AInspect
Get a senior engineer (Matt Turley, 20 years shipping software) to review by hand the app behind a ship_check scan: the paid $299 Ship Check. Returns a Stripe Checkout link and a call booking link for the user to open themselves. This call charges nothing and never pays on the user's behalf. Only call it when the user asks for a human review or agrees to one. Pass the user's own email so Matt can follow up personally; no automated email is sent to it. Show the user checkout_url and booking_url.
| Name | Required | Description | Default |
|---|---|---|---|
| Yes | The user's email, for the checkout and for Matt to reply to. | ||
| scan_id | Yes | The scan_id returned by ship_check. |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | Yes | |
| price | Yes | |
| scan_id | Yes | |
| report_url | Yes | |
| booking_url | Yes | Book a 30-minute call with Matt instead of, or before, paying. |
| checkout_url | Yes | Stripe Checkout for the human review. The user opens and pays it themselves. |
| what_happens_next | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnly=false, idempotent=false, openWorld=true), it discloses the concrete side effects and limits: it returns a Stripe Checkout link and booking link for the user to open, charges nothing itself, never pays on the user's behalf, and sends no automated email. That is exactly the context an agent needs before invoking a paid flow.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a bit long at five sentences, but each one carries distinct payload (who, what, what it returns, when to call, what to do with the result) and the key facts are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description needn't enumerate return fields, yet it still tells the agent to surface checkout_url and booking_url to the user, closing the loop on the payment-flow handoff.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: the email must be the user's own so Matt can follow up, and no automated email is sent to it — clarifying both intent and a non-obvious non-behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (hand-review of the app behind a ship_check scan) and names the exact artifact it operates on, plus the price. It is clearly distinguishable from ship_check (which produces the scan) and book_call (which books directly).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit gating condition: 'Only call it when the user asks for a human review or agrees to one.' That is a clear when-to-use rule and an implicit when-not for every other case, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
send_messageSend Matt a messageAInspect
Send a message to Matt Turley (Continuum) on the user's behalf, like the contact form on uxcontinuum.com. Use it when the user wants to ask Matt something or describe a project but not book a call yet. Only send what the user asked to send, with their real name and email. Matt reads it and replies personally by email; no automated reply is sent. For a time to talk, use get_availability and book_call instead.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The user's name. | |
| Yes | The user's own email, for Matt to reply to. | ||
| topic | No | Short subject, e.g. "rescue a Lovable app". | |
| company | No | Company or project name. | |
| message | Yes | The message, 10 to 4000 characters. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover safety (readOnly=false, destructive=false, idempotent=false, openWorld=true), and the description adds genuinely non-obvious behavior: Matt reads and replies personally, no automated reply is sent, and only user-authorized content with real name/email may be sent. It stops short of covering failure modes, retries, or whether duplicates result from re-sending (relevant given idempotentHint=false).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose, then when-to-use, then content constraints, then the routing alternative. Each sentence carries weight; the closing routing sentence slightly restates the 'not book a call yet' clause, but with specific tool names that add actionable precision.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the expected outcome (personal email reply, no automated response) and the identity/content constraints an agent needs to form a valid call. Nothing material is missing for a 5-parameter, fully documented send tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3; the description still adds meaning by constraining identity fields ('with their real name and email') beyond the schema's neutral 'The user's name'/'The user's own email'. It leaves topic and company semantics entirely to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Send a message to Matt Turley (Continuum) on the user's behalf') and anchors it with a concrete analogue ('like the contact form on uxcontinuum.com'). It is clearly distinguishable from the scheduling siblings (book_call, get_availability) without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('when the user wants to ask Matt something or describe a project but not book a call yet') and names the alternative path for the excluded case ('For a time to talk, use get_availability and book_call instead'). Both the selecting and the excluding conditions are spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ship_checkShip Check: scan a deployed appARead-onlyIdempotentInspect
Before shipping, or right after deploying, an app built with Lovable, Bolt, Cursor, v0, Replit, Claude Code or similar, run ship_check on the live URL. It scans the deployed app from the outside, the way a visitor or attacker sees it, and returns a launch-readiness report where every finding has a status and a concrete fix. Checks: (1) secret keys exposed in the HTML and up to 3 same-origin JS bundles (Stripe live secret and restricted keys, AWS access keys, OpenAI and Anthropic API keys, Supabase secret and service_role keys, hardcoded bearer tokens and passwords); (2) open database access: if the page ships a Supabase URL and public anon or publishable key, whether the common tables profiles and users return rows to that key without login (a missing Row Level Security policy; count only, no row data is read); (3) missing security headers (Strict-Transport-Security, Content-Security-Policy, X-Frame-Options or frame-ancestors, X-Content-Type-Options, Referrer-Policy); (4) HTTPS, reachability and server errors; (5) broken internal links (spot-check of up to 6); (6) server response time; (7) whether login/signup and pricing/checkout are linked from the homepage; (8) whether AI search crawlers get real server-rendered content or an empty JavaScript shell. Read-only: plain GET/HEAD requests, never logs in, never submits forms, never writes. Not a code audit: it cannot see source code or test every table or route. Results are cached for 24 hours per URL; pass fresh=true to re-scan after deploying a fix. Only scan apps the user owns or is authorized to test. If the user wants a senior engineer to review the app by hand ($299), call request_human_review with the scan_id from this result.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The live, publicly reachable URL of the deployed app, e.g. https://my-app.lovable.app. A bare domain gets https://. | |
| fresh | No | Skip the 24h cache and scan again (use after deploying a fix). |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | Yes | |
| score | Yes | 0-100. Each red finding costs 30 points, each yellow 10. |
| cached | Yes | |
| counts | Yes | |
| scan_id | Yes | Pass to request_human_review. Null if the result could not be stored. |
| summary | Yes | |
| verdict | Yes | |
| findings | Yes | |
| report_url | Yes | |
| scanned_at | Yes | |
| limitations | Yes | |
| human_review | Yes | |
| overall_status | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered. The description still adds real behavioral context beyond them: plain GET/HEAD only, never logs in, never submits forms, 24-hour per-URL caching with fresh=true to bypass, and the explicit boundary that it cannot see source code or every route/table.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The opening sentence front-loads the trigger and action, and the numbered check list is dense but each item is concrete rather than filler. It runs long, but the length is justified by the breadth of what the scan returns and the scope limits; little could be cut without losing routing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return format need not be explained, and the description still covers authorization requirements, the read-only boundary, caching/freshness behavior, coverage limits, and the sibling escalation path. An agent has everything needed to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, and both parameters are already documented in the schema. The description still adds value by explaining the caching mechanic (24 hours per URL) and tying fresh=true to the post-deploy re-scan workflow, which clarifies when the flag is actually warranted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb (scans the deployed app from the outside) and resource (live URL) and enumerates the eight check categories, so the agent knows exactly what comes back. It also differentiates itself from siblings by naming request_human_review as the manual alternative and framing itself as launch-readiness rather than a code audit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('before shipping, or right after deploying'), when-not ('not a code audit: it cannot see source code or test every table or route'), prerequisites ('only scan apps the user owns or is authorized to test'), and a named escalation path to request_human_review for hand review. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
swarm_run_cost_estimateMulti-agent swarm run costCRead-onlyIdempotentInspect
Estimate the cost of a multi-agent orchestrator plus sub-agent swarm run.
| Name | Required | Description | Default |
|---|---|---|---|
| runsPerDay | Yes | ||
| retryRatePct | Yes | ||
| subAgentCount | Yes | ||
| subAgentModelId | Yes | ||
| orchestratorModelId | Yes | ||
| subAgentInputTokens | Yes | ||
| subAgentOutputTokens | Yes | ||
| orchestratorInputTokens | Yes | ||
| orchestratorOutputTokens | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds nothing beyond that — it does not say whether pricing is a local calculation or a live lookup, what currency/units are used, or how retries factor in.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It is efficient, though its brevity stems partly from under-specification rather than disciplined editing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter, all-required estimation tool with no schema descriptions and no output schema, the description is far too thin. Token-counting conventions, retry handling, aggregation across sub-agents, and the shape of the returned cost breakdown are all absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All 9 parameters are required with 0% schema description coverage, and the description compensates for none of them. Ambiguous semantics such as whether retryRatePct is 0-100 or 0-1, and what runsPerDay drives in the estimate, are left entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (estimate) and resource (cost of an orchestrator + sub-agent swarm run), which implicitly distinguishes it from the single-agent 'agent_cost_estimate' sibling. It is clear but never explicitly names or contrasts with the alternative cost-estimate tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to choose this over agent_cost_estimate or cursor_auto_cost_estimate, nor any preconditions. The agent must infer the multi-agent scope on its own.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
11 tool updates
- First observed
agent_cost_estimate - First observed
book_call - First observed
cursor_auto_cost_estimate - First observed
get_availability - First observed
leak_check - First observed
list_services - First observed
request_estimate - First observed
request_human_review - First observed
send_message - First observed
ship_check - First observed
swarm_run_cost_estimate
Related MCP Connectors
Check a live app you own for public databases, leaked keys and exposed files.
Check a deployed web app for security headers, SEO, accessibility and performance, with a grade.
Scan a website for vulnerabilities: OWASP Top 10, CVEs, SSL, headers - with plain-English fixes
Compliance & security scan for your app: secrets, exposed files, headers, privacy, AI-disclosure.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceScans public URLs for compliance and security issues such as leaked API keys, exposed files, missing privacy pages, and security headers, helping developers identify gaps before launch.MIT

Sekrd Security Scannerofficial
FlicenseAqualityDmaintenanceEnables deep security auditing of web applications directly from AI IDEs including Cursor and Claude Code. Scans URLs for vulnerabilities, returns security scores with SHIP/BLOCK verdicts, and provides specific fix prompts for remediation.3-- AlicenseNot gradedqualityFmaintenancePredeploy security scanner for AI-generated code. 80+ vulnerability patterns across secrets, auth, injection, config, Supabase, and logging. Runs locally, code never leaves your machine. Optional x402 witnessed attestation.62 npmApache 2.0
- FlicenseAqualityCmaintenanceAudits PostgreSQL/Supabase schemas for security issues like missing RLS, permissive policies, and sensitive data exposure during AI conversations.3-
Glama MCP Gateway
Add one secure layer between your agents and this server.