web-data-toolkit
Server Details
YouTube transcripts (video, channel, search), Google Trends, Google Play and App Store reviews.
- Status
- Healthy
- Uptime
- 95.0% over 22 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-06-18
- URL
- Repository
- leekung125/web-data-toolkit-mcp
- GitHub Stars
- 0
TDQS
Scored across 5 tools
Each tool targets a distinct data source or granularity: iOS reviews, Google Play reviews, Google Trends, channel-level transcripts, and video-level transcripts. The only mild overlap is between the two YouTube transcript tools, but their input modes (channel/search vs. specific video URLs) are clear enough to avoid serious misselection.
All tool names follow the same lowercase snake_case pattern of platform_entity: app_store_reviews, google_play_reviews, google_trends, youtube_channel_transcripts, youtube_transcript. The pattern is predictable and readable, with only a minor singular/plural difference between the two YouTube tools.
Five tools is a well-scoped size for a web data toolkit. Each tool covers a meaningful data source without redundancy or bloat, and the count is neither too thin nor too heavy.
The toolkit covers its apparent domain well: both major app store review sources, search trends, and YouTube transcripts at both channel and video level. Minor gaps exist, such as no pagination for channel transcripts and no explicit metadata-only fetch, but agents can generally accomplish core workflows without dead ends.
Available Tools
5 toolsapp_store_reviewsApp Store customer reviews for an iOS app, by country and sort order.BInspect
Rating, title, text, author, app version and the review date for any iOS App Store app, plus the app's own metadata row (name, seller, price, average rating, rating count).
| Name | Required | Description | Default |
|---|---|---|---|
| app | Yes | App Store URL or numeric app id such as 389801252. | |
| max | No | Maximum reviews per app, 1-500. Default 100. | |
| raw | No | Return every field as JSON instead of a compact summary. Default false. | |
| sort | No | mostRecent (default) or mostHelpful. | |
| country | No | Two-letter country code. Default us. | |
| maxRating | No | 1-5. Keep only reviews at or below this rating. | |
| minRating | No | 1-5. Keep only reviews at or above this rating. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure, but it only lists return fields and never states that this is a read-only operation, what happens with invalid app IDs, whether auth is needed, or any rate limits. The output-focused text gives no insight into side effects or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence with no filler, front-loading the key review fields before the metadata row. It is efficient, though it is a dense noun-phrase list rather than a verb-led, easily scannable sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The schema fully documents the parameters, and the description compensates for the missing output schema by naming the returned fields. However, it does not describe the result structure (e.g., reviews array versus metadata object), pagination, or any operational behavior, leaving some ambiguity for a tool with no annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All seven parameters are fully described in the input schema with types and meanings, so schema coverage is 100%. The description adds no input parameter semantics beyond what the schema already provides; it only clarifies the output contents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly enumerates the resource and output fields — review rating, title, text, author, app version, review date, plus the app's metadata row — making it evident what the tool returns. However, it lacks an action verb such as 'fetch' or 'list,' so the operation itself is implied by the tool name and title rather than explicitly stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for any iOS App Store app' provides some usage context and implicitly differentiates this tool from the sibling google_play_reviews. It does not explicitly describe when to use this tool versus alternatives or mention exclusions, so the guidance remains inferred rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
google_play_reviewsGoogle Play reviews for an Android app, by country and language.AInspect
Rating, text, author, thumbs-up count, app version and the developer's reply for any Google Play app.
| Name | Required | Description | Default |
|---|---|---|---|
| app | Yes | Google Play URL or package name such as com.whatsapp. | |
| max | No | Maximum reviews per app. Default 100. | |
| raw | No | Return every field as JSON instead of a compact summary. Default false. | |
| lang | No | Language code. Default en. | |
| rating | No | 1-5 to keep only that star rating. | |
| country | No | Two-letter country code. Default us. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It lists output fields but omits any mention of read-only nature, authentication, rate limits, pagination, or behavior on invalid inputs. The description does not contradict annotations (none exist) but fails to add meaningful behavioral context beyond the data fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the key output fields. There is no wasted text or redundancy. It is appropriately sized for a straightforward retrieval tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With six parameters and no output schema, the description should convey the output format (a list of reviews), sorting behavior, and how filters like country and language affect results. It only lists fields and does not mention that results are a list or that it honors the max and raw parameters beyond what the schema states. The description is adequate but leaves some operational details unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each parameter (app, max, raw, lang, rating, country) is already documented with defaults and meanings. The description adds no additional parameter semantics; it only enumerates return fields, which is useful but not param-specific. The baseline of 3 is appropriate given the schema's thoroughness.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource (Google Play app reviews) and the specific data fields returned (rating, text, author, thumbs-up count, app version, developer's reply). The title reinforces the scope (Android apps, by country and language). It is easily distinguished from the sibling app_store_reviews, which targets Apple's App Store.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for Google Play apps but does not explicitly state when to use this tool versus alternatives like app_store_reviews. There is no 'use this when' or 'use that when' guidance, leaving the agent to infer from the name and title. The context is clear but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
google_trendsGoogle Trends interest over time, interest by region and related queries.AInspect
For up to five keywords: relative search interest over a timeframe, interest broken down by region, and the top and rising related queries. No browser, straight from the Trends endpoints.
| Name | Required | Description | Default |
|---|---|---|---|
| geo | No | Country or region code such as US or GB. Empty means worldwide. | |
| raw | No | Return every field as JSON instead of a compact summary. Default false. | |
| keyword | Yes | Search term. Comma-separate up to 5 to compare. | |
| related | No | Include top and rising related queries. Default true. | |
| byRegion | No | Include the interest-by-region breakdown. Default true. | |
| timeframe | No | Any Trends range, e.g. now 7-d, today 3-m, today 5-y. Default today 12-m. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It adds useful context beyond the schema: 'No browser, straight from the Trends endpoints' signals an API-backed path with no browser automation, and the five-keyword limit is behaviorally relevant. However, it does not disclose authorization requirements, rate limits, or the meaning of 'relative' interest, leaving several behavioral aspects implicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no wasted words. The first sentence front-loads the three core outputs and the keyword limit; the second adds a distinguishing implementation trait. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-oriented tool with six parameters and no output schema, the description covers the main return categories and the key constraint (up to five keywords). It does not explain the response shape or potential error cases, but given the simple domain and the schema's high coverage, the remaining gaps are minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds some context by mapping its output list to the parameters (region breakdown, related queries), and it states the five-keyword limit that mirrors the keyword parameter. It does not add format details for geo or timeframe beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly enumerates the tool's three output types — relative search interest over time, interest by region, and related queries — tied to a precise resource (Google Trends). It is clearly distinguishable from the sibling tools, which concern app reviews and YouTube transcripts, and the 'For up to five keywords' scope sharpens the purpose further.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case: when Google Trends data for up to five keywords is needed. There is no explicit when-to-use versus alternatives, but because all siblings are from different domains, the lack of named exclusions is not a major gap. It does not state any constraints like minimum keywords or when to use related versus byRegion toggles.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
youtube_channel_transcriptsSearch YouTube, or list a channel or playlist, and return every video with its transcript.AInspect
Enumerates the newest videos of a channel, @handle or playlist and returns each video's metadata and transcript in a single call. Pass "search: your query" instead of a channel to go straight from a question to transcripts without knowing which creator to ask - the top results come back with full transcripts. Useful for building a corpus on a topic or from one creator.
| Name | Required | Description | Default |
|---|---|---|---|
| max | No | How many of the newest videos to fetch. Default 25, max 500. | |
| raw | No | Return every field as JSON instead of a compact summary. Default false. | |
| lang | No | Preferred caption language. Default en. | |
| source | Yes | @handle, channel URL or id, a playlist URL, or "search: your query" to search YouTube. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It does disclose key traits like returning the newest videos, returning metadata plus transcripts in one call, and search returning top results. However, it omits operational details such as rate limits, authentication requirements, behavior when transcripts are unavailable, or error handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise at three sentences and front-loads the core capability before explaining the search mode and use case. The final sentence about corpus building is useful but slightly optional; otherwise there is no wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema and no annotations, so the description must be self-sufficient. It covers the main source modes and the batch nature, but it does not mention the single-video alternative youtube_transcript, transcript availability caveats, return format details beyond 'metadata and transcript', or potential limits/errors. This leaves important edge cases to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds meaningful semantics beyond the schema by explaining the 'search: your query' source syntax and when to use it. Other parameters (max, raw, lang) are adequately documented in the schema itself.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action and resource: 'Enumerates the newest videos of a channel, @handle or playlist and returns each video's metadata and transcript in a single call.' It also explains the 'search:' mode, clearly distinguishing this bulk transcript tool from the single-video sibling youtube_transcript.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives practical guidance: use search when you don't know the creator, and use the tool for building a corpus from one topic or creator. It does not explicitly tell the agent when not to use it or point to youtube_transcript for single-video needs, so it falls just short of full exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
youtube_transcriptGet the full transcript of one or more YouTube videos as clean text.AInspect
Returns the transcript text for each video plus its language, whether the captions were auto-generated, word count, title and channel. Comma-separate up to 50 video URLs or ids. Videos with no captions come back with a status instead of text.
| Name | Required | Description | Default |
|---|---|---|---|
| raw | No | Return every field as JSON instead of a compact summary. Default false. | |
| url | Yes | YouTube URL or 11-character video id. Comma-separate up to 50. | |
| lang | No | Preferred language codes in order, comma separated. Default en. | |
| segments | No | Include timestamped segments as well as full text. | |
| translateTo | No | Translate the transcript into this language code. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden. It discloses output composition, the auto-generated captions indicator, and the edge case where no captions exist. It does not cover error conditions or output formatting details, but for a read-only transcript tool this is strong disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, no filler, and every clause carries useful information. The output contents are listed first, followed by input format and an edge case. This is appropriately sized and well structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a public read operation with no output schema, the description covers the main return fields, batch input rules, and a notable failure behavior. It could mention the exact status values or error handling, but the description is sufficient for an agent to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters. The description reinforces the URL batch limit and adds context about missing captions, but does not add new meaning for raw, lang, segments, or translateTo.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb and resource: it returns the transcript text for each video, plus language, auto-generation status, word count, title, and channel. This clearly distinguishes it from channel-level or review tools like youtube_channel_transcripts and app_store_reviews.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear invocation context: pass up to 50 comma-separated video URLs or IDs, and videos without captions return a status. It does not explicitly name alternatives or exclusion conditions, but the scope is specific enough that an agent can infer when to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
- Added
app_store_reviews - Changed
google_play_reviews2 fields changed- added
Input schema / properties / rawAdded value: +{ + "description": "Return every field as JSON instead of a compact summary. Default false.", + "type": "boolean" +} - removed
Input schema / rawRemoved value: -{ - "description": "Return every field as JSON instead of a compact summary. Default false.", - "type": "boolean" -}
- Changed
google_trends2 fields changed- added
Input schema / properties / rawAdded value: +{ + "description": "Return every field as JSON instead of a compact summary. Default false.", + "type": "boolean" +} - removed
Input schema / rawRemoved value: -{ - "description": "Return every field as JSON instead of a compact summary. Default false.", - "type": "boolean" -}
- Changed
youtube_channel_transcripts2 fields changed- added
Input schema / properties / rawAdded value: +{ + "description": "Return every field as JSON instead of a compact summary. Default false.", + "type": "boolean" +} - removed
Input schema / rawRemoved value: -{ - "description": "Return every field as JSON instead of a compact summary. Default false.", - "type": "boolean" -}
- Changed
youtube_transcript2 fields changed- added
Input schema / properties / rawAdded value: +{ + "description": "Return every field as JSON instead of a compact summary. Default false.", + "type": "boolean" +} - removed
Input schema / rawRemoved value: -{ - "description": "Return every field as JSON instead of a compact summary. Default false.", - "type": "boolean" -}
4 tool updates
- First observed
google_play_reviews - First observed
google_trends - First observed
youtube_channel_transcripts - First observed
youtube_transcript
Related MCP Connectors
YouTube transcripts, Google Hotels prices, Google Ads Transparency, Google Trends and Threads posts.
YouTube transcripts, open jobs from careers pages, app store reviews and X posts for AI agents.
Web data tools: Threads, Yelp, YouTube/TikTok transcripts, Google Trends, Airbnb, Jumia prices.
YouTube data for AI agents: channels, videos, transcripts, comments, search. Video research.
Related MCP Servers
- FlicenseNot gradedqualityBmaintenanceEnables searching and retrieving video transcripts, metadata, channel info, playlists, comments, trending videos, engagement analytics, chapters, SponsorBlock clean transcripts, and most-replayed heatmaps.22 npm-
- AlicenseAqualityAmaintenanceMCP server for extracting structured intelligence from YouTube channels and videos — transcripts, topics, and competitive signals for AI-powered research workflows.1242 npmMIT
- FlicenseBqualityNot gradedmaintenanceEnables extraction and processing of YouTube video transcripts from individual videos, channels, and playlists. Supports transcript search, batch processing, multiple output formats (JSON, text, SRT, VTT), and bulk operations across multiple videos.1134 npm-
- AlicenseAqualityAmaintenanceEnables AI agents to search YouTube like a database, retrieve transcripts and channel data, and mine videos for claims, numbers, and demand signals — without needing an API key.650 npmMIT
Glama MCP Gateway
Add one secure layer between your agents and this server.