Skip to main content
Glama
razvangirgiz

wazap-mcp

by razvangirgiz

Get the media of a WhatsApp message

get_media
Read-only

Retrieve and save media from a WhatsApp message, including audio transcripts, photos, or files, to a specified directory.

Instructions

What a message's media holds: a voice note or audio as its transcript (kept once made; an API provider bills it), or with save_to its file; a photo attached as an image; any file saved at path on the machine running wazap. MEDIA_UNAVAILABLE: WhatsApp no longer has it.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
save_toNoAbsolute directory; default <data-dir>/media. A recording: its file, and no new transcript
languageNoWhat is spoken, e.g. "ro"
account_idNoAccount id
message_idYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
mimeNo
pathNo
sizeNo
typeYes
senderYes
captionYes
filenameNoThe name it was saved under
account_idYes
message_idYes
transcriptNo
image_attachedNoThe photo, or a small preview of it, is attached
original_filenameYes
transcript_unavailableNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv1.0.3

TDQS

B3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds valuable behavior beyond the annotations: transcripts are 'kept once made' and 'an API provider bills it' (a cost/caching trait), save_to writes media to local disk, and MEDIA_UNAVAILABLE is described as a possible error when WhatsApp no longer retains the media. These details are not in the annotations and are context-rich. Nothing contradicts the readOnlyHint/destructiveHint flags.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short (about 55 words) and contains no fluff, but it is structured as a dense, semicolon-packed run-on that is hard to parse quickly. The MEDIA_UNAVAILABLE sentence is a useful standalone component, yet the overall ordering muddles the main action. It is concise without being cleanly structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema and annotations available, the description does not need to explain return values. However, it fails to clarify the role of account_id, the meaning of message_id, or when to use language vs save_to. It does cover the media type possibilities and the MEDIA_UNAVAILABLE error, but an agent would still have to infer significant context from parameter names and external knowledge.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, so the heavy lifting is done by the schema. The description adds a little context for save_to (its file) and the transcript/language relation, but the required message_id has no schema description and the description does not compensate. This is an acceptable baseline-3 case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explains what a message's media can hold (transcript, file, photo, saved path) but never states a clear verb phrase like 'fetches' or 'returns' – it is phrased as a noun fragment. The title 'Get the media of a WhatsApp message' is clearer than the description. It implies media retrieval but does not explicitly differentiate from siblings like get_message.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives such as get_message, read_messages, or search. No exclusions or 'instead of' notes are present; usage is only weakly implied by the resource being a message's media. With many siblings, this is a notable gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.