Skip to main content
Glama

Musevate MCP server

Generate video from a prompt, from an image, or from a set of reference images, through one endpoint that routes across the field of AI video models.

Endpoint: https://musevate.com/api/mcp (streamable-http) Bridge: musevate-mcp in this repository — a local stdio server, for clients that cannot speak streamable HTTP Registry: com.musevate/mcp on the official MCP registry Website: musevate.com

The server is hosted. There are two ways to reach it, and the first is better whenever your client supports it.

Connecting directly

Add https://musevate.com/api/mcp as a custom connector. Every method requires an account, so the first request is refused with 401 and a WWW-Authenticate header pointing at the OAuth metadata -- your client reads that, registers itself, and sends you to musevate.com to sign in.

Clients that offer a choice should pick "always required". Clients that prefer a static credential can send an API key from Settings instead:

Authorization: Bearer mv_live_...

Both routes reach the same tools. Neither is preferred.

By client

Claude (web, desktop or mobile) — Settings, Connectors, Add custom connector, and paste the endpoint. The sign-in happens in the browser.

Claude Code

claude mcp add --transport http musevate https://musevate.com/api/mcp

ChatGPT — Settings, Connectors, add a custom connector with the same URL.

Anything else that speaks streamable HTTP — point it at the endpoint. If the client has no OAuth support, give it the header instead:

{
  "mcpServers": {
    "musevate": {
      "type": "http",
      "url": "https://musevate.com/api/mcp",
      "headers": { "Authorization": "Bearer mv_live_..." }
    }
  }
}

Related MCP server: MCP Veo 3 Video Generation Server

Connecting through the local bridge

Some clients only know how to launch a command and talk to it over a pipe. For those, musevate-mcp is a stdio server that forwards to the same endpoint:

{
  "mcpServers": {
    "musevate": {
      "command": "npx",
      "args": ["-y", "github:musevate/MCP"],
      "env": { "MUSEVATE_API_KEY": "mv_live_..." }
    }
  }
}

It installs and builds itself from this repository. An npm release under the name musevate-mcp is coming; until it lands, the line above is the one that works, and it will keep working afterwards.

Create the key at musevate.com/settings. The bridge cannot run the OAuth flow -- that needs a browser, which a pipe does not have -- so it takes a key and nothing else.

It holds no logic of its own: no pricing, no retry rules, no opinion about what the tools do. Adding any would create a second place where "what does this cost" is decided, and two such places drift apart. It moves bytes and gets the framing right, which is most of what a bridge gets wrong:

  • A notification is answered with silence, never with a response to a request the client never made.

  • Every answer is re-serialised to exactly one line, so a pretty-printed body is not read as several malformed messages.

  • A gateway error page, an empty body or a dead network becomes a JSON-RPC error carrying the id the client is waiting on -- never raw HTML on the protocol channel, and never a hang.

Environment:

Variable

Default

MUSEVATE_API_KEY

—

required; --api-key also works

MUSEVATE_MCP_URL

https://musevate.com/api/mcp

for testing against another deployment

MUSEVATE_TIMEOUT_MS

180000

above the server's own 120s ceiling

Checking it by hand

The transport is plain JSON-RPC over POST; there is no session to establish and no stream to hold open.

curl -s https://musevate.com/api/mcp \
  -H 'content-type: application/json' \
  -H 'accept: application/json, text/event-stream' \
  -H "authorization: Bearer $MUSEVATE_API_KEY" \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'

Without the header that returns 401 and the WWW-Authenticate pointing at /.well-known/oauth-protected-resource, which is exactly what a connector follows to start the sign-in.

examples/generate.mjs runs the whole flow — balance, quote, generate, poll — in about a hundred lines and no dependencies:

export MUSEVATE_API_KEY=mv_live_...
node examples/generate.mjs "a slow dolly through a neon-lit street at night"
node examples/generate.mjs "the logo turns to face us" https://example.com/logo.png

Tools

Tool

Cost

What it does

list_models

free

Every model available, with maximum length, the sizes it offers, whether it can generate sound, and what it costs in credits

quote_video

free

Prices a generation without running it, and names the model that would serve it

generate_video

spends credits

Returns immediately with an id; rendering takes about a minute

check_video

free

Status, and once finished a link valid for one hour

check_balance

free

Credits remaining

Images

generate_video and quote_video take images as https URLs, because a tool call is a JSON object and cannot carry bytes:

  • imageUrl -- the picture to animate. Supplying one makes the request image-to-video: the image is the first frame and the prompt describes what happens to it.

  • endImageUrl -- the frame to finish on. Needs imageUrl. Only some models publish a field for one.

  • referenceImageUrls -- up to thirty, in the order the prompt addresses them as @Image1, @Image2. This is reference-to-video and is not combined with imageUrl.

The mode is derived from what you send rather than set separately, so it cannot disagree with the images.

Prefer an image over describing one. A model asked to draw a specific logo draws something close to it and no closer; given the logo, it carries it.

Cost, and not spending it twice

quote_video is free and is the right way to answer "what would this cost". generate_video is the one that charges, and says so in its own description so that an assistant reads it before calling.

An identical call repeated within two minutes returns the generation the first one started rather than starting a second. To make a deliberate second take, change the request or pass a new requestId.

Pricing

Credits, prepaid. No subscription required. Every generation is priced before it runs and refunded automatically if it fails for a technical reason.

Details at musevate.com/pricing.

Building the bridge

npm install
npm test      # builds, then runs the suite

src/bridge.ts is the whole translation, as one function over strings, so the cases that are awkward to reach through a real pipe are reachable from a test. src/index.ts holds only what needs a process: the pipe, the environment, the exit.

What this repository is, and is not

This is the local bridge, the manifest published to the registry, and the documentation for using the hosted server.

The product source is not here, and the bridge is not a reimplementation of it: routing, pricing, the margin floor and the refund rules all live in the service.

The MIT licence covers these files -- the bridge, the documentation and the manifest. It is not a licence to the Musevate service, which is a paid product governed by its own terms.

Available Tools

5 tools
check_balanceA

How many credits the account has left. Free to call.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Free to call' adds useful cost/access context, and 'how many credits the account has left' implies a read-only query returning a balance. Still, it does not disclose the exact response format or explicitly confirm there are no side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two very short sentences with zero filler. The core message is front-loaded and every word contributes to the agent's understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a trivial zero-parameter tool, the description gives the key fact an agent needs: the tool reports how many credits remain. 'Free to call' adds cost context. A slightly more explicit return format would be nice, but the description is sufficient at this complexity level.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so the baseline of 4 applies. The description does not need to explain parameter meaning because there is nothing to explain.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies a credit/balance-checking operation and is easily distinguished from sibling tools like generate_video or check_video. It lacks an explicit imperative verb like 'Get' or 'Inspect', but the meaning is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

For a zero-parameter standalone balance check there is no obvious alternative tool, and 'Free to call' signals that cost is not a concern. However, the description does not explicitly state when to use it or whether any authentication is required, leaving some guidance implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

check_videoA

Check a generation started by generate_video. Returns its status and, once finished, a link to the video that is valid for one hour. Free to call.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe id generate_video returned.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses cost behavior ('Free to call'), link expiry ('valid for one hour'), and the fact that it returns status before completion. Minor omissions like status values or explicit polling guidance are acceptable for a simple check tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences front-load the purpose and then state the return behavior and cost. Every phrase earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter status check, the description provides the essential context: what it checks, what it returns, link validity, and cost. A small gap is the lack of detail on status values or polling behavior, but nothing blocks correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single required parameter is fully described in the schema ('The id generate_video returned'), and the description reinforces that origin. With 100% schema coverage, the description adds no significant new parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the specific operation ('Check a generation'), the resource it acts on ('started by generate_video'), and the key return value (status and link). This clearly distinguishes it from siblings like generate_video or list_models.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It ties the tool explicitly to a prior generate_video call, giving clear when-to-use context. It does not list exclusions or alternatives, but the intended workflow is unambiguous from the description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_videoA

Generate a video from a prompt, from an image, or from a set of reference images. THIS SPENDS THE ACCOUNT'S CREDITS — call quote_video first and confirm the cost with the person you are working for before calling this. Returns immediately with an id; rendering takes about a minute. Use check_video to collect the result. Retrying is safe: an identical call repeated within two minutes returns the generation the first call started rather than starting a second one, so to deliberately make a second take either change the request or pass a new requestId.

ParametersJSON Schema
NameRequiredDescriptionDefault
audioNo
promptYesWhat the video should show.
qualityNo
imageUrlNohttps URL of an image to animate. Supplying one makes this image-to-video: the picture is the first frame and the prompt describes what happens to it. JPEG, PNG or WebP, up to 10 MB. Prefer this over describing a logo, a product or a face in words — a model asked to draw a specific logo will draw something close to it and no closer.
modelSlugNo
requestIdNoYour own id for this attempt. Send the same one when retrying and you get the same generation back instead of a second charge. Send a new one to deliberately start another.
resolutionNo
aspectRatioNo
endImageUrlNohttps URL of the frame to finish on. Needs imageUrl. Only some models publish a field for one; list_models with mode image_to_video shows which.
durationSecondsNo
referenceImageUrlsNohttps URLs of reference images, in the order the prompt addresses them as @Image1, @Image2 and so on. This is reference-to-video and is not combined with imageUrl. Order is preserved exactly as sent.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing side effects and behavior. It clearly states that the operation spends the account's credits, returns immediately with an id, and takes about a minute to render. It also discloses idempotent retry behavior (identical calls within two minutes return the same generation). It does not mention potential failure modes or authorization requirements, but the core behavioral profile is covered thoroughly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, dense paragraph but is well-organized. It front-loads the core purpose and the critical cost warning, then follows with the immediate-return behavior, the companion tool for collection, and the idempotency rule. Every sentence carries needed information, and there is no fluff. It is appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter tool with no output schema and no annotations, the description covers the essential operational context: how to initiate, what to expect in terms of timing, how to retrieve results, and how to avoid extra charges. It explains the variants (prompt/image/reference) and the special relationship between imageUrl and endImageUrl with respect to model support. It omits details on some parameters, but the missing ones are largely self-explanatory from their names, so the overall picture is sufficient for an agent to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 45%, so the description must compensate. It does explain key parameters: imageUrl (how it becomes the first frame, size limits, when to prefer it), referenceImageUrls (ordering with @Image1, @Image2, not combined with imageUrl), and requestId (idempotency usage). However, it offers no explanation for audio, quality, modelSlug, resolution, aspectRatio, or durationSeconds, which are left to their names alone. The partial coverage earns a 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Generate a video from a prompt, from an image, or from a set of reference images,' which states a specific verb, the resource (video), and the three distinct input modes. This clearly distinguishes it from siblings like check_video and quote_video by naming the exact action. It leaves no ambiguity about what 'generate' means here.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly instructs to call quote_video first and confirm the cost with a human, names check_video as the way to collect the result, and provides guidance on when to prefer imageUrl over a text prompt (for logos, products, faces). It also explains how to force a second take by changing the request or passing a new requestId. This is textbook when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_modelsA

List the AI video models available, with what each can actually do: maximum length, the sizes it offers, whether it can generate sound, and what it costs in credits for a five second clip at 720p. Use this before quoting or generating so the settings you pick are ones a model can serve. This server generates in every mode listed except lip_sync: pass imageUrl for image-to-video, or referenceImageUrls for reference-to-video.

ParametersJSON Schema
NameRequiredDescriptionDefault
modeNoOnly list models that can serve this mode. Defaults to text_to_video.
needsAudioNoOnly list models that can generate sound with the video.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it delivers: it discloses what the returned capabilities are, the credit cost basis, and a server behavior detail about modes and image inputs. The 'every mode listed except lip_sync' phrase is slightly ambiguous and could be clearer, but it is a substantive behavioral disclosure beyond what the schema states.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler; the core purpose is front-loaded and the usage guidance follows immediately. The only weak point is the slightly awkward 'except lip_sync' clause, but overall the description is compact and readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description compensates by spelling out exactly what the listing returns: maximum length, sizes, sound support, and credit cost. For a low-complexity list tool, this is sufficient; error behavior and further edge cases are not necessary for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the mode and needsAudio parameters are already well documented. The description adds useful context about cost basis and image input behavior, but those details are not tightly tied to the two parameters of this specific tool, so the added semantic value is moderate rather than transformative.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') with a clear resource and object ('the AI video models available') and enumerates concrete capabilities (max length, sizes, sound, credit cost). It also orients the agent relative to its siblings ('before quoting or generating'), making it easy to tell apart from quote_video and generate_video.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit when-to-use instruction: 'Use this before quoting or generating so the settings you pick are ones a model can serve.' It does not explicitly name alternative tools or state when not to use it, but the timing guidance is clear enough for an agent to route correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

quote_videoA

Price a generation without running it. Returns the credit cost and the model that would be used. Free, and the right way to answer 'what would this cost' — generate_video is the one that spends credits. Images are not downloaded for a quote, so the price reflects the kind of request rather than the pictures themselves.

ParametersJSON Schema
NameRequiredDescriptionDefault
audioNoGenerate sound with the video.
promptYesWhat the video should show.
qualityNoWhich tier the router should choose from. Default quality.
imageUrlNohttps URL of an image to animate. Supplying one makes this image-to-video: the picture is the first frame and the prompt describes what happens to it. JPEG, PNG or WebP, up to 10 MB. Prefer this over describing a logo, a product or a face in words — a model asked to draw a specific logo will draw something close to it and no closer.
modelSlugNoPin a specific model instead of letting the router choose.
resolutionNo720p, 1080p, 1440p or 4k. Default 720p.
aspectRatioNo
endImageUrlNohttps URL of the frame to finish on. Needs imageUrl. Only some models publish a field for one; list_models with mode image_to_video shows which.
durationSecondsNoClip length in seconds. Default 5.
referenceImageUrlsNohttps URLs of reference images, in the order the prompt addresses them as @Image1, @Image2 and so on. This is reference-to-video and is not combined with imageUrl. Order is preserved exactly as sent.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations at all, the description carries the burden of behavioral disclosure. It clearly states the operation is free, performs no generation, and does not download images. This gives an agent a solid safety model: quoting is non-destructive and non-spending. It could add rate-limit or auth context, but what's here is meaningful and accurate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, no filler: the first states the core action, the second gives selection guidance, the third explains a subtle behavioral detail. Every sentence earns its place, and the key purpose is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 10 parameters and no output schema, the description gives enough high-level context to select and invoke it correctly: it returns cost and model, it is free, and it does not generate or download. The parameter details are handled by the schema. A slightly deeper note about how quoting interacts with modelSlug or imageUrl would push it higher, but no essential fact for usage is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 90%, so the schema already documents the parameters thoroughly. The description adds a high-level framing ('credit cost and model') but does not go deeper into individual parameters. That is acceptable given the rich schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb-resource pair ('Price a generation') and immediately distinguishes this tool from the sibling that spends credits ('generate_video is the one that spends credits'). It also states what it returns—credit cost and the model that would be used—so an agent can tell exactly what this tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says this is 'the right way to answer what would this cost' and contrasts it with generate_video, which is the credit-spending alternative. The guidance is directly usable for tool selection, even without weighing all siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 5 tool updatesv0.1.0
    • First observedcheck_balance
    • First observedcheck_video
    • First observedgenerate_video
    • First observedlist_models
    • First observedquote_video

TDQS

A4.2/5.0

Scored across 5 tools

Disambiguation5/5

Each tool targets a distinct action: listing model capabilities, checking balance, quoting a cost, generating a video, and checking generation status. There is no meaningful overlap or ambiguity between them.

Naming Consistency5/5

All tool names follow a consistent snake_case verb_noun pattern: list_models, check_video, check_balance, quote_video, generate_video. The naming is predictable and clear across the entire set.

Tool Count5/5

Five tools is well-scoped for a video generation service. Each tool covers a necessary step in the workflow—discovering options, checking balance, quoting, generating, and retrieving results—without unnecessary bloat.

Completeness5/5

The tool surface covers the full lifecycle needed for paid video generation: model discovery, pricing, balance checking, submission, and result retrieval. Generation is idempotent, and the workflow has no obvious dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers