AetherWave Studio
Server Details
One tool surface for music, image, video, and audio generation across Suno, Grok Imagine, Seedance, Kling, Hailuo, Wan, VEO, Ideogram, and GPT Image 2. Generate, edit, upscale, reframe, and master through one credit pool. Connect in one click with OAuth, no API key required.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP
- URL
Glama MCP Gateway
Connect through Glama MCP Gateway for full control over tool access and complete visibility into every call.
Full call logging
Every tool call is logged with complete inputs and outputs, so you can debug issues and audit what your agents are doing.
Tool access control
Enable or disable individual tools per connector, so you decide what your agents can and cannot do.
Managed credentials
Glama handles OAuth flows, token storage, and automatic rotation, so credentials never expire on your clients.
Usage analytics
See which tools your agents call, how often, and when, so you can understand usage patterns and catch anomalies.
Tool Definition Quality
Average 4.6/5 across 16 of 16 tools scored.
Each tool targets a distinct media operation (image, video, audio, listing, mastering, etc.) with clear boundaries. Even similar tools like generate_image and edit_image are differentiated by their primary intent (creation vs. modification) and model selection guidance.
All tools follow the 'aetherwave_verb_noun' pattern consistently, using snake_case. Verbs and nouns are descriptive and predictable (e.g., generate_image, list_video_models, remove_background_video).
16 tools is slightly above the ideal range (3-15) but remains well-scoped for a multimedia generation platform covering image, video, audio, and user management. Each tool serves a distinct purpose, and no obvious bloat exists.
The tool surface covers core creation, editing, listing, and enhancement workflows for images, videos, and audio. Minor gaps exist (e.g., no delete tool, no get-single-creation tool), but the essential lifecycle is well-covered.
Available Tools
16 toolsaetherwave_balanceCheck credit balanceRead-onlyInspect
Returns the current AetherWave credit balance for the API key. Use this BEFORE a generation to confirm sufficient credits, especially for video which can cost 30-300+ credits depending on model/duration/resolution.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
aetherwave_edit_imageEdit image with AI (I2I)Inspect
Edits an existing image guided by a text prompt. Pass a public imageUrl plus a prompt describing the change ("add a moon to the sky", "swap the background for a neon city", "make it look like a comic panel"). Submits, polls, and returns the edited image URL(s). Default model is 'grok-imagine-i2i' (6 cr per call, returns 2 variations, ~30s, best cost-to-quality on standard edits). Other I2I-capable models: 'seedream-v4-edit', 'wan-2.5-spicy-i2i', 'flux-kontext-pro', 'qwen-image-edit', 'gpt-image-1.5-i2i' (slow, ~5min). Use list_image_models for full lineup. Note: source URLs with spaces or parentheses may fail upstream; prefer clean URLs.
Model selection guide for edits
Default: grok-imagine-i2i (6 cr per call, returns 2 variations = 3 cr/image effective, fast ~30s, strong general-purpose edit quality).
Pick a different model when:
Need a single deterministic output, or 4K resolution ->
seedream-v4-edit(7 cr per image, supports 1K/2K/4K, multi-image up to 6)Subtle edits / preserve composition / character consistency ->
flux-kontext-proorflux-kontext-maxNSFW edits ->
wan-2.5-spicy-i2iHighest quality, time is not a concern (~5 min OK) ->
gpt-image-1.5-i2iorgrok-imagine-quality-i2i(16 cr @ 1K, 22 cr @ 2K)Stylized / artistic transformation ->
midjourney-i2i
If the user simply says "edit this image" with no other signal, default to grok-imagine-i2i.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Model ID. Defaults to 'grok-imagine-i2i' (3 cr/image effective, 2 outputs). Other options: 'seedream-v4-edit', 'wan-2.5-spicy-i2i', 'flux-kontext-pro', 'qwen-image-edit', 'gpt-image-1.5-i2i', 'grok-imagine-quality-i2i'. Use list_image_models for the full list. | |
| prompt | Yes | Text description of the edit (e.g. 'replace the sky with sunset clouds'). | |
| quality | No | Quality preset for models that support it (e.g. GPT Image 2). | |
| imageUrl | Yes | Public URL of the source image to edit. Must be a real, fetchable URL. | |
| maxImages | No | Number of variations to return for multi-output models. | |
| resolution | No | Output resolution. Tiered-pricing models accept '1K' / '2K'. | |
| aspectRatio | No | Output aspect ratio (e.g. '1:1', '16:9'). Defaults to the source ratio for most models. | |
| renderingSpeed | No | Rendering speed preset for models that support it. | |
| negative_prompt | No | What to avoid in the output (supported by some models). |
aetherwave_generate_imageGenerate image (Grok Imagine, GPT Image 2, Seedream V4, Wan, Imagen 4, Nano Banana, Ideogram V3, Z-Image Turbo)Inspect
Generates one or more images from a text prompt (T2I) or a text prompt + reference image(s) (I2I). Submits the job, polls until terminal, and returns the final image URLs. Default model is 'grok-imagine-t2i' (fast, 6 images per generation, 5 credits). Use list_image_models to see the full lineup with pricing. For I2I, pass referenceImages as an array of public image URLs and pick a model with I2I support (e.g. 'grok-imagine-i2i', 'wan-2.5-spicy-i2i').
Model selection guide (when the user does not specify a model)
Default: grok-imagine-t2i (5 cr, 6 outputs per call, fast, general purpose).
Strong recommendation: when a single high-quality output is what's wanted (most agent / one-shot workflows), prefer gpt-image-2-t2i (9 cr @ 1K / higher @ 2K, single deterministic image, best general quality across realism, illustration, typography, and composition; supports up to 2K resolution and most aspect ratios including auto). This is the front-runner for serious creative output where you don't need to pick from 6 variations.
Pick a different model when the prompt has these signals:
"single best result" / "one image" / production / no time to pick from variations ->
gpt-image-2-t2i(9 cr, 1 output, top general quality)"photoreal" / "photo of" / "realistic" ->
gpt-image-2-t2i(9 cr, best general realism) orimagen-4(12 cr, very high quality) orz-image-turbo(3 cr, fastest)"highest quality" / "premium" / no budget ->
gpt-image-2-t2iat 2K, orgrok-imagine-quality-t2i(16 cr @ 1K, 22 cr @ 2K), orimagen-4-ultraText inside the image (signs, posters, typography) ->
ideogram-v3-t2i(best in class) orgpt-image-2-t2i(also strong)Artistic / painterly / stylized ->
midjourney-t2iAlbum art / cover art ->
gpt-image-2-t2ifor one strong image;grok-imagine-t2ifor 6 variations to choose from;seedream-v4-t2iif 4K wantedLogo or design with embedded text ->
ideogram-v3-t2iNSFW / adult / explicit ->
wan-2.5-spicy-t2i(auto-tags creation as 18+; routes to adult gallery)Cheapest possible / quick test ->
z-image-turbo(3 cr)Multiple variations to compare -> keep
grok-imagine-t2i(6 outputs default) or usenumImageson a multi-output model
For I2I (reference image provided): prefer the dedicated aetherwave_edit_image tool for "change something in this image" intent. Use aetherwave_generate_image with I2I models only when you specifically want style transfer (midjourney-i2i), premium quality (grok-imagine-quality-i2i), or adult content (wan-2.5-spicy-i2i).
Always pass an explicit aspectRatio (e.g. "1:1" for square album art, "16:9" for video thumbnails, "9:16" for shorts/reels). Some upstream providers reject submissions with no aspect ratio.
Ask the user only when:
The prompt contradicts itself (e.g., "highest quality but cheapest")
The user requested "the best model" with no context, surface 2-3 options with tradeoffs
A single generation would cost more than 20 credits and the user has not confirmed
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Seed for deterministic generation (supported by some models). | |
| model | No | Model ID. Defaults to 'grok-imagine-t2i'. Use list_image_models for the full list. | |
| prompt | Yes | Text description of the image to generate. | |
| numImages | No | Number of images for models that support multiple outputs. | |
| resolution | No | Output resolution. Most models accept '1K' or '2K'; some accept '480p'/'720p'. | |
| aspectRatio | No | Aspect ratio (e.g. '1:1', '16:9', '9:16'). Pass this explicitly when possible; some upstream providers reject submissions without an aspect ratio. Default ratios vary by model. | |
| negative_prompt | No | What to avoid in the output (supported by some models). | |
| referenceImages | No | Array of public image URLs for image-to-image generation. Required when using an I2I model. A single URL string is also accepted (wrapped as a one-element array). |
aetherwave_generate_musicGenerate music (Suno)Inspect
Generates AI music via Suno. Returns two tracks per submission. Default model is V5.5 (newest, best quality). For instrumental output set instrumental: true. Music gen typically takes 30-90s - this tool polls with up to a 6-minute budget. Note: the title param is advisory for instrumentals - Suno often writes its own title from the prompt content for instrumental generations. Transient GENERATE_AUDIO_FAILED errors are common; retry once before degrading the model version.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | Suno model version. Defaults to V5_5 (current best). | |
| title | No | Optional title for the generated tracks. | |
| lyrics | No | Custom lyrics. If omitted, Suno will generate lyrics from the prompt (unless instrumental=true). | |
| prompt | Yes | Style/mood/topic description. E.g. 'Lo-fi ambient track, rain sounds, warm pads' or 'High-energy synthwave with driving bass'. | |
| instrumental | No | If true, no vocals. Default false. |
aetherwave_generate_videoGenerate video (Grok Imagine, Wan 2.7, Hailuo 02, Seedance, Kling 2.6, VEO 3.1, Happy Horse)Inspect
Generates a short-form video from a text prompt (T2V) or a text prompt + starting image (I2V). Submits, polls, and returns the final video URL. Default model is 'grok-imagine-t2v' (fast, 4-6 cr/s, with built-in KIE -> fal.ai fallback). Use list_video_models for the full lineup with credit cost per second. I2V models (e.g. 'grok-imagine-i2v', 'seedance-pro-i2v') require a public imageUrl. Video generation can take 30s to several minutes; this tool polls with up to an 8-minute budget.
Model selection guide for videos (when the user does not specify a model)
Default: grok-imagine-t2v (4-6 cr/s, fast, has KIE -> fal.ai fallback for redundancy. Best general-purpose).
Pick a different model when the prompt has these signals:
"highest quality" / "premium" / broadcast / commercial ->
veo3.1-qualityorveo3-quality(Google's flagship, fixed 350-560 cr for 8s, 3-5 min)"fast premium" / quick high-quality ->
veo3-fastorveo3.1-fast(84 cr fixed for 8s)Cinematic camera moves / dolly / pan ->
seedance-pro-t2v(3-10 cr/s) orkling-3.0-pro-t2v(26 cr/s)Realistic human motion / faces ->
hailuo-2.3-pro-i2v(I2V, supply imageUrl)Talking head / lip sync ->
kling-avatar-pro(23 cr/s) orinfinitalk(5-17 cr/s)Anime / stylized / fantasy ->
wan-2.7-t2vNSFW / adult ->
wan-22-nsfw-i2v(I2V only; auto-tags adult)Animate this exact image -> any I2V variant (
grok-imagine-i2v,seedance-pro-i2v,hailuo-2.3-pro-i2v)First + last frame interpolation ->
seedance-pro-i2vwith bothimageUrl+endImageUrlCheapest test ->
hailuo-2.0-standard@ 512p (3 cr/s, ~18 cr for 6s) orgrok-imagine-t2v@ 480p (4 cr/s, ~24 cr for 6s)Clip 12-15s ->
grok-imagine-t2v(accepts up to 15s)True 4K ->
kling-3.0-4k-t2v(94 cr/s, expensive but native 4K)
Audio in generated video: grok-imagine-t2v, seedance-pro-t2v, and the VEO 3.x family include audio at base cost (no surcharge). Kling 2.6 and Kling 3.0 are the outliers — they price audio as a +50-100% surcharge (Kling 2.6 doubles the cost, Kling 3.0 Pro adds ~46%). Default to Grok / Seedance / VEO when sound matters and you don't want to think about audio pricing.
Cost framing: resolution and duration drive cost more than model choice. A 6-second 480p Grok generation costs ~24 cr; the same prompt at 1080p Seedance 2 is ~858 cr (35x more). Pick the lowest acceptable resolution + duration first.
For I2V models: imageUrl is required. For first+last-frame models, pass endImageUrl too.
Ask the user only when:
Single generation would cost more than 100 credits and they haven't confirmed
They asked for "the best" with no other signal; surface 2-3 options with cost ranges
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Moderation mode for Grok Imagine. Defaults to 'normal'. | |
| model | No | Model ID. Defaults to 'grok-imagine-t2v'. Use list_video_models for the full list. | |
| prompt | Yes | Text description of the video scene. | |
| duration | No | Duration in seconds. Grok Imagine accepts 6-15; other models have their own ranges (see list_video_models). | |
| imageUrl | No | Public URL of starting image. Required for I2V models. | |
| resolution | No | Output resolution. Default depends on model. | |
| aspectRatio | No | Aspect ratio (e.g. '16:9', '9:16', '1:1'). | |
| endImageUrl | No | Public URL of ending image. Supported by some I2V models (first+last frame). |
aetherwave_list_image_modelsList available image modelsRead-onlyInspect
Returns every image-generation model AetherWave supports, with its credit cost, default aspect ratio, supported inputs (T2I vs I2I), and any model-specific options. Call this before generate_image when you don't know the right model ID. The model key (e.g. 'grok-imagine-t2i') is what you pass as model to generate_image.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
aetherwave_list_master_presetsList available audio mastering presetsRead-onlyInspect
Returns every AI mastering preset AetherWave supports, with target LUFS, tags, descriptions, and difficulty level. Call this before master_audio when you don't know which preset fits the track. 12 presets total covering streaming, hip hop, EDM, pop, rock, lo-fi, R&B, acoustic, cinematic, podcast, gentle, and loud-and-punchy mastering styles. Each preset has a target LUFS value (e.g. -14 for streaming, -9 for loud) so you can match the user's distribution target.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
aetherwave_list_my_creationsList my AetherWave gallery itemsRead-onlyInspect
Returns items from the authenticated user's gallery — images, videos, audio tracks they've generated on AetherWave. Useful for agent workflows like 'find my last 5 images and reframe them all to 9:16' or 'list my recent songs and master each one'. Supports pagination and type filtering. Each item includes id, type, prompt, model, contentUrl, thumbnailUrl, createdAt, isFavorite, visibility, rating, and type-specific fields (duration for audio/video, width/height for images).
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | Filter to a single media type. Omit for all types. | |
| limit | No | Max items to return. Defaults to 100, max 500. | |
| offset | No | Pagination offset. Defaults to 0. | |
| favoritesOnly | No | If true, only return items marked as favorite. |
aetherwave_list_video_modelsList available video modelsRead-onlyInspect
Returns every video-generation model AetherWave supports (Grok Imagine, Wan 2.7, Hailuo 02, Seedance Pro/Lite, Kling 2.6 with audio, VEO 3.1, Happy Horse, etc.) with per-second credit cost, supported durations, resolutions, aspect ratios, and whether the model needs an input image (I2V). Call this before generate_video when you don't know the right model ID.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
aetherwave_master_audioMaster an audio track (AI mastering)Inspect
Submits an audio file for AI mastering and returns the mastered URL synchronously (route polls the Python service internally; expect 30s-5min). Useful as a final polish step after music generation. Cost: 20 credits per track. Producer, Mogul, and Ultimate plans get mastering free. Output is WAV (~50MB per 3-minute track, lossless for redistribution). Pick a preset to steer the mastering style; call aetherwave_list_master_presets for the full live list (12 presets including streaming, loud, gentle, hip_hop, edm, pop, rock, lofi, rnb, acoustic, cinematic, podcast). Each preset has a target LUFS value so you can match the distribution target.
| Name | Required | Description | Default |
|---|---|---|---|
| preset | Yes | Mastering preset name. Must be one of: 'streaming', 'loud', 'gentle', 'hip_hop', 'edm', 'pop', 'rock', 'lofi', 'rnb', 'acoustic', 'cinematic', 'podcast'. Call aetherwave_list_master_presets for full metadata (target LUFS, description, tags). | |
| audioUrl | Yes | Public URL to the source audio file (MP3 or WAV). | |
| trackTitle | No | Optional title for the mastered output (used in gallery row label). |
aetherwave_reframe_imageReframe image to a new aspect ratio (Ideogram V3 Reframe)Inspect
Reframes an image to a new aspect ratio by intelligently outpainting the edges. Pass a public imageUrl and the target aspectRatio ('16:9', '9:16', '1:1', '4:3', '3:4', etc.). Three speed tiers: 'turbo' (5 cr, fast), 'balanced' (10 cr, default), 'quality' (14 cr, slowest, best edges). Returns the reframed image URL.
| Name | Required | Description | Default |
|---|---|---|---|
| speed | No | Rendering speed. 'turbo'=5cr, 'balanced'=10cr (default), 'quality'=14cr. | |
| imageUrl | Yes | Public URL of the source image. | |
| aspectRatio | Yes | Target aspect ratio (e.g. '16:9', '9:16', '1:1', '4:3', '3:4', '21:9'). |
aetherwave_reframe_videoReframe video to a new aspect ratio (Luma Ray 2 Flash)Inspect
Reframes a video to a new aspect ratio by intelligently outpainting/cropping the edges. Pass a public videoUrl and target reframeAspectRatio. 17 credits per second. Optional reframePrompt lets you steer the new edge content (e.g. 'extend the sky with sunset clouds'). Returns the reframed video URL (R2-hosted).
| Name | Required | Description | Default |
|---|---|---|---|
| videoUrl | Yes | Public URL of the source video (MP4). | |
| reframePrompt | No | Optional prompt to steer the new edge content. | |
| reframeAspectRatio | Yes | Target aspect ratio. |
aetherwave_remove_backgroundRemove background from image (Recraft + fal.ai BiRefNet v2 fallback)Inspect
Strips the background from an image, returning a PNG with transparent alpha. Pass a public imageUrl. Useful for product shots, character cutouts, logo isolation, or compositing onto a new background. ~5 credits per image. Recraft is the primary provider; on outage the tool auto-falls back to fal.ai BiRefNet v2 so single-image calls never silently fail. Works best on photographic subjects (people, products, animals); transparent-PNG inputs have no foreground to segment.
| Name | Required | Description | Default |
|---|---|---|---|
| imageUrl | Yes | Public URL of the source image. |
aetherwave_remove_background_videoRemove background from videoInspect
Strips the background from a video frame-by-frame using rembg (u2netp) on AetherWave's Python service. Pass a public videoUrl. Choose bgType: "transparent" for an alpha-channel WebM output (compositing) or bgType: "color" with a customColor hex for a solid replacement. 2 credits per second. Slowest tool in the surface (per-frame processing); a 6s clip takes ~4 min, a 30s clip ~15-20 min. Works best on subjects with clear edges (people, products). Returns the processed video URL (R2-hosted).
| Name | Required | Description | Default |
|---|---|---|---|
| bgType | No | 'transparent' = alpha WebM output (default). 'color' = solid replacement using customColor. | |
| videoUrl | Yes | Public URL of the source video (MP4). | |
| customColor | No | Hex color for solid background when bgType='color' (e.g. '#00ff00'). Default green. |
aetherwave_upscale_imageUpscale image (Topaz)Inspect
Upscales a source image using Topaz's high-fidelity upscaler. Pass a public imageUrl and an upscaleFactor. Credit cost depends on the source resolution × factor; small images cost less than large ones at the same factor. Returns the upscaled image URL.
| Name | Required | Description | Default |
|---|---|---|---|
| imageUrl | Yes | Public URL of the source image. | |
| upscaleFactor | No | Upscale multiplier. Defaults to '2x'. '8x' is heavy; use only on small sources. |
aetherwave_upscale_videoUpscale video (Atlas Video Upscaler)Inspect
Upscales a source video to 1080p or 2K using Atlas. Pass a public videoUrl and the target resolution. Cost is per-second (7 cr/s @ 1080p, 9 cr/s @ 2K). Atlas-side limits: clips up to 53s at 1080p, 23s at 2K, source must be <=30fps. Returns the upscaled video URL (R2-hosted).
| Name | Required | Description | Default |
|---|---|---|---|
| videoUrl | Yes | Public URL of the source video (MP4). | |
| targetResolution | No | Target output resolution. Defaults to '1080p'. '2k' is more expensive and limited to ~23s clips. |
Claim this connector by publishing a /.well-known/glama.json file on your server's domain with the following structure:
{
"$schema": "https://glama.ai/mcp/schemas/connector.json",
"maintainers": [{ "email": "your-email@example.com" }]
}The email address must match the email associated with your Glama account. Once published, Glama will automatically detect and verify the file within a few minutes.
Control your server's listing on Glama, including description and metadata
Access analytics and receive server usage reports
Get monitoring and health status updates for your server
Feature your server to boost visibility and reach more users
For users:
Full audit trail – every tool call is logged with inputs and outputs for compliance and debugging
Granular tool control – enable or disable individual tools per connector to limit what your AI agents can do
Centralized credential management – store and rotate API keys and OAuth tokens in one place
Change alerts – get notified when a connector changes its schema, adds or removes tools, or updates tool definitions, so nothing breaks silently
For server owners:
Proven adoption – public usage metrics on your listing show real-world traction and build trust with prospective users
Tool-level analytics – see which tools are being used most, helping you prioritize development and documentation
Direct user feedback – users can report issues and suggest improvements through the listing, giving you a channel you would not have otherwise
The connector status is unhealthy when Glama is unable to successfully connect to the server. This can happen for several reasons:
The server is experiencing an outage
The URL of the server is wrong
Credentials required to access the server are missing or invalid
If you are the owner of this MCP connector and would like to make modifications to the listing, including providing test credentials for accessing the server, please contact support@glama.ai.
Discussions
No comments yet. Be the first to start the discussion!
Related MCP Servers
AlicenseAqualityBmaintenanceMonitors brand mentions, citations, sentiment, competitor share of voice, and GEO performance across AI search engines.Last updated1MIT- Alicense-qualityAmaintenanceProvides regulatory and compliance intelligence from free government sources, including rules, recalls, enforcement actions, and comment deadlines, classified by industry and severity.Last updatedMIT
- AlicenseAqualityAmaintenanceCompetitive intelligence platform with 24 tools. Monitor competitor pricing, content, positioning, tech stacks, and AI visibility — track how ChatGPT, Claude, and Gemini rank your brand.Last updated332MIT
- Flicense-qualityCmaintenanceEnables users to monitor and analyze competitor LinkedIn posts, extract insights, generate original post concepts, and create graphics-ready content with metadata.Last updated