novelai-mcp
This server lets MCP clients (Claude, Cursor, etc.) generate and edit images using NovelAI Diffusion models.
Text-to-image generation with V5 (Full/Curated) and V4.5 models
Custom sizes, multi-character prompts with positions, in-image text, transparent backgrounds, comic panels, furry/background modes
Image-to-image redrawing/restyling with adjustable strength
Inpainting with masks to redraw specific image areas
Director Tools: background removal, lineart, sketch, colorize, emotion, declutter
Neural upscaling
Vibe Transfer to create images in a reference image's style/palette
Danbooru/NovelAI tag suggestions
Generated images are saved to disk and returned to the client
Supports negative prompts, seeds, samplers, and output directories
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@novelai-mcpGenerate a portrait of a girl with silver hair and blue eyes in a sunflower field."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
NovelAI MCP Server
An MCP server that lets Claude, Cursor and other MCP clients generate and edit images with NovelAI.
Supports NovelAI Diffusion V5 (Full and Curated) and V4.5: text-to-image, img2img, inpainting, Director Tools, Vibe Transfer, upscaling and tag suggestions.
This is an unofficial project. It is not affiliated with or endorsed by NovelAI / Anlatan.
Tools
Tool | What it does |
| Text-to-image. Size presets or custom size, multi-character prompts with positions, in-image text, transparent background, comic panels, furry and background modes. |
| Redraws an existing image from a new prompt. |
| Redraws the masked area of an image. White in the mask = repaint, black = keep. |
| Director Tools: |
| Neural upscale. NovelAI chooses the factor. |
| Generates a new image in the style and palette of a reference image. |
| Danbooru/NovelAI tag autocomplete. |
Generated images are saved to disk and also returned to the client, so the model can see them.
Related MCP server: NovelAI MCP Server
Requirements
A NovelAI subscription and a Persistent API Token: NovelAI → Settings → Account → Persistent API Tokens.
uv. It runs the server without a manual install.
Cost
Requests spend your Anlas the same way the NovelAI website does. Large sizes, more steps, upscale and Vibe Transfer cost more. Check your Anlas balance before running many generations. The server does not limit spending.
Setup
Claude Code
claude mcp add novelai -e NOVELAI_API_KEY=your_token -- uvx novelai-mcpClaude Desktop
Edit claude_desktop_config.json:
Windows:
%APPDATA%\Claude\claude_desktop_config.jsonmacOS:
~/Library/Application Support/Claude/claude_desktop_config.json
{
"mcpServers": {
"novelai": {
"command": "uvx",
"args": ["novelai-mcp"],
"env": {
"NOVELAI_API_KEY": "your_token"
}
}
}
}Restart Claude Desktop after editing.
Cursor
Add the same mcpServers block to .cursor/mcp.json or to Settings → MCP.
Install from GitHub instead of PyPI
Replace "args": ["novelai-mcp"] with:
"args": ["--from", "git+https://github.com/DesmondFox/novelai-mcp", "novelai-mcp"]Configuration
Variable | Required | Default | Meaning |
| yes | — | Persistent API Token. |
| no |
| Where images are saved. Each tool also accepts |
Usage
Ask the model in plain language. It picks the tool and parameters. Examples:
"Generate a portrait of a girl with silver hair and blue eyes in a sunflower field."
"Make a chibi cat girl holding a coffee mug, transparent background."
"Draw a tavern scene: a knight girl on the left, a mage boy on the right."
"Cyberpunk storefront at night with a neon sign that says NEON DREAMS."
"Take
~/Pictures/NovelAI/t2i_20260101_120000_abc123.pngand change her expression to smug.""Remove the background from that image, then upscale it."
"Suggest tags for 'kimono'."
Useful parameters of novelai_generate_image:
Parameter | Default | Notes |
|
| Also |
|
|
|
| — | Custom size, rounded to a multiple of 64. Both must be set. |
| 23 | Same default as the NovelAI website. |
| 5.0 | Prompt guidance. |
| random | Set it to reproduce an image. The used seed is returned. |
|
|
|
| — | List of |
| — | Text to draw inside the image. |
| false | Transparent background. |
| false | Adds comic/manga panel tags. |
Development
git clone https://github.com/DesmondFox/novelai-mcp
cd novelai-mcp
pip install -e ".[dev]"
pytest
ruff check .For local development you can put the key in a .env file (see .env.example) instead of the client config.
License
Available Tools
7 toolsnovelai_augmentB
NovelAI Director Tools for image manipulations: remove background ('bg-removal'), extract line art ('lineart'), convert to sketch ('sketch'), colorize sketch ('colorize'), change facial expressions/emotions ('emotion'), or declutter ('declutter').
| Name | Required | Description | Default |
|---|---|---|---|
| tool | Yes | ||
| defry | No | ||
| prompt | No | ||
| emotion | No | ||
| image_path | Yes | ||
| output_dir | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description should disclose behavioral details. It only lists available manipulations and gives no information about side effects, file handling, requirements, or reversibility. This is a significant gap for a tool that likely modifies images.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that efficiently communicates the core purpose and valid operations without wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Without output schema or annotations, the description is too thin for a 6-parameter tool. It doesn't cover input/output paths, optional parameters, or operational details, leaving an agent to guess on several fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage, the description must explain parameters. It adds meaning by enumerating valid values for 'tool' and referencing 'emotion', but leaves prompt, defry, and output_dir unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a specific action ('image manipulations') and enumerates distinct operations with their string values. It differentiates from siblings by scope, though it doesn't name sibling tools explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage through the list of operations (e.g., use for background removal), but doesn't provide explicit when-to-use vs alternatives like novelai_inpaint, nor conditions for selecting this tool over others.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
novelai_generate_imageB
Generate anime/art images with NovelAI Diffusion models (prioritizing V5 Full/Curated, V4.5). Supports natural language, Danbooru tags, multilingual prompts, transparent background, and custom dimensions.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | ||
| size | No | portrait | |
| model | No | nai-diffusion-5-full | |
| scale | No | ||
| steps | No | ||
| width | No | ||
| height | No | ||
| prompt | Yes | ||
| sampler | No | k_euler_ancestral | |
| uc_preset | No | light | |
| allow_text | No | ||
| characters | No | ||
| furry_mode | No | ||
| output_dir | No | ||
| cfg_rescale | No | ||
| render_text | No | ||
| transparent | No | ||
| comic_panels | No | ||
| quality_toggle | No | ||
| background_mode | No | ||
| negative_prompt | No | ||
| dynamic_thresholding | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions supported prompt styles and output options like transparent background and custom dimensions, but it does not disclose side effects (e.g., file saving, output format, API costs, rate limits) or any destructive actions. The agent is left uninformed about what happens after generation, such as whether the image is returned, saved to disk, or requires further steps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly written sentence that front-loads the core purpose and highlights the most salient features. There is no filler, and every clause adds relevant information about capabilities. This is appropriately concise for a tool with such a broad feature set.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (22 parameters, no output schema), the description is significantly under-specified. It does not explain the return value, how the generated image is delivered, or the meaning of advanced parameters. The absence of any output schema means the description should at least hint at the result format, but it does not. An agent would struggle to correctly configure the tool beyond the basics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does mention transparent background and custom dimensions, which map to the 'transparent' and 'width'/'height' parameters, and it hints at prompt flexibility (natural language, Danbooru tags, multilingual). However, the description does not explain the meaning or effect of most parameters (e.g., scale, steps, sampler, cfg_rescale, dynamic_thresholding), leaving the agent to guess for 18+ parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Generate') and resource ('anime/art images') and specifies the models used. It is distinct from sibling tools like upscale, img2img, and inpaint, which have different purposes. The capability list (natural language, Danbooru tags, transparent background, custom dimensions) makes the tool's scope unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention scenarios where img2img, inpaint, or vibe_transfer would be more appropriate, nor does it state any exclusions or prerequisites. The only hint is the mention of model priorities, which is about model selection, not usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
novelai_img2imgA
Modify or restyle an existing image using NovelAI Image-to-Image. Applies new prompt with adjustable change strength (0.0 to 1.0) and optional noise.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | ||
| size | No | ||
| model | No | nai-diffusion-5-full | |
| noise | No | ||
| scale | No | ||
| steps | No | ||
| prompt | Yes | ||
| sampler | No | k_euler_ancestral | |
| strength | No | ||
| image_path | Yes | ||
| output_dir | No | ||
| negative_prompt | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden of behavioral disclosure. It mentions adjustable change strength (0.0 to 1.0) and optional noise, which gives some insight into behavior, but it does not disclose important behavioral traits such as whether the original image is modified in place, output file handling, or potential side effects. The description is not misleading but is incomplete for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, two sentences, and front-loads the core purpose. It mentions the key adjustable parameters (strength and noise) without excessive detail. However, it could be slightly more structured by explicitly listing the most important parameters or usage context, but it is appropriately sized for a tool description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 12 parameters, no output schema, and no annotations, the description is incomplete. It does not explain the meaning of most parameters, the expected output format, or how the tool handles file paths and outputs. An agent would need to infer or guess the semantics of many parameters, making this insufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the 12 parameters. It explains the core concept of strength (change strength 0.0 to 1.0) and noise, but does not explain other key parameters like seed, size, model, scale, steps, sampler, output_dir, or negative_prompt. The description adds some meaning beyond the schema but leaves most parameters unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool modifies or restyles an existing image using NovelAI Image-to-Image, with a specific verb ('Modify or restyle') and resource ('existing image'). It distinguishes itself from siblings like novelai_generate_image (which creates new images) and novelai_inpaint (which edits specific regions), making its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for modifying/restyling existing images, which differentiates it from generation tools, but it does not explicitly state when to use this tool versus alternatives like novelai_inpaint or novelai_vibe_transfer. It mentions adjustable change strength and optional noise, giving some context, but lacks explicit when-to-use or when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
novelai_inpaintB
Inpaint / redraw specific parts of an image using a mask. Allows replacing clothes, faces, hair, background or fixing anatomy.
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | ||
| model | No | nai-diffusion-5-full | |
| scale | No | ||
| steps | No | ||
| prompt | Yes | ||
| sampler | No | k_euler_ancestral | |
| strength | No | ||
| mask_path | Yes | ||
| image_path | Yes | ||
| output_dir | No | ||
| negative_prompt | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of disclosing behavior, but it only states the high-level redrawing operation. It does not explain output behavior, whether images are written to output_dir, how the mask is interpreted, or any side effects or requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler. The concrete examples of use cases add useful information without bloating the text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with 11 parameters, 0% schema coverage, no annotations, and no output schema, this description is too thin for an agent to confidently select and invoke it correctly. It lacks guidance on required inputs, output behavior, parameter roles, and how it differs from closely related siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description needed to compensate by explaining the 11 parameters, but it only indirectly hints at image, mask, and prompt through the phrase 'using a mask'. It adds no meaning for seed, scale, steps, sampler, strength, negative_prompt, model, or output_dir, leaving the bare schema property names to carry the semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Inpaint / redraw') and resource ('specific parts of an image using a mask') and gives concrete use cases like replacing clothes, faces, hair, background, or fixing anatomy. It clearly states what the tool does, though it does not explicitly contrast it with sibling tools like novelai_img2img or novelai_generate_image.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'specific parts of an image using a mask' implies when this tool should be used, but there is no explicit guidance about when to prefer it over novelai_img2img, novelai_generate_image, or other siblings. No exclusions or alternative routing hints are provided; the agent must infer the use case from the name and mask requirement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
novelai_suggest_tagsB
Search and suggest valid Danbooru/NovelAI tags matching a query string.
| Name | Required | Description | Default |
|---|---|---|---|
| lang | No | en | |
| model | No | nai-diffusion-5-full | |
| query | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavioral traits on its own. It only says the tool searches and suggests valid tags, but doesn't clarify whether results mix Danbooru and NovelAI tags, how many results are returned, ordering rules, pagination, or any rate limits. For a tool with an empty annotation block, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that contains no filler. It fits the purpose, resource, and filter into eleven words, making it easy to parse. There is nothing extraneous to remove.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple and has an output schema, so return-value details are covered. The critical missing pieces are the semantics of lang and model and any behavioral limits. However, since query is required and lang/model have defaults, the tool is minimally callable. It is adequate but not complete for an agent needing to fine-tune suggestions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for all three parameters. It only hints at 'query' via 'query string,' while 'lang' and 'model' are left entirely unexplained. An agent must rely on parameter names and defaults, which is insufficient for customizing behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Search and suggest valid Danbooru/NovelAI tags' with a clear filter ('matching a query string'). This immediately distinguishes it from the sibling tools, which are all image generation/editing operations. Even with no title, the tool's role is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit when-to-use or when-not-to-use guidance, and it does not name any alternative tool. However, the sibling context (all image-focused) implies this is the tag-lookup tool for prompt building, and the word 'suggest' hints at autocomplete usage. The guidance is implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
novelai_upscaleB
Upscale an image with NovelAI's neural upscaler.
| Name | Required | Description | Default |
|---|---|---|---|
| model | No | nai-diffusion-5-full | |
| image_path | Yes | ||
| output_dir | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It only states the core action and does not reveal side effects, output behavior, output location semantics, or any requirements/limitations. The agent learns what the tool does but not how its behavior manifests beyond that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler. Every word earns its place, and the primary action is stated immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, no output schema, and 3 parameters, the description is too thin. An agent is not told what the tool returns, whether it writes to output_dir, what happens when output_dir is null, or what model options are available. The basic purpose is clear, but the operational envelope is incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the missing parameter documentation, but it does not. The schema gives titles and a default model value, but the model value 'nai-diffusion-5-full' and the meaning of output_dir null are unexplained. 'Image Path' is self-evident from its title, giving only marginal semantic value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('upscale') with a specific resource ('an image') and names the provider ('NovelAI's neural upscaler'). It is immediately distinguishable from sibling tools like novelai_generate_image or novelai_inpaint, which clearly do different operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the obvious use case: use when an image needs to be upscaled. However, it provides no explicit guidance about when to prefer this tool over siblings, nor any exclusions or alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
novelai_vibe_transferC
Generate a new image influenced by the style, colors, and atmosphere of a reference image (Vibe Transfer).
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | ||
| size | No | portrait | |
| model | No | nai-diffusion-4-5-full | |
| scale | No | ||
| steps | No | ||
| prompt | Yes | ||
| sampler | No | k_euler_ancestral | |
| output_dir | No | ||
| negative_prompt | No | ||
| reference_strength | No | ||
| reference_image_path | Yes | ||
| information_extracted | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description must disclose side effects, expected outputs, and operational details by itself. It only repeats the basic purpose and does not mention that an image will be written to disk, what model settings apply, or how the reference image is processed. For a multi-parameter generation tool this is a material transparency gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single sentence is concise, but for a 12-parameter image-generation tool it is severely under-specified. It front-loads the main idea yet omits essential operational and parameter context, so the brevity is not effective.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (12 params, no output schema, no annotations, no enums, several influencing knobs like reference_strength and information_extracted), a one-line description is far from enough for an agent to invoke it correctly. Critical information about output behavior, defaults, and parameter interactions is absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0% and 12 parameters, the description contributes no parameter-level meaning. The only hinted concept is 'reference image,' which matches reference_image_path, but prompt, reference_strength, information_extracted, sampler, scale, and others are left entirely unexplained. The description does not compensate for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a generative action ('Generate a new image') and the distinctive mechanism ('influenced by the style, colors, and atmosphere of a reference image'), which separates it from sibling tools like plain generation. It lacks an explicit contrast with similar image-to-image siblings, but the Vibe Transfer concept is reasonably distinct.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance about when to choose this tool over siblings such as novelai_img2img or novelai_generate_image. The phrase 'influenced by style... atmosphere' implies a use case, but there are no explicit when-to-use or when-not-to-use conditions, nor mention of alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.0- First observed
novelai_augment - First observed
novelai_generate_image - First observed
novelai_img2img - First observed
novelai_inpaint - First observed
novelai_suggest_tags - First observed
novelai_upscale - First observed
novelai_vibe_transfer
TDQS
Scored across 7 tools
Each tool has a distinct purpose: upscaling, tag suggestion, generation, img2img, inpainting, augmentations, and vibe transfer. However, novelai_img2img and novelai_vibe_transfer both involve modifying or generating from an existing image, which could cause some confusion. The descriptions clarify the difference (strength vs style influence), but a user might initially hesitate.
All tools follow a consistent 'novelai_' prefix followed by a clear verb or action (generate, img2img, inpaint, augment, vibe_transfer). The naming is uniform and predictable, making it easy to understand the tool's function from its name.
7 tools is an appropriate scope for an image generation and editing MCP server. Each tool covers a distinct aspect of the workflow: generation, editing (img2img, inpaint), augmentation, style transfer, and utilities (upscale, tag suggestions). This is neither too thin nor too heavy.
The tool set covers the core image generation and editing workflows: generating new images, refining them, inpainting, style transfer, and enhancements. A notable gap is the lack of explicit image-to-image with mask (which is partially covered by inpaint) and perhaps batch processing, but the provided tools are sufficient for most common use cases. Minor gap: no explicit 'delete' or 'get' tools, but those are not typical for generation servers.
Maintenance
Related MCP Connectors
Generate AI images, videos, music, SFX & speech in any AI assistant. Results appear inline in chat.
Generate AI images, video, music, and sound effects, and upscale them, from any MCP client.
LLM chat, text tools, image generation, editing, batch image jobs, and asynchronous video generation
Image and video AI tools and your own pipelines, run from any AI assistant.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables AI assistants to generate high-quality images using Google's Gemini and Imagen models with support for multiple aspect ratios, dynamic model selection, and direct file saving capabilities.MIT
- AlicenseNot gradedqualityBmaintenanceEnables AI assistants to generate images using NovelAI, supporting text-to-image, image-to-image, and tag suggestions.3MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to generate and edit images using multiple models through OpenRouter, with features like style presets, batch operations, and variations.22 npm1MIT
- AlicenseAqualityDmaintenanceEnables AI image generation via Antigravity (Google Gemini) and OpenAI DALL-E 3, supporting text-to-image, image editing, multiple outputs, and character consistency.510 npm9MIT