VLLM MCP Server
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| PYTHONPATH | No | Python path for module resolution | src |
| VLLM_MCP_HOST | No | Server host (optional) | localhost |
| VLLM_MCP_PORT | No | Server port (optional) | 8080 |
| OPENAI_API_KEY | No | Your OpenAI API key | |
| OPENAI_BASE_URL | No | OpenAI base URL (optional) | https://api.openai.com/v1 |
| DASHSCOPE_API_KEY | Yes | Your Dashscope API key | |
| VLLM_MCP_LOG_LEVEL | No | Log level (optional) | INFO |
| VLLM_MCP_TRANSPORT | No | Transport type (optional) | stdio |
| OPENAI_DEFAULT_MODEL | No | Default OpenAI model to use | gpt-4o |
| DASHSCOPE_DEFAULT_MODEL | No | Default Dashscope model to use | qwen-vl-plus |
| OPENAI_SUPPORTED_MODELS | No | Comma-separated list of supported OpenAI models | gpt-4o,gpt-4o-mini,gpt-4-turbo,gpt-4-vision-preview |
| DASHSCOPE_SUPPORTED_MODELS | No | Comma-separated list of supported Dashscope models | qwen-vl-plus,qwen-vl-max,qwen-vl-chat,qwen2-vl-7b-instruct,qwen2-vl-72b-instruct |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Server capabilities have not been inspected yet.
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| generate_multimodal_responseC | Generate response from multimodal model. |
| list_available_providersB | List available model providers and their configurations. |
| validate_multimodal_requestB | Validate if a multimodal request is supported. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose with no overlap. generate_multimodal_response handles actual generation, list_available_providers provides configuration information, and validate_multimodal_request performs validation checks. An agent can easily distinguish between these three functions.
All three tools follow a consistent verb_noun naming pattern (generate_multimodal_response, list_available_providers, validate_multimodal_request). The naming is uniform, predictable, and clearly communicates each tool's function without any style mixing or deviations.
With only 3 tools, this server feels somewhat thin for a multimodal generation service. While the tools cover core functionality, typical MCP servers in this domain would include additional operations like model management, conversation history, or specialized generation modes. The count is borderline but functional.
The tool set covers the essential workflow: checking available providers, validating requests, and generating responses. However, there are minor gaps such as no tool for managing conversation context, handling streaming responses, or providing model-specific configuration options that would enhance the agent's ability to work effectively with multimodal models.