DINO-X Image Detection MCP Server
The DINO-X MCP Server enables comprehensive image analysis and detection for multimodal applications.
Full-Scene Object Detection: Identify all objects in an image with categories, counts, coordinates, and optional descriptions
Text-Prompted Object Detection: Locate specific objects using English noun prompts (e.g.,
person.car) with bounding boxes and descriptionsHuman Pose Estimation: Detect 17 keypoints per person for body posture and movement analysis
Detection Visualization: Generate annotated images with bounding boxes and labels, saved locally (STDIO mode only)
Structured Outputs: Provides detailed object data suitable for VQA and multi-step reasoning tasks
Flexible Input Sources: Supports
https://URLs andfile://URIs with common image formats (JPG, JPEG, WebP, PNG)Multiple Runtime Modes: Available in local (STDIO) and remote (Streamable HTTP) modes for deployment flexibility
Provides visualization capabilities for object detection results, allowing bounding boxes, keypoints, and other visual markers to be overlaid on the original image for better presentation of analysis results.
Enables running the DINO-X MCP server, which provides tools for fine-grained object detection and image understanding in AI applications.
Used as the package manager for installing and building the DINO-X MCP server project.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@DINO-X Image Detection MCP Serverdetect all objects in this image of a city street"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
DINO-X MCP Server
English | 中文
DINO-X Official MCP Server — powered by the DINO-X and Grounding DINO models — brings fine-grained object detection and image understanding to your multimodal applications.
Why DINO-X MCP?
With DINO-X MCP, you can:
Fine-Grained Understanding: Full image detection, object detection, and region-level descriptions.
Structured Outputs: Get object categories, counts, locations, and attributes for VQA and multi-step reasoning tasks.
Composable: Works seamlessly with other MCP servers to build end-to-end visual agents or automation pipelines.
Related MCP server: Vision MCP Server
Transport Modes
DINO-X MCP supports two transport modes:
Feature | STDIO (default) | Streamable HTTP |
Runtime | Local | Local or Cloud |
Transport | Standard I/O | HTTP (streaming responses) |
Input source |
|
|
Visualization | Supported (saves annotated images locally) | Not supported (for now) |
Quick Start
1. Prepare an MCP client
Any MCP-compatible client works, e.g.:
2. Get your API key
Apply on the DINO-X platform: Request API Key (new users get free quota).
3. Configure MCP
Option A: Official Hosted Streamable HTTP (Recommended)
Add to your MCP client config and replace with your API key:
{
"mcpServers": {
"dinox-mcp": {
"url": "https://mcp.deepdataspace.com/mcp?key=your-api-key"
}
}
}Option B: Use the NPM package locally (STDIO)
Install Node.js first
Download the installer from nodejs.org
Or use command:
# macOS / Linux
curl -o- https://raw.githubusercontent.com/nvm-sh/nvm/v0.40.1/install.sh | bash
# or
wget -qO- https://raw.githubusercontent.com/nvm-sh/nvm/v0.40.1/install.sh | bash
# load nvm into current shell (choose the one you use)
source ~/.bashrc || true
source ~/.zshrc || true
# install and use LTS Node.js
nvm install --lts
nvm use --lts
# Windows (one of the following)
winget install OpenJS.NodeJS.LTS
# or with Chocolatey (in admin PowerShell)
iwr -useb https://raw.githubusercontent.com/chocolatey/chocolatey/master/chocolateyInstall/InstallChocolatey.ps1 | iex
choco install nodejs-lts -yConfigure your MCP client:
{
"mcpServers": {
"dinox-mcp": {
"command": "npx",
"args": ["-y", "@deepdataspace/dinox-mcp"],
"env": {
"DINOX_API_KEY": "your-api-key-here",
"IMAGE_STORAGE_DIRECTORY": "/path/to/your/image/directory"
}
}
}
}Note: Replace your-api-key-here with your real key.
Option C: Run from source locally
Make sure Node.js is installed (see Option B), then:
# clone
git clone https://github.com/IDEA-Research/DINO-X-MCP.git
cd DINO-X-MCP
# install deps
npm install
# build
npm run buildConfigure your MCP client:
{
"mcpServers": {
"dinox-mcp": {
"command": "node",
"args": ["/path/to/DINO-X-MCP/build/index.js"],
"env": {
"DINOX_API_KEY": "your-api-key-here",
"IMAGE_STORAGE_DIRECTORY": "/path/to/your/image/directory"
}
}
}
}CLI Flags & Environment Variables
Common flags
--http: start in Streamable HTTP mode (otherwise STDIO by default)--stdio: force STDIO mode--dinox-api-key=...: set API key--enable-client-key: allow API key via URL?key=(Streamable HTTP only)--port=8080: HTTP port (default 3020)
Environment variables
DINOX_API_KEY(required/conditionally required): DINO-X platform API keyIMAGE_STORAGE_DIRECTORY(optional, STDIO): directory to save annotated imagesAUTH_TOKEN(optional, HTTP): if set, client must sendAuthorization: Bearer <token>
Examples:
# STDIO (local)
node build/index.js --dinox-api-key=your-api-key
# Streamable HTTP (server provides a shared API key)
node build/index.js --http --dinox-api-key=your-api-key
# Streamable HTTP (custom port)
node build/index.js --http --dinox-api-key=your-api-key --port=8080
# Streamable HTTP (require client-provided API key via URL)
node build/index.js --http --enable-client-keyClient config when using ?key=:
{
"mcpServers": {
"dinox-mcp": {
"url": "http://localhost:3020/mcp?key=your-api-key"
}
}
}Using AUTH_TOKEN with a gateway that injects Authorization: Bearer <token>:
AUTH_TOKEN=my-token node build/index.js --http --enable-client-keyClient example with supergateway:
{
"mcpServers": {
"dinox-mcp": {
"command": "npx",
"args": [
"-y",
"supergateway",
"--streamableHttp",
"http://localhost:3020/mcp?key=your-api-key",
"--oauth2Bearer",
"my-token"
]
}
}
}Tools
Capability | Tool ID | Transport | Input | Output |
Full-scene object detection |
| STDIO / HTTP | Image URL | Category + bbox + (optional) captions |
Text-prompted object detection |
| STDIO / HTTP | Image URL + English nouns (dot-separated for multiple, e.g., | Target object bbox + (optional) captions |
Human pose estimation |
| STDIO / HTTP | Image URL | 17 keypoints + bbox + (optional) captions |
Visualization |
| STDIO only | Image URL + detection results array | Local path to annotated image |
🎬 Use Cases
🎯 Scenario | 📝 Input | ✨ Output |
Detection & Localization | 💬 Prompt: | |
Object Counting | 💬 Prompt: | |
Feature Detection | 💬 Prompt: | |
Attribute Reasoning | 💬 Prompt: | |
Full Scene Detection | 💬 Prompt: |
|
Pose Analysis | 💬 Prompt: |
FAQ
Supported image sources?
STDIO:
file://andhttps://Streamable HTTP:
https://only
Supported image formats?
jpg, jpeg, webp, png
Development & Debugging
Use watch mode to auto-rebuild during development:
npm run watchUse MCP Inspector for debugging:
npm run inspectorLicense
Apache License 2.0
Available Tools
4 toolsdetect-all-objectsB
Analyze an image to detect all identifiable objects, returning the category, count, coordinate positions and detailed descriptions for each object.
| Name | Required | Description | Default |
|---|---|---|---|
| imageFileUri | Yes | URI of the input image. Preferred for remote or local files. Must start with 'https://' or 'file://'. | |
| includeDescription | Yes | Whether to return a description of the objects detected in the image, but will take longer to process. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions that including descriptions 'will take longer to process,' which adds some context about performance impact. However, it doesn't address other important behavioral aspects like error handling, rate limits, authentication requirements, or what happens with invalid inputs, which are significant gaps for a tool with no annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that efficiently conveys the tool's purpose, action, and output without any redundant information. It's appropriately sized and front-loaded, with every word serving a clear purpose in explaining what the tool does.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (image analysis with 2 parameters), no annotations, and no output schema, the description is partially complete. It clearly states what the tool does and what it returns, but lacks details on behavioral traits, error conditions, and output format specifics. The absence of an output schema means the description should ideally explain return values more thoroughly, which it doesn't do beyond listing output categories.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, so the schema already fully documents both parameters (imageFileUri and includeDescription). The description adds minimal value beyond the schema by briefly mentioning that includeDescription affects processing time, but doesn't provide additional syntax, format details, or usage context for the parameters. This meets the baseline expectation when schema coverage is high.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Analyze an image to detect all identifiable objects') and distinguishes it from siblings by emphasizing 'all identifiable objects' rather than specific types like human poses or text-based detection. It provides the verb+resource+scope combination that makes the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus the sibling tools (detect-human-pose-keypoints, detect-objects-by-text, visualize-detection-result). It doesn't mention alternatives, exclusions, or specific contexts where this tool is preferred over others, leaving the agent with no comparative usage information.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
detect-human-pose-keypointsC
Detects 17 keypoints for each person in an image, supporting body posture and movement analysis.
| Name | Required | Description | Default |
|---|---|---|---|
| imageFileUri | Yes | URI of the input image. Preferred for remote or local files. Must start with 'https://' or 'file://'. | |
| includeDescription | Yes | Whether to return a description of the objects detected in the image, but will take longer to process. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden but offers minimal behavioral insight. It mentions the output (17 keypoints per person) and implies analysis use, but lacks details on performance (e.g., speed, accuracy), limitations (e.g., image quality requirements), or side effects. The description doesn't contradict annotations, but it's insufficient for a mutation-like detection tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core purpose without unnecessary details. Every word contributes to understanding the tool's function, making it appropriately concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (detection with 2 parameters) and lack of annotations and output schema, the description is incomplete. It doesn't explain return values (e.g., keypoint coordinates, confidence scores), error handling, or prerequisites (e.g., image format support), leaving significant gaps for agent usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents both parameters. The description adds no parameter-specific information beyond what's in the schema, such as explaining how 'includeDescription' relates to keypoints or image processing. This meets the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: detecting 17 keypoints per person in an image for posture/movement analysis. It specifies the verb ('detects'), resource ('keypoints'), and scope ('each person in an image'), but doesn't explicitly differentiate from sibling tools like 'detect-all-objects' or 'detect-objects-by-text', which likely serve different detection purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention sibling tools or contexts where human pose detection is preferred over general object detection, leaving the agent to infer usage based on tool names alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
detect-objects-by-textA
Analyze an image based on a text prompt to identify and count specific objects, and return detailed descriptions of the objects and their 2D coordinates.
| Name | Required | Description | Default |
|---|---|---|---|
| imageFileUri | Yes | URI of the input image. Preferred for remote or local files. Must start with 'https://' or 'file://'. | |
| textPrompt | Yes | Nouns of target objects (English only, avoid adjectives). Use periods to separate multiple categories (e.g., 'person.car.traffic light'). | |
| includeDescription | Yes | Whether to return a description of the objects detected in the image, but will take longer to process. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions that including descriptions will 'take longer to process', which is useful behavioral context. However, it doesn't disclose other critical traits like potential rate limits, error conditions, authentication needs, or what happens with invalid inputs. For a tool with no annotation coverage, this leaves significant gaps in understanding its behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that efficiently conveys the core functionality without unnecessary words. It's front-loaded with the main action and outcome, making it easy to understand at a glance. Every part of the sentence contributes directly to explaining what the tool does.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that there's no output schema and no annotations, the description should ideally provide more context about return values and behavioral constraints. While it mentions what will be returned (descriptions and coordinates), it doesn't specify the format or structure of the output. For a tool with 3 required parameters and no structured output documentation, the description is adequate but leaves room for improvement in completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters thoroughly. The description doesn't add any meaningful semantic information beyond what's in the schema descriptions (e.g., it doesn't explain the format of returned coordinates or how object counts are calculated). Baseline score of 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('analyze an image based on a text prompt'), the resource ('image'), and the outcome ('identify and count specific objects, return detailed descriptions and 2D coordinates'). It distinguishes from sibling tools like 'detect-all-objects' by specifying text-based filtering and from 'detect-human-pose-keypoints' by focusing on object detection rather than pose analysis.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage through its purpose statement but doesn't explicitly state when to use this tool versus alternatives like 'detect-all-objects' or 'visualize-detection-result'. It mentions the text prompt requirement, which suggests this tool is for targeted object detection, but lacks clear guidance on scenarios where other tools might be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
visualize-detection-resultB
Visualize detection results by drawing bounding boxes and labels on the original image. Images are saved to the directory specified by IMAGE_STORAGE_DIRECTORY environment variable.
| Name | Required | Description | Default |
|---|---|---|---|
| imageFileUri | Yes | URI of the input image. Preferred for remote or local files. Must start with 'https://' or 'file://'. | |
| detections | Yes | Array of detection results with name and bbox information. | |
| fontSize | No | Font size for labels (default: 24) | |
| boxThickness | No | Thickness of bounding box lines (default: 4) | |
| showLabels | No | Whether to show category labels (default: true) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden but provides minimal behavioral information. It mentions that 'Images are saved to the directory specified by IMAGE_STORAGE_DIRECTORY environment variable' which is useful context about output location, but doesn't disclose important behavioral traits like whether this is a read-only operation, what happens if the directory doesn't exist, what file format is used, or whether the original image is modified versus a copy being created.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately concise with two clear sentences. The first sentence states the core purpose, and the second provides important behavioral context about where results are saved. There's no wasted text, though it could be slightly more structured with clearer separation of concerns.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 parameters, no annotations, and no output schema, the description is insufficient. It doesn't explain what the tool returns (success/failure indicators, output file paths, error conditions), doesn't mention dependencies on sibling tools, and provides minimal behavioral context. The description leaves too many important questions unanswered for effective agent use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the baseline is 3. The description adds no additional parameter information beyond what's already documented in the schema. It doesn't explain relationships between parameters, provide examples of valid detections arrays, or offer guidance on appropriate values for fontSize and boxThickness.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Visualize detection results by drawing bounding boxes and labels on the original image') and distinguishes it from sibling detection tools by focusing on visualization rather than detection. It specifies the exact resource being modified (the original image with annotations added).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (like needing detection results first), doesn't explain relationships to sibling detection tools, and offers no context about appropriate use cases or limitations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
Each tool has a clearly distinct purpose: detect-all-objects performs general object detection, detect-human-pose-keypoints focuses on human pose analysis, detect-objects-by-text enables text-guided detection, and visualize-detection-result handles visualization. There is no overlap in functionality, making tool selection straightforward for an agent.
All tool names follow a consistent verb_noun pattern with hyphens (e.g., detect-all-objects, detect-human-pose-keypoints). The naming is predictable and readable throughout, with no deviations in style or convention.
With 4 tools, the server is well-scoped for image detection tasks. Each tool earns its place by covering distinct aspects: detection (general, pose-specific, text-guided) and visualization, avoiding bloat while providing essential functionality.
The tool set covers core detection workflows (general, pose, text-guided) and visualization, with no obvious dead ends. A minor gap exists in lacking tools for modifying or deleting detection results, but agents can work around this by re-running detections or handling data externally.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Roboflow computer vision for AI agents: datasets, annotation, versioning, workflows, inference.
E2LLM gives your AI eyes and hands in a real browser: structured perception (SiFR) plus action.
Turn any LLM multimodal; generate images, voices, videos, 3D models, music, and more.
Enable language models to perform advanced AI-powered web scraping with enterprise-grade reliabili…
Related MCP Servers
- AlicenseNot gradedqualityNot gradedmaintenanceEnables LLMs to capture and analyze screenshots of your screen, windows, or regions with smart detection capabilities. Features natural language queries, automatic window targeting, and text enhancement for UI debugging and visual inspection.2
- AlicenseAqualityDmaintenanceEnables AI agents to analyze images, extract text, compare images, and analyze video through any OpenAI-compatible vision model.411920MIT
- AlicenseBqualityBmaintenanceMCP server providing 26 visual tools for text-only LLMs, enabling description, coordinate location, OCR, annotation, cropping/zooming, anomaly scanning, and computer control with switchable VLM backends.279MIT
- AlicenseNot gradedqualityCmaintenanceEnables text-only LLMs to understand images by converting them into text descriptions, supporting multiple vision backends like cloud APIs, local models, and OCR engines.1MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/IDEA-Research/DINO-X-MCP'
If you have feedback or need assistance with the MCP directory API, please join our Discord server