Skip to main content
Glama
IDEA-Research

DINO-X Image Detection MCP Server

DINO-X MCP Server

License npm version npm downloads PRs Welcome MCP Badge GitHub stars

English | 中文

DINO-X Official MCP Server — powered by the DINO-X and Grounding DINO models — brings fine-grained object detection and image understanding to your multimodal applications.

Why DINO-X MCP?

With DINO-X MCP, you can:

  • Fine-Grained Understanding: Full image detection, object detection, and region-level descriptions.

  • Structured Outputs: Get object categories, counts, locations, and attributes for VQA and multi-step reasoning tasks.

  • Composable: Works seamlessly with other MCP servers to build end-to-end visual agents or automation pipelines.

Related MCP server: Vision MCP Server

Transport Modes

DINO-X MCP supports two transport modes:

Feature

STDIO (default)

Streamable HTTP

Runtime

Local

Local or Cloud

Transport

Standard I/O

HTTP (streaming responses)

Input source

file:// and https://

https:// only

Visualization

Supported (saves annotated images locally)

Not supported (for now)

Quick Start

1. Prepare an MCP client

Any MCP-compatible client works, e.g.:

2. Get your API key

Apply on the DINO-X platform: Request API Key (new users get free quota).

3. Configure MCP

Add to your MCP client config and replace with your API key:

{
  "mcpServers": {
    "dinox-mcp": {
      "url": "https://mcp.deepdataspace.com/mcp?key=your-api-key"
    }
  }
}

Option B: Use the NPM package locally (STDIO)

Install Node.js first

  • Download the installer from nodejs.org

  • Or use command:

# macOS / Linux
curl -o- https://raw.githubusercontent.com/nvm-sh/nvm/v0.40.1/install.sh | bash
# or
wget -qO- https://raw.githubusercontent.com/nvm-sh/nvm/v0.40.1/install.sh | bash

# load nvm into current shell (choose the one you use)
source ~/.bashrc || true
source ~/.zshrc  || true

# install and use LTS Node.js
nvm install --lts
nvm use --lts

# Windows (one of the following)
winget install OpenJS.NodeJS.LTS
# or with Chocolatey (in admin PowerShell)
iwr -useb https://raw.githubusercontent.com/chocolatey/chocolatey/master/chocolateyInstall/InstallChocolatey.ps1 | iex
choco install nodejs-lts -y

Configure your MCP client:

{
  "mcpServers": {
    "dinox-mcp": {
      "command": "npx",
      "args": ["-y", "@deepdataspace/dinox-mcp"],
      "env": {
        "DINOX_API_KEY": "your-api-key-here",
        "IMAGE_STORAGE_DIRECTORY": "/path/to/your/image/directory"
      }
    }
  }
}

Note: Replace your-api-key-here with your real key.

Option C: Run from source locally

Make sure Node.js is installed (see Option B), then:

# clone
git clone https://github.com/IDEA-Research/DINO-X-MCP.git
cd DINO-X-MCP

# install deps
npm install

# build
npm run build

Configure your MCP client:

{
  "mcpServers": {
    "dinox-mcp": {
      "command": "node",
      "args": ["/path/to/DINO-X-MCP/build/index.js"],
      "env": {
        "DINOX_API_KEY": "your-api-key-here",
        "IMAGE_STORAGE_DIRECTORY": "/path/to/your/image/directory"
      }
    }
  }
}

CLI Flags & Environment Variables

  • Common flags

    • --http: start in Streamable HTTP mode (otherwise STDIO by default)

    • --stdio: force STDIO mode

    • --dinox-api-key=...: set API key

    • --enable-client-key: allow API key via URL ?key= (Streamable HTTP only)

    • --port=8080: HTTP port (default 3020)

  • Environment variables

    • DINOX_API_KEY (required/conditionally required): DINO-X platform API key

    • IMAGE_STORAGE_DIRECTORY (optional, STDIO): directory to save annotated images

    • AUTH_TOKEN (optional, HTTP): if set, client must send Authorization: Bearer <token>

    Examples:

# STDIO (local)
node build/index.js --dinox-api-key=your-api-key

# Streamable HTTP (server provides a shared API key)
node build/index.js --http --dinox-api-key=your-api-key

# Streamable HTTP (custom port)
node build/index.js --http --dinox-api-key=your-api-key --port=8080

# Streamable HTTP (require client-provided API key via URL)
node build/index.js --http --enable-client-key

Client config when using ?key=:

{
  "mcpServers": {
    "dinox-mcp": {
      "url": "http://localhost:3020/mcp?key=your-api-key"
    }
  }
}

Using AUTH_TOKEN with a gateway that injects Authorization: Bearer <token>:

AUTH_TOKEN=my-token node build/index.js --http --enable-client-key

Client example with supergateway:

{
  "mcpServers": {
    "dinox-mcp": {
      "command": "npx",
      "args": [
        "-y",
        "supergateway",
        "--streamableHttp",
        "http://localhost:3020/mcp?key=your-api-key",
        "--oauth2Bearer",
        "my-token"
      ]
    }
  }
}

Tools

Capability

Tool ID

Transport

Input

Output

Full-scene object detection

detect-all-objects

STDIO / HTTP

Image URL

Category + bbox + (optional) captions

Text-prompted object detection

detect-objects-by-text

STDIO / HTTP

Image URL + English nouns (dot-separated for multiple, e.g., person.car)

Target object bbox + (optional) captions

Human pose estimation

detect-human-pose-keypoints

STDIO / HTTP

Image URL

17 keypoints + bbox + (optional) captions

Visualization

visualize-detection-result

STDIO only

Image URL + detection results array

Local path to annotated image

🎬 Use Cases

🎯 Scenario

📝 Input

✨ Output

Detection & Localization

💬 Prompt:Detect and visualize the fire areas in the forest 🖼️ Input Image:1-1

1-2

Object Counting

💬 Prompt:Please analyze thiswarehouse image, detectall the cardboard boxes,count the total number🖼️ Input Image:2-1

Feature Detection

💬 Prompt:Find all red carsin the image🖼️ Input Image:4-1

4-2

Attribute Reasoning

💬 Prompt:Find the tallest personin the image, describetheir clothing🖼️ Input Image:5-1

5-2

Full Scene Detection

💬 Prompt:Find the fruit withthe highest vitamin Ccontent in the image🖼️ Input Image:6-1

6-3Answer: Kiwi fruit (93mg/100g)

Pose Analysis

💬 Prompt:Please analyze whatyoga pose this is🖼️ Input Image:3-1

3-3

FAQ

  • Supported image sources?

    • STDIO: file:// and https://

    • Streamable HTTP: https:// only

  • Supported image formats?

    • jpg, jpeg, webp, png

Development & Debugging

Use watch mode to auto-rebuild during development:

npm run watch

Use MCP Inspector for debugging:

npm run inspector

License

Apache License 2.0

Available Tools

4 tools
detect-all-objectsB

Analyze an image to detect all identifiable objects, returning the category, count, coordinate positions and detailed descriptions for each object.

ParametersJSON Schema
NameRequiredDescriptionDefault
imageFileUriYesURI of the input image. Preferred for remote or local files. Must start with 'https://' or 'file://'.
includeDescriptionYesWhether to return a description of the objects detected in the image, but will take longer to process.

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions that including descriptions 'will take longer to process,' which adds some context about performance impact. However, it doesn't address other important behavioral aspects like error handling, rate limits, authentication requirements, or what happens with invalid inputs, which are significant gaps for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that efficiently conveys the tool's purpose, action, and output without any redundant information. It's appropriately sized and front-loaded, with every word serving a clear purpose in explaining what the tool does.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (image analysis with 2 parameters), no annotations, and no output schema, the description is partially complete. It clearly states what the tool does and what it returns, but lacks details on behavioral traits, error conditions, and output format specifics. The absence of an output schema means the description should ideally explain return values more thoroughly, which it doesn't do beyond listing output categories.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, so the schema already fully documents both parameters (imageFileUri and includeDescription). The description adds minimal value beyond the schema by briefly mentioning that includeDescription affects processing time, but doesn't provide additional syntax, format details, or usage context for the parameters. This meets the baseline expectation when schema coverage is high.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Analyze an image to detect all identifiable objects') and distinguishes it from siblings by emphasizing 'all identifiable objects' rather than specific types like human poses or text-based detection. It provides the verb+resource+scope combination that makes the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus the sibling tools (detect-human-pose-keypoints, detect-objects-by-text, visualize-detection-result). It doesn't mention alternatives, exclusions, or specific contexts where this tool is preferred over others, leaving the agent with no comparative usage information.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

detect-human-pose-keypointsC

Detects 17 keypoints for each person in an image, supporting body posture and movement analysis.

ParametersJSON Schema
NameRequiredDescriptionDefault
imageFileUriYesURI of the input image. Preferred for remote or local files. Must start with 'https://' or 'file://'.
includeDescriptionYesWhether to return a description of the objects detected in the image, but will take longer to process.

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden but offers minimal behavioral insight. It mentions the output (17 keypoints per person) and implies analysis use, but lacks details on performance (e.g., speed, accuracy), limitations (e.g., image quality requirements), or side effects. The description doesn't contradict annotations, but it's insufficient for a mutation-like detection tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the core purpose without unnecessary details. Every word contributes to understanding the tool's function, making it appropriately concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (detection with 2 parameters) and lack of annotations and output schema, the description is incomplete. It doesn't explain return values (e.g., keypoint coordinates, confidence scores), error handling, or prerequisites (e.g., image format support), leaving significant gaps for agent usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents both parameters. The description adds no parameter-specific information beyond what's in the schema, such as explaining how 'includeDescription' relates to keypoints or image processing. This meets the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: detecting 17 keypoints per person in an image for posture/movement analysis. It specifies the verb ('detects'), resource ('keypoints'), and scope ('each person in an image'), but doesn't explicitly differentiate from sibling tools like 'detect-all-objects' or 'detect-objects-by-text', which likely serve different detection purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention sibling tools or contexts where human pose detection is preferred over general object detection, leaving the agent to infer usage based on tool names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

detect-objects-by-textA

Analyze an image based on a text prompt to identify and count specific objects, and return detailed descriptions of the objects and their 2D coordinates.

ParametersJSON Schema
NameRequiredDescriptionDefault
imageFileUriYesURI of the input image. Preferred for remote or local files. Must start with 'https://' or 'file://'.
textPromptYesNouns of target objects (English only, avoid adjectives). Use periods to separate multiple categories (e.g., 'person.car.traffic light').
includeDescriptionYesWhether to return a description of the objects detected in the image, but will take longer to process.

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It mentions that including descriptions will 'take longer to process', which is useful behavioral context. However, it doesn't disclose other critical traits like potential rate limits, error conditions, authentication needs, or what happens with invalid inputs. For a tool with no annotation coverage, this leaves significant gaps in understanding its behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that efficiently conveys the core functionality without unnecessary words. It's front-loaded with the main action and outcome, making it easy to understand at a glance. Every part of the sentence contributes directly to explaining what the tool does.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that there's no output schema and no annotations, the description should ideally provide more context about return values and behavioral constraints. While it mentions what will be returned (descriptions and coordinates), it doesn't specify the format or structure of the output. For a tool with 3 required parameters and no structured output documentation, the description is adequate but leaves room for improvement in completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters thoroughly. The description doesn't add any meaningful semantic information beyond what's in the schema descriptions (e.g., it doesn't explain the format of returned coordinates or how object counts are calculated). Baseline score of 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('analyze an image based on a text prompt'), the resource ('image'), and the outcome ('identify and count specific objects, return detailed descriptions and 2D coordinates'). It distinguishes from sibling tools like 'detect-all-objects' by specifying text-based filtering and from 'detect-human-pose-keypoints' by focusing on object detection rather than pose analysis.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage through its purpose statement but doesn't explicitly state when to use this tool versus alternatives like 'detect-all-objects' or 'visualize-detection-result'. It mentions the text prompt requirement, which suggests this tool is for targeted object detection, but lacks clear guidance on scenarios where other tools might be more appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

visualize-detection-resultB

Visualize detection results by drawing bounding boxes and labels on the original image. Images are saved to the directory specified by IMAGE_STORAGE_DIRECTORY environment variable.

ParametersJSON Schema
NameRequiredDescriptionDefault
imageFileUriYesURI of the input image. Preferred for remote or local files. Must start with 'https://' or 'file://'.
detectionsYesArray of detection results with name and bbox information.
fontSizeNoFont size for labels (default: 24)
boxThicknessNoThickness of bounding box lines (default: 4)
showLabelsNoWhether to show category labels (default: true)

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden but provides minimal behavioral information. It mentions that 'Images are saved to the directory specified by IMAGE_STORAGE_DIRECTORY environment variable' which is useful context about output location, but doesn't disclose important behavioral traits like whether this is a read-only operation, what happens if the directory doesn't exist, what file format is used, or whether the original image is modified versus a copy being created.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately concise with two clear sentences. The first sentence states the core purpose, and the second provides important behavioral context about where results are saved. There's no wasted text, though it could be slightly more structured with clearer separation of concerns.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 parameters, no annotations, and no output schema, the description is insufficient. It doesn't explain what the tool returns (success/failure indicators, output file paths, error conditions), doesn't mention dependencies on sibling tools, and provides minimal behavioral context. The description leaves too many important questions unanswered for effective agent use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema description coverage, the baseline is 3. The description adds no additional parameter information beyond what's already documented in the schema. It doesn't explain relationships between parameters, provide examples of valid detections arrays, or offer guidance on appropriate values for fontSize and boxThickness.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Visualize detection results by drawing bounding boxes and labels on the original image') and distinguishes it from sibling detection tools by focusing on visualization rather than detection. It specifies the exact resource being modified (the original image with annotations added).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (like needing detection results first), doesn't explain relationships to sibling detection tools, and offers no context about appropriate use cases or limitations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A3.6/5.0
Disambiguation5/5

Each tool has a clearly distinct purpose: detect-all-objects performs general object detection, detect-human-pose-keypoints focuses on human pose analysis, detect-objects-by-text enables text-guided detection, and visualize-detection-result handles visualization. There is no overlap in functionality, making tool selection straightforward for an agent.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern with hyphens (e.g., detect-all-objects, detect-human-pose-keypoints). The naming is predictable and readable throughout, with no deviations in style or convention.

Tool Count5/5

With 4 tools, the server is well-scoped for image detection tasks. Each tool earns its place by covering distinct aspects: detection (general, pose-specific, text-guided) and visualization, avoiding bloat while providing essential functionality.

Completeness4/5

The tool set covers core detection workflows (general, pose, text-guided) and visualization, with no obvious dead ends. A minor gap exists in lacking tools for modifying or deleting detection results, but agents can work around this by re-running detections or handling data externally.

Maintenance

ActivityNo data
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/IDEA-Research/DINO-X-MCP'

If you have feedback or need assistance with the MCP directory API, please join our Discord server