Skip to main content
Glama
ymeng98

ddddocr Smithery MCP Server

by ymeng98
README.md
# ddddocr Smithery MCP Server

A Model Context Protocol (MCP) server for ddddocr that can be deployed on [Smithery](https://smithery.ai/), providing OCR and CAPTCHA recognition capabilities to AI agents.

## Features

- **OCR Recognition**: Extract text from images with high accuracy
- **Text Detection**: Identify and locate text regions in images  
- **Slide CAPTCHA Solving**: Match sliding puzzle pieces and find positions
- **Color Filtering**: Process images with specific color filters
- **Probability Output**: Get confidence scores for OCR results

## Quick Start

### Deploy on Smithery

1. Fork this repository to your GitHub account
2. Visit [Smithery](https://smithery.ai/) and connect your GitHub account
3. Deploy from your forked repository
4. Use the provided Smithery URL in your Claude Desktop configuration

### Local Development

```bash
# Install dependencies
npm install

# Build the project
npm run build

# Run in development mode
npm run dev

# Run tests
npm test
```

## Usage

This MCP server provides the following tools:

### `ocr_recognize`
Extract text content from images.

**Parameters:**
- `image` (required): Base64 encoded image data
- `probability` (optional): Return confidence scores
- `charset_range` (optional): Limit character set (e.g., "0123456789")
- `color_filter` (optional): Apply color filters
- `png_fix` (optional): Fix transparent PNG images

### `text_detection`
Detect text regions and bounding boxes in images.

**Parameters:**
- `image` (required): Base64 encoded image data

### `slide_match`
Match sliding CAPTCHA pieces to find correct positions.

**Parameters:**
- `target_image` (required): Base64 encoded puzzle piece
- `background_image` (required): Base64 encoded background with gap
- `simple_target` (optional): Whether target has transparency

### `slide_comparison`
Compare images to find sliding distance for CAPTCHA solving.

**Parameters:**
- `target_image` (required): Base64 encoded image with gap
- `background_image` (required): Base64 encoded complete image

## Configuration

Add this server to your Claude Desktop configuration:

```json
{
  "mcpServers": {
    "ddddocr": {
      "command": "npx",
      "args": ["-y", "@smithery/ddddocr-mcp@latest"]
    }
  }
}
```

Or if deployed on Smithery:

```json
{
  "mcpServers": {
    "ddddocr": {
      "command": "npx",
      "args": ["-y", "@smithery/cli", "run", "your-deployment-url"]
    }
  }
}
```

## Architecture

This server acts as a bridge between MCP clients and the ddddocr service:

1. **MCP Layer**: Handles protocol communication with AI agents
2. **Service Layer**: Manages ddddocr process lifecycle
3. **API Layer**: Communicates with ddddocr HTTP endpoints
4. **Processing Layer**: Handles image processing and result formatting

## Requirements

- Node.js 18+
- ddddocr executable (automatically downloaded in Docker)
- Sufficient memory for image processing (recommend 512MB+)

## Security

- Uses non-root user in Docker container
- Validates all input parameters
- Implements proper error handling
- No persistent storage of user images

## License

MIT - See LICENSE file for details

## Contributing

1. Fork the repository
2. Create a feature branch
3. Make your changes
4. Add tests if applicable
5. Submit a pull request

## Support

For issues and questions:
- Check the [GitHub Issues](https://github.com/ymeng98/ddddocr-smithery-mcp/issues)
- Review [ddddocr documentation](https://github.com/86maid/ddddocr)
- Visit [Smithery Documentation](https://smithery.ai/docs)

TDQS

B3.1/5.0

Scored across 4 tools

Disambiguation2/5

ocr_recognize and text_detection are fairly distinguishable, but slide_match and slide_comparison both target sliding captcha solving with descriptions that seem to describe the same task. The boundary between matching pieces and comparing images for a sliding distance is unclear.

Naming Consistency3/5

All names use snake_case and have recognizable prefixes, but verb/noun forms are inconsistent: ocr_recognize uses a verb while text_detection and slide_comparison use nouns. slide_match and slide_comparison follow a common slide_ pattern, yet the overall set is not uniformly structured.

Tool Count4/5

Four tools is a reasonable size for an OCR and captcha-focused server. However, slide_match and slide_comparison appear redundant, so the count is acceptable but not perfectly lean.

Completeness4/5

The core OCR recognition, text detection, and sliding captcha workflows are covered. Minor gaps include lack of a dedicated captcha classification tool or preprocessing utilities, but they are likely workarounds given the stated scope.

Maintenance

ActivityInactive
ResponsivenessNo issues