ddddocr Smithery MCP Server
# ddddocr Smithery MCP Server
A Model Context Protocol (MCP) server for ddddocr that can be deployed on [Smithery](https://smithery.ai/), providing OCR and CAPTCHA recognition capabilities to AI agents.
## Features
- **OCR Recognition**: Extract text from images with high accuracy
- **Text Detection**: Identify and locate text regions in images
- **Slide CAPTCHA Solving**: Match sliding puzzle pieces and find positions
- **Color Filtering**: Process images with specific color filters
- **Probability Output**: Get confidence scores for OCR results
## Quick Start
### Deploy on Smithery
1. Fork this repository to your GitHub account
2. Visit [Smithery](https://smithery.ai/) and connect your GitHub account
3. Deploy from your forked repository
4. Use the provided Smithery URL in your Claude Desktop configuration
### Local Development
```bash
# Install dependencies
npm install
# Build the project
npm run build
# Run in development mode
npm run dev
# Run tests
npm test
```
## Usage
This MCP server provides the following tools:
### `ocr_recognize`
Extract text content from images.
**Parameters:**
- `image` (required): Base64 encoded image data
- `probability` (optional): Return confidence scores
- `charset_range` (optional): Limit character set (e.g., "0123456789")
- `color_filter` (optional): Apply color filters
- `png_fix` (optional): Fix transparent PNG images
### `text_detection`
Detect text regions and bounding boxes in images.
**Parameters:**
- `image` (required): Base64 encoded image data
### `slide_match`
Match sliding CAPTCHA pieces to find correct positions.
**Parameters:**
- `target_image` (required): Base64 encoded puzzle piece
- `background_image` (required): Base64 encoded background with gap
- `simple_target` (optional): Whether target has transparency
### `slide_comparison`
Compare images to find sliding distance for CAPTCHA solving.
**Parameters:**
- `target_image` (required): Base64 encoded image with gap
- `background_image` (required): Base64 encoded complete image
## Configuration
Add this server to your Claude Desktop configuration:
```json
{
"mcpServers": {
"ddddocr": {
"command": "npx",
"args": ["-y", "@smithery/ddddocr-mcp@latest"]
}
}
}
```
Or if deployed on Smithery:
```json
{
"mcpServers": {
"ddddocr": {
"command": "npx",
"args": ["-y", "@smithery/cli", "run", "your-deployment-url"]
}
}
}
```
## Architecture
This server acts as a bridge between MCP clients and the ddddocr service:
1. **MCP Layer**: Handles protocol communication with AI agents
2. **Service Layer**: Manages ddddocr process lifecycle
3. **API Layer**: Communicates with ddddocr HTTP endpoints
4. **Processing Layer**: Handles image processing and result formatting
## Requirements
- Node.js 18+
- ddddocr executable (automatically downloaded in Docker)
- Sufficient memory for image processing (recommend 512MB+)
## Security
- Uses non-root user in Docker container
- Validates all input parameters
- Implements proper error handling
- No persistent storage of user images
## License
MIT - See LICENSE file for details
## Contributing
1. Fork the repository
2. Create a feature branch
3. Make your changes
4. Add tests if applicable
5. Submit a pull request
## Support
For issues and questions:
- Check the [GitHub Issues](https://github.com/ymeng98/ddddocr-smithery-mcp/issues)
- Review [ddddocr documentation](https://github.com/86maid/ddddocr)
- Visit [Smithery Documentation](https://smithery.ai/docs)TDQS
Scored across 4 tools
ocr_recognize and text_detection are fairly distinguishable, but slide_match and slide_comparison both target sliding captcha solving with descriptions that seem to describe the same task. The boundary between matching pieces and comparing images for a sliding distance is unclear.
All names use snake_case and have recognizable prefixes, but verb/noun forms are inconsistent: ocr_recognize uses a verb while text_detection and slide_comparison use nouns. slide_match and slide_comparison follow a common slide_ pattern, yet the overall set is not uniformly structured.
Four tools is a reasonable size for an OCR and captcha-focused server. However, slide_match and slide_comparison appear redundant, so the count is acceptable but not perfectly lean.
The core OCR recognition, text detection, and sliding captcha workflows are covered. Minor gaps include lack of a dedicated captcha classification tool or preprocessing utilities, but they are likely workarounds given the stated scope.