Skip to main content
Glama
baobab-tech

Text Classification MCP Server (Model2Vec)

by baobab-tech
README.md
# Text Classification MCP Server (Model2Vec)

**A powerful Model Context Protocol (MCP) server that provides comprehensive text classification tools using fast static embeddings from Model2Vec (Minish Lab).**

## ๐Ÿ› ๏ธ Complete MCP Tools & Resources

This server provides **6 essential tools**, **2 resources**, and **1 prompt template** for text classification:

### ๐Ÿท๏ธ Classification Tools
- **`classify_text`** - Classify single text with confidence scores  
- **`batch_classify`** - Classify multiple texts simultaneously

### ๐Ÿ“ Category Management Tools  
- **`add_custom_category`** - Add individual custom categories
- **`batch_add_custom_categories`** - Add multiple categories at once
- **`list_categories`** - View all available categories  
- **`remove_categories`** - Remove unwanted categories

### ๐Ÿ“Š Resources
- **`categories://list`** - Access category list programmatically
- **`model://info`** - Get model and system information

### ๐Ÿ’ฌ Prompt Templates
- **`classification_prompt`** - Ready-to-use classification prompt template

## ๐Ÿš€ Key Features

- **Zero-install**: Just `uv run` โ€” dependencies are declared inline (PEP 723)
- **Multiple Transports**: Supports stdio (local), HTTP/SSE, and Streamable HTTP
- **Fast Classification**: Uses efficient static embeddings from Model2Vec
- **10 Default Categories**: Technology, business, health, sports, entertainment, politics, science, education, travel, food
- **Custom Categories**: Add your own categories with descriptions
- **Batch Processing**: Classify multiple texts at once
- **Resource Endpoints**: Access category lists and model information
- **Prompt Templates**: Built-in prompts for classification tasks

## ๐Ÿ“‹ Installation

### Prerequisites
- Python 3.10+
- [`uv`](https://docs.astral.sh/uv/) package manager

### Quick Setup
No separate install step needed โ€” dependencies are declared inline in the script (PEP 723) and resolved automatically by `uv`.

## ๐Ÿƒโ€โ™‚๏ธ Running the Server

### Stdio Transport (Default)
```bash
uv run text_classifier_server.py
```

### HTTP/SSE Transport
```bash
# SSE on default port 8000
uv run text_classifier_server.py --http

# SSE on custom port
uv run text_classifier_server.py --http 9000
```

### Streamable HTTP Transport
```bash
uv run text_classifier_server.py --streamable-http
```

## ๐Ÿ”ง Configuration

### For Claude Desktop

#### Stdio Transport (Local)
Add to `~/Library/Application Support/Claude/claude_desktop_config.json`:
```json
{
  "mcpServers": {
    "text-classifier": {
      "command": "uv",
      "args": ["run", "/path/to/text_classifier_server.py"]
    }
  }
}
```

#### HTTP Transport (Remote)
Start the server with `uv run text_classifier_server.py --http`, then add:
```json
{
  "mcpServers": {
    "text-classifier": {
      "url": "http://localhost:8000/sse"
    }
  }
}
```

### For Claude Code
```bash
claude mcp add text-classifier -- uv run /Users/olivier/DEV/mcp-text-classifier/text_classifier_server.py
```

## ๐Ÿ› ๏ธ Available Tools

### classify_text
Classify a single text into predefined categories with confidence scores.

**Parameters:**
- `text` (string): The text to classify
- `top_k` (int, optional): Number of top categories to return (default: 3)

**Returns:** JSON with predictions, confidence scores, and category descriptions

**Example:**
```python
classify_text("Apple announced new AI features", top_k=3)
```

### batch_classify
Classify multiple texts simultaneously for efficient processing.

**Parameters:**
- `texts` (list): List of texts to classify
- `top_k` (int, optional): Number of top categories per text (default: 1)

**Returns:** JSON with batch classification results

**Example:**
```python
batch_classify(["Tech news", "Sports update", "Business report"], top_k=2)
```

### add_custom_category
Add a new custom category for classification.

**Parameters:**
- `category_name` (string): Name of the new category
- `description` (string): Description to generate the category embedding

**Returns:** JSON with operation result

**Example:**
```python
add_custom_category("automotive", "Cars, vehicles, transportation, automotive industry")
```

### batch_add_custom_categories
Add multiple custom categories in a single operation for efficiency.

**Parameters:**
- `categories_data` (list): List of dictionaries with 'name' and 'description' keys

**Returns:** JSON with batch operation results

**Example:**
```python
batch_add_custom_categories([
    {"name": "automotive", "description": "Cars, vehicles, transportation"},
    {"name": "music", "description": "Music, songs, artists, albums, concerts"}
])
```

### list_categories
List all available categories and their descriptions.

**Parameters:** None

**Returns:** JSON with all categories and their descriptions

### remove_categories
Remove one or multiple categories from the classification system.

**Parameters:**
- `category_names` (list): List of category names to remove

**Returns:** JSON with removal results for each category

**Example:**
```python
remove_categories(["automotive", "custom_category"])
```

## ๐Ÿ“š Available Resources

- **`categories://list`**: Get list of available categories with metadata
- **`model://info`**: Get information about the loaded Model2Vec model and system status

## ๐Ÿ’ฌ Available Prompts

- **`classification_prompt`**: Template for text classification tasks with context and instructions

**Parameters:**
- `text` (string): The text to classify

**Returns:** Formatted prompt for classification with available categories listed

## ๐Ÿงช Testing

### Test with MCP Inspector
```bash
npx @modelcontextprotocol/inspector uv run text_classifier_server.py
```

## ๐Ÿ” Troubleshooting

### Model download fails
```bash
# Manual model download
uv run python -c "from model2vec import StaticModel; StaticModel.from_pretrained('minishlab/potion-base-8M')"
```

## ๐Ÿ“– Technical Details

- **Model**: `minishlab/potion-base-8M` from Model2Vec
- **Similarity**: Cosine similarity between text and category embeddings
- **Performance**: ~30MB model, fast inference with static embeddings
- **Protocol**: MCP specification 2024-11-05
- **Transports**: stdio, HTTP+SSE, Streamable HTTP

## ๐Ÿค Contributing

1. Fork the repository
2. Create a feature branch
3. Add tests for new functionality
4. Submit a pull request

## ๐Ÿ“„ License

MIT License - see LICENSE file for details.

## ๐Ÿ™ Acknowledgments

- [Model2Vec](https://github.com/MinishLab/model2vec) by Minish Lab for fast static embeddings
- [Anthropic](https://anthropic.com) for the Model Context Protocol specification
- [FastMCP](https://github.com/jlowin/fastmcp) for the excellent Python MCP framework

---

**Need help?** Check the troubleshooting section or open an issue in the repository.

TDQS

A3.6/5.0

Scored across 6 tools

Disambiguation5/5

Each tool has a distinct purpose: adding categories (single/batch), classifying text (single/batch), listing, and removing categories. No overlap or ambiguity.

Naming Consistency5/5

All tools use snake_case and follow a consistent verb_noun pattern (e.g., add_custom_category, batch_classify, list_categories). No mixing of styles.

Tool Count5/5

Six tools is well-scoped for a text classification server: category CRUD (add/batch add, list, remove) and classification (single/batch). Neither too few nor too many.

Completeness4/5

Covers core functionality: category management and classification. Missing an update category tool is a minor gap, but agents can work around it by removing and re-adding.

Maintenance

ActivityInactive
ResponsivenessNo issues