Skip to main content
Glama
WOODSEE-DIGI

Qwen3 MCP Server

by WOODSEE-DIGI
README.md
# Qwen3 MCP Server

A Model Context Protocol (MCP) server ecosystem providing access to multiple AI models optimized for different tasks: code generation, vision analysis, and complex reasoning.

## šŸš€ Quick Start

```bash
# Automated setup
./setup.sh

# Start default server
python src/main.py

# Or use ephemeral model switching
ask-qwen3 "Write a Python function"    # Code generation
ask-vision "Analyze this image"        # Visual analysis  
ask-ministral "Solve this equation"     # Complex reasoning
```

## šŸ“š Documentation

### Essential Guides
- **[Setup Guide](docs/SETUP.md)** - Complete installation and configuration
- **[Usage Guide](docs/USAGE.md)** - Workflows, examples, and best practices  
- **[Models Reference](docs/MODELS.md)** - Model capabilities and configurations
- **[Agent Guide](AGENTS.md)** - Warp agent integration guidance

### Quick Navigation
- šŸ—ļø **Getting Started**: [Setup Guide](docs/SETUP.md) → [Usage Guide](docs/USAGE.md)
- šŸ¤– **Model Selection**: See [Models Reference](docs/MODELS.md#model-selection-guide)
- šŸ”§ **Troubleshooting**: Check [Setup Guide](docs/SETUP.md#troubleshooting) or [Usage Guide](docs/USAGE.md#troubleshooting-usage-issues)
- šŸŽÆ **Specific Tasks**: Browse [Usage Guide](docs/USAGE.md#workflows-and-use-cases)

## 🌟 Features

### Multi-Model Ecosystem
- **Qwen3-Coder-Next**: Code generation, debugging, technical writing
- **Qwen3-VL-8B**: Image analysis, UI review, document OCR
- **Qwen3-30B**: Complex reasoning with thinking mode
- **Ministral-3-14B**: Mathematical reasoning and logical analysis

### Flexible Hosting
- **Ollama**: Local model serving (recommended)
- **HTTP API**: Remote model endpoints 
- **Transformers**: Direct model loading
- **Ephemeral Switching**: Dynamic model selection

### Developer Experience
- **MCP Compliance**: Full Model Context Protocol support
- **Shell Integration**: Quick aliases and commands
- **Warp Integration**: Native Warp agent support
- **Multi-Transport**: stdio and HTTP transports
- **Thinking Mode**: Detailed reasoning visualization

## šŸŽÆ Use Cases

| Task | Recommended Model | Command |
|------|-------------------|--------|
| Code Review | Qwen3-Coder | `ask-qwen3 "Review this code"` |
| UI Analysis | Qwen3-Vision | `ask-vision "Analyze this screenshot"` |
| Math Problems | Ministral | `ask-ministral "Solve step-by-step"` |
| System Design | Qwen3-30B | `python src/main.py --enable-thinking` |
| Document OCR | Qwen3-Vision | `ask-vision "Extract text from image"` |
| Algorithm Design | Qwen3-Coder | `ask-qwen3 "Implement data structure"` |

## ⚔ Quick Commands

### Model Switching
```bash
mcp-qwen3     # Code-focused development
mcp-vision    # Visual analysis tasks
mcp-ministral # Reasoning and mathematics
mcp-all       # Enable all models
mcp-clean     # Reset to clean state
```

### One-Shot Tasks
```bash
ask-qwen3 "Write a REST API endpoint"
ask-vision "What's wrong with this UI?"
ask-ministral "Prove this theorem"
```

### Server Management
```bash
# Start with specific model
python src/main.py --model-method ollama --ollama-model qwen3:30b-a3b

# Start with HTTP endpoint
python src/main.py --model-method http --http-model qwen/qwen3-coder-next

# Enable debug logging
python src/main.py --log-level DEBUG
```

## šŸ”§ System Requirements

- **Python**: 3.10+ (3.12+ recommended)
- **Memory**: 16GB+ RAM (32GB+ for 30B model)
- **Network**: Access to HTTP endpoints or Ollama service
- **OS**: macOS, Linux, Windows
- **Optional**: CUDA-compatible GPU for Transformers method

## 🚦 Health Check

```bash
# Check system status
mcp-list

# Test specific model
ask-ministral "Hello, are you working?"

# Verify endpoints
curl -s http://localhost:1234/v1/models
```

## šŸ“ Project Structure

```
qwen3-mcp-server/
ā”œā”€ā”€ docs/                  # šŸ“š Comprehensive documentation
│   ā”œā”€ā”€ SETUP.md          # Installation and configuration
│   ā”œā”€ā”€ USAGE.md          # Usage patterns and examples
│   └── MODELS.md         # Model reference and capabilities
ā”œā”€ā”€ src/                   # šŸ”§ Core implementation
│   ā”œā”€ā”€ main.py           # Entry point and CLI
│   ā”œā”€ā”€ server.py         # MCP server implementation  
│   ā”œā”€ā”€ model_interface.py # Model hosting abstractions
│   └── config.py         # Configuration management
ā”œā”€ā”€ config/                # āš™ļø Model configurations
│   ā”œā”€ā”€ qwen3-coder-http.json
│   ā”œā”€ā”€ qwen3-vl-8b-http.json
│   └── ministral-3-14b-reasoning-http.json
ā”œā”€ā”€ scripts/               # šŸ¤– Automation scripts
│   └── switch-model.sh   # Model switching logic
ā”œā”€ā”€ AGENTS.md             # šŸ¤– Warp agent guidance
ā”œā”€ā”€ setup.sh              # šŸš€ Automated setup
└── requirements.txt      # šŸ“¦ Python dependencies


## šŸ“„ License

MIT License - see [LICENSE](LICENSE) file for details.

## šŸ™ Acknowledgments

- [Model Context Protocol](https://modelcontextprotocol.io/) by Anthropic
- [Qwen Team](https://github.com/QwenLM) for the Qwen3 models  
- [Ollama](https://ollama.ai/) for local model hosting
- [Mistral AI](https://mistral.ai/) for the Ministral reasoning model