vio
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@vioresolve a medium lane model for my planner agent"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Venice Inference Organizer (VIO)
Live Venice.ai model catalog, five quality/cost lanes, HTTP + MCP resolve service, and an optional keys dashboard.
VIO is a directory, not a proxy. Apps and agents still call https://api.venice.ai/api/v1. On startup they ask VIO which model IDs currently fit their slots, then use those IDs against Venice.
Proposal: PROPOSAL.md
What it does
Refresh Venice
GET /modelsandGET /models/traitson a scheduleTag every model inside its own modality into
extra_low…extra_highElect a primary plus fallbacks per type × lane
Resolve one slot or a whole app bundle over HTTP or MCP
Let users pin or retarget lanes without hard-coding model IDs
Show a keys dashboard: each key, last 6 characters, and what it is used for
Optionally join Venice Key Manager if that tool is installed on the same host
Let bound agents update or cycle their own inference keys
Related MCP server: AI Model Advisor MCP Server
Lanes (per modality)
Lane | Intent |
| Cheapest viable |
| Fast daily driver |
| Balanced default |
| Specialist (code, reasoning, long context) |
| Best allowed after privacy + budget |
Cheap text is never ranked against cheap video.
Quick start
git clone https://github.com/maximusmaximus/venice-inference-organizer.git
cd venice-inference-organizer
cp .env.example .env
# set VENICE_API_KEY (inference is enough for catalog refresh)
pip install -e .
vio serveDefault bind is loopback (127.0.0.1:8787). Do not publish this port to the internet without a service token.
curl -H "Authorization: Bearer $VIO_SERVICE_TOKEN" \
"http://127.0.0.1:8787/v1/resolve?app=example-agent&slot=planner&lane=medium"MCP
VIO exposes catalog, resolve, prefs, and key-binding tools. It does not reimplement the official Venice MCP (@veniceai/mcp-server). Pair that package if you want chat/image/video tools.
{
"mcpServers": {
"vio": {
"command": "python",
"args": ["-m", "vio.mcp_server"],
"env": {
"VIO_SERVICE_TOKEN": "change-me",
"VENICE_API_KEY": "your-inference-key"
}
}
}
}Optional Key Manager
If Venice Key Manager is running on this host (default http://127.0.0.1:8660), the VIO dashboard joins its key inventory and categories with VIO consumer bindings. If it is not installed, VIO lists keys from Venice when VENICE_ADMIN_KEY is set.
Security
Never commit
.env, live tokens, or full API secretsDashboard and
GET /v1/keysshowlast6CharsonlyCreate/cycle returns a secret once; VIO stores
key_id+ last 6Leaf agents receive
INFERENCEkeys only
License
MIT. See LICENSE.
This server cannot be deployed
Maintenance
Related MCP Connectors
AI model routing on your own vendor keys: pick the best model per prompt, or route and run it.
Connect MCP clients to 2,000+ AI models without managing provider API keys.
Durable, user-controlled goals and governed plans for AI agents.
Cost-optimized LLM model routing recommendations for autonomous AI agents
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceDiscovers LLM models in real time from cloud providers and local Ollama instances, returning compatibility profiles and live pricing so AI agents can route tasks to the cheapest viable model without breaking tool calls or context clipping.10MIT
- AlicenseAqualityDmaintenanceEnables AI agents to discover, compare, and select the best AI models across multiple providers based on pricing, performance, and capabilities, with real-time cost estimation and benchmarking.9671 npm1MIT
- FlicenseNot gradedqualityCmaintenanceEnables real-time access to LLM pricing, benchmarks, deprecation alerts, and cost optimization for over 30 models across 8 providers, allowing AI agents to make cost-effective model selections.-
- FlicenseNot gradedqualityBmaintenanceEnables AI agents to discover and query live model information from OpenRouter, HuggingFace, and Ollama Cloud, ensuring recommendations are based on current data.-