Skip to main content
Glama

Venice Inference Organizer (VIO)

Live Venice.ai model catalog, five quality/cost lanes, HTTP + MCP resolve service, and an optional keys dashboard.

VIO is a directory, not a proxy. Apps and agents still call https://api.venice.ai/api/v1. On startup they ask VIO which model IDs currently fit their slots, then use those IDs against Venice.

Proposal: PROPOSAL.md

What it does

  • Refresh Venice GET /models and GET /models/traits on a schedule

  • Tag every model inside its own modality into extra_low … extra_high

  • Elect a primary plus fallbacks per type × lane

  • Resolve one slot or a whole app bundle over HTTP or MCP

  • Let users pin or retarget lanes without hard-coding model IDs

  • Show a keys dashboard: each key, last 6 characters, and what it is used for

  • Optionally join Venice Key Manager if that tool is installed on the same host

  • Let bound agents update or cycle their own inference keys

Related MCP server: AI Model Advisor MCP Server

Lanes (per modality)

Lane

Intent

extra_low

Cheapest viable

low

Fast daily driver

medium

Balanced default

high

Specialist (code, reasoning, long context)

extra_high

Best allowed after privacy + budget

Cheap text is never ranked against cheap video.

Quick start

git clone https://github.com/maximusmaximus/venice-inference-organizer.git
cd venice-inference-organizer
cp .env.example .env
# set VENICE_API_KEY (inference is enough for catalog refresh)
pip install -e .
vio serve

Default bind is loopback (127.0.0.1:8787). Do not publish this port to the internet without a service token.

curl -H "Authorization: Bearer $VIO_SERVICE_TOKEN" \
  "http://127.0.0.1:8787/v1/resolve?app=example-agent&slot=planner&lane=medium"

MCP

VIO exposes catalog, resolve, prefs, and key-binding tools. It does not reimplement the official Venice MCP (@veniceai/mcp-server). Pair that package if you want chat/image/video tools.

{
  "mcpServers": {
    "vio": {
      "command": "python",
      "args": ["-m", "vio.mcp_server"],
      "env": {
        "VIO_SERVICE_TOKEN": "change-me",
        "VENICE_API_KEY": "your-inference-key"
      }
    }
  }
}

Optional Key Manager

If Venice Key Manager is running on this host (default http://127.0.0.1:8660), the VIO dashboard joins its key inventory and categories with VIO consumer bindings. If it is not installed, VIO lists keys from Venice when VENICE_ADMIN_KEY is set.

Security

  • Never commit .env, live tokens, or full API secrets

  • Dashboard and GET /v1/keys show last6Chars only

  • Create/cycle returns a secret once; VIO stores key_id + last 6

  • Leaf agents receive INFERENCE keys only

License

MIT. See LICENSE.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Discovers LLM models in real time from cloud providers and local Ollama instances, returning compatibility profiles and live pricing so AI agents can route tasks to the cheapest viable model without breaking tool calls or context clipping.
    10
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables real-time access to LLM pricing, benchmarks, deprecation alerts, and cost optimization for over 30 models across 8 providers, allowing AI agents to make cost-effective model selections.
    -