Skip to main content
Glama
anteew
by anteew

RanchHand — OpenAI-compatible MCP Server (Architecture)

RanchHand is a minimal MCP server that fronts an OpenAI-style API. It works great with Ollama's OpenAI-compatible endpoints (http://localhost:11434/v1) and should work with other OpenAI-compatible backends.

Features

  • Tools:

    • openai_models_list → GET /v1/models

    • openai_chat_completions → POST /v1/chat/completions

    • openai_embeddings_create → POST /v1/embeddings

    • Optional HTTP ingest on localhost:41414 (bind 127.0.0.1):

      • POST /ingest/slack (index: chunk + embed + upsert in in-memory store)

      • POST /query (kNN query with embeddings)

      • GET /profiles | POST /profiles (role defaults: embed, summarizers, reranker, chunking)

      • POST /answer (retrieve + generate answer with bracketed citations)

  • Config via env:

    • OAI_BASE (default http://localhost:11434/v1)

    • OAI_API_KEY (optional; some backends ignore it, Ollama allows any value)

    • OAI_DEFAULT_MODEL (fallback model name, e.g. llama3:latest)

    • OAI_TIMEOUT_MS (optional request timeout)

Development

Linting

This project uses ESLint to maintain code quality and consistency.

# Run the linter to check for issues
npm run lint

# Automatically fix linting issues where possible
npm run lint:fix

The linting rules enforce:

  • Consistent code style (single quotes, semicolons, 2-space indentation)

  • Error prevention (no unused variables, no undefined variables)

  • Modern JavaScript practices (const/let instead of var, arrow functions)

CI will automatically run linting checks on all pull requests.

Testing

This repo uses Vitest for unit tests. External network calls are mocked, so tests run deterministically without Ollama or internet access.

Commands:

# Run tests once
npm test

# TDD: watch mode
npm run test:watch

# With coverage report
npm run test:coverage

Coverage thresholds are configured in vitest.config.mjs (initial targets):

  • Lines/Statements ≥ 60%

  • Functions ≥ 55%

  • Branches ≥ 50%

These thresholds indicate the minimum proportion of code exercised by tests. They are a guardrail, not a guarantee of correctness. We can raise them as the test suite grows.

Notes:

  • Tests live in tests/**/*.test.js

  • Use vi.spyOn/vi.mock to stub fetch and other external calls

  • For CI stability, avoid real network calls in tests

Run (standalone)

# Example with Ollama running locally
export OAI_BASE=http://localhost:11434/v1
export OAI_DEFAULT_MODEL=llama3:latest
node server.mjs

HTTP Ingest Service

node http.mjs
# Binds to 127.0.0.1:41414
# Shared secret is created at ~/.threadweaverinc/auth/shared_secret.txt on first run

Example request:

SECRET=$(cat ~/.threadweaverinc/auth/shared_secret.txt)
curl -s -X POST http://127.0.0.1:41414/ingest/slack \
  -H "Content-Type: application/json" \
  -H "X-Ranchhand-Token: $SECRET" \
  -d '{
    "namespace":"slack:T123:C456",
    "channel":{"teamId":"T123","channelId":"C456"},
    "items":[{"ts":"1234.5678","text":"Hello world","userName":"Dan"}]
  }'

Query:

SECRET=$(cat ~/.threadweaverinc/auth/shared_secret.txt)
curl -s -X POST http://127.0.0.1:41414/query \
  -H "Content-Type: application/json" \
  -H "X-Ranchhand-Token: $SECRET" \
  -d '{
    "namespace":"slack:T123:C456",
    "query":"hello",
    "topK": 5,
    "withText": true
  }'

Answer with citations:

SECRET=$(cat ~/.threadweaverinc/auth/shared_secret.txt)
curl -s -X POST http://127.0.0.1:41414/answer \
  -H "Content-Type: application/json" \
  -H "X-Ranchhand-Token: $SECRET" \
  -d '{
    "namespace":"slack:T123:C456",
    "query":"What did Dan say about hello?",
    "topK": 3
  }'

Profiles:

curl -s http://127.0.0.1:41414/profiles
curl -s -X POST http://127.0.0.1:41414/profiles \
  -H "Content-Type: application/json" \
  -d '{ "embed": { "model": "nomic-embed-text:latest" }, "chunking": { "chunk_tokens": 512 } }'

MCP Tools

  • openai_models_list

    • Input: {}

    • Output: OpenAI-shaped { data: [{ id, object, ... }] }

  • openai_chat_completions

    • Input: { model?: string, messages: [{ role: 'user'|'system'|'assistant', content: string }], temperature?, top_p?, max_tokens? }

    • Output: OpenAI-shaped chat completion response (single-shot; streaming TBD)

  • openai_embeddings_create

    • Input: { model?: string, input: string | string[] }

    • Output: OpenAI-shaped embeddings response

Claude/Codex (MCP)

Point your MCP config to:

{
  "mcpServers": {
    "ranchhand": {
      "command": "node",
      "args": ["/absolute/path/to/server.mjs"],
      "env": { "OAI_BASE": "http://localhost:11434/v1", "OAI_DEFAULT_MODEL": "llama3:latest" }
    }
  }
}

Notes

  • Streaming chat completions are not implemented yet (single response per call). If your backend requires streaming, we can add an incremental content pattern that MCP clients can consume.

  • RanchHand passes through OpenAI-style payloads and shapes outputs to be OpenAI-compatible, but exact metadata (usage, token counts) depends on the backend.

  • HTTP ingest is currently an acknowledgment stub (counts + sample). Chunking/embedding/upsert will be wired next; design is pluggable for local store or Qdrant.

Available Tools

3 tools
openai_chat_completionsD

Create chat completion (POST /v1/chat/completions).

ParametersJSON Schema
NameRequiredDescriptionDefault
max_tokensNo
messagesYes
modelNo
streamNo
temperatureNo
top_pNo

TDQS

D1.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It only states the action ('create chat completion') without mentioning any behavioral traits such as authentication requirements, rate limits, cost implications, response format, or error handling. This is inadequate for a tool that likely involves API calls with significant operational considerations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with a single sentence that directly states the tool's purpose. There is no wasted language or unnecessary elaboration, making it front-loaded and efficient in structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of an OpenAI chat completions API tool with 6 parameters, no annotations, and no output schema, the description is severely incomplete. It fails to explain the tool's behavior, parameter usage, or expected outcomes, leaving critical gaps for an AI agent to understand and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 0%, meaning none of the 6 parameters (max_tokens, messages, model, stream, temperature, top_p) are documented in the schema. The description adds no information about what these parameters mean, their expected formats, or how they affect the chat completion. This leaves all parameters semantically undefined.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Create chat completion (POST /v1/chat/completions)' restates the name/title with minimal elaboration. It specifies the verb 'create' and resource 'chat completion', but lacks specificity about what a chat completion entails or how it differs from sibling tools like embeddings creation or model listing. This is a tautological description that provides little additional insight beyond the tool name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like openai_embeddings_create or openai_models_list. It does not mention any context, prerequisites, or exclusions for usage. This absence of guidance leaves the agent without direction on appropriate tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openai_embeddings_createC

Create embeddings (POST /v1/embeddings).

ParametersJSON Schema
NameRequiredDescriptionDefault
inputYes
modelNo

TDQS

C2.1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must fully disclose behavioral traits. It only states the action and endpoint, failing to describe critical aspects such as authentication needs, rate limits, response format, or potential side effects (e.g., if it's a read-only or mutating operation). This leaves significant gaps in understanding how the tool behaves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with a single sentence, front-loaded with the key action. There is no wasted text, making it efficient in structure, though this brevity contributes to gaps in other dimensions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of creating embeddings (a mutating operation with parameters), no annotations, no output schema, and 0% schema coverage, the description is incomplete. It lacks essential details about behavior, parameters, and outputs, making it insufficient for an AI agent to use the tool effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, meaning the input schema provides no descriptions for parameters. The description does not add any meaning beyond the schema, failing to explain what 'input' and 'model' parameters represent, their expected formats, or examples. With 2 parameters and no compensation in the description, this is inadequate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the action ('Create embeddings') and the API endpoint ('POST /v1/embeddings'), which clarifies the verb and resource. However, it lacks specificity about what embeddings are or how they differ from sibling tools like chat completions or model listing, making it somewhat vague in distinguishing its unique purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like openai_chat_completions or openai_models_list. It does not mention any context, prerequisites, or exclusions, leaving the agent without clear usage instructions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

openai_models_listB

List models from OpenAI-compatible backend (GET /v1/models).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states the tool lists models via a GET request, implying a read-only operation, but doesn't cover aspects like authentication needs, rate limits, error handling, or the format of returned data. This leaves significant gaps in understanding how the tool behaves in practice.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that front-loads the core action ('List models') and includes essential technical details (the backend and endpoint). There is no wasted verbiage, making it highly concise and well-structured for quick understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations and output schema, the description is incomplete for a tool that interacts with an external API. It doesn't explain what the return value looks like (e.g., list of model objects), potential errors, or authentication requirements, which are critical for an AI agent to use this tool effectively in real-world scenarios.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description appropriately doesn't discuss parameters, focusing instead on the tool's purpose and endpoint. This aligns with the baseline expectation for tools with no parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('List models') and the resource ('from OpenAI-compatible backend'), with the specific API endpoint ('GET /v1/models') providing technical context. It distinguishes from siblings like 'openai_chat_completions' and 'openai_embeddings_create' by focusing on model listing rather than chat or embedding operations, though it doesn't explicitly name these alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for retrieving available models from an OpenAI-compatible API, but it doesn't provide explicit guidance on when to use this tool versus alternatives (e.g., for checking model availability before making chat completions). No exclusions or prerequisites are mentioned, leaving usage context somewhat inferred rather than clearly defined.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.0.0
    • First observedopenai_chat_completions
    • First observedopenai_embeddings_create
    • First observedopenai_models_list

TDQS

C2.6/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct purpose targeting different OpenAI API endpoints: chat completions, embeddings creation, and model listing. There is no overlap or ambiguity between these functions.

Naming Consistency4/5

The naming follows a consistent pattern with 'openai_' prefix and descriptive suffixes, though there's a minor deviation: 'openai_chat_completions' uses plural while 'openai_embeddings_create' uses singular verb form. Overall, the pattern is predictable and readable.

Tool Count3/5

With only 3 tools, the set feels thin for a server named 'RanchHand' which suggests broader functionality. While these cover core OpenAI operations, the limited scope may not fully represent what the server name implies.

Completeness3/5

For an OpenAI-compatible backend, the tools cover chat, embeddings, and models, but there are notable gaps in other common operations like image generation, audio processing, file operations, or fine-tuning endpoints. The surface is functional but incomplete for comprehensive OpenAI API coverage.

Related MCP Connectors