Skip to main content
Glama
comet-ml

Opik MCP Server

by comet-ml

Opik MCP Server

The official Model Context Protocol (MCP) server for Opik, the open-source LLM observability and evaluation platform, built by Comet. Plug your AI host (Claude Code, Cursor, VS Code Copilot, Codex, opencode, or any MCP client) directly into your Opik workspace: read traces, log scores, and save prompt versions, all from the chat.

Built for LLM engineers who already run Opik and want to drive it from the same AI assistant they code with.

You:    "Which traces in project 'demo' failed today?"
Claude: → list(entity_type="trace", project_name="demo") → "Three traces failed…"

You:    "Score trace 7f2e… 0.9 on helpfulness with reason 'great recovery'."
Claude: → write(score.create) → done

Quick start

One command registers the server with the AI clients on your machine, installs the Opik skill pack, and verifies the connection. It needs uv and no Opik SDK:

uvx opik mcp configure

It detects Claude Code, Cursor, VS Code Copilot, Codex and opencode, and sets up the server that fits your Opik:

Your Opik

Server

Transport

Sign-in

Opik Cloud (www.comet.com)

the hosted server, run by Comet

Streamable HTTP

in the browser (OAuth); no API key

Self-hosted Comet, open-source Opik

the local server, this package

stdio

env vars; an API key only where the deployment needs one

Clients load MCP servers when a session starts, so start a new session afterwards. Without a terminal, as from a coding agent or a script, name the client: uvx opik mcp configure --ai-client claude-code (or codex, cursor, vscode, opencode). Run that way it needs an existing ~/.opik.config or OPIK_API_KEY in the environment; without either, use the commands below.

Setup guide, troubleshooting and FAQ: comet.com/docs/opik/mcp-server.


Related MCP server: dap-mcp

Opik Cloud: the hosted server

Comet runs the server at https://www.comet.com/opik/api/v1/mcp. Your client connects over HTTP and opens a browser sign-in the first time. There is nothing to install, no API key, and no workspace to set: the server works in the workspace you pick when you sign in. After adding it, start a new session and ask: "list my Opik projects".

Add to Cursor Install in VS Code

Claude Code

claude mcp add --scope user --transport http opik-mcp https://www.comet.com/opik/api/v1/mcp
claude mcp login opik-mcp

--scope user makes the server available in every project; without it, Claude Code registers it for the current directory only. claude mcp login opens the sign-in; /mcp → Authenticate in a session does the same. Over SSH, claude mcp login opik-mcp --no-browser prints the sign-in URL to open on your own machine; the last step needs an interactive terminal (ssh -t).

Codex

codex mcp add opik-mcp --url https://www.comet.com/opik/api/v1/mcp

add opens the sign-in; codex mcp login opik-mcp opens it again. In ~/.codex/config.toml the same server is:

[mcp_servers.opik-mcp]
url = "https://www.comet.com/opik/api/v1/mcp"

Cursor

Use the button above, or add to ~/.cursor/mcp.json:

{
  "mcpServers": {
    "opik-mcp": {
      "url": "https://www.comet.com/opik/api/v1/mcp"
    }
  }
}

VS Code Copilot

Use the button above, or add to .vscode/mcp.json in your workspace or to your user mcp.json (MCP: Open User Configuration):

{
  "servers": {
    "opik-mcp": {
      "type": "http",
      "url": "https://www.comet.com/opik/api/v1/mcp"
    }
  }
}

Other clients

Any client that takes a URL can use the hosted server:

npx add-mcp https://www.comet.com/opik/api/v1/mcp --name opik-mcp

A client that can only start local commands can reach it through npx -y mcp-remote https://www.comet.com/opik/api/v1/mcp.

Opik Cloud with an API key

Where nobody can complete the browser sign-in, as when an agent runs unattended from a script or a client has no MCP OAuth support, use an API key from comet.com/api/my/settings/ instead, with the local server pointed at Opik Cloud:

claude mcp add --scope user opik-mcp \
  --env OPIK_API_KEY="$OPIK_API_KEY" \
  --env OPIK_WORKSPACE=<workspace> \
  -- uvx opik-mcp

The other clients take the same two variables in their env block. Set OPIK_WORKSPACE to the segment after comet.com/opik/ in your Opik URL (https://www.comet.com/opik/acme-ai/projects → acme-ai). Left out, the server sends default, which Comet resolves to your account's default workspace, so reads can come from the wrong workspace without an error.


Self-hosted and open-source Opik: the local server

opik-mcp runs on your machine: the client starts it with uvx opik-mcp and talks to it over stdio. Install uv once; it fetches the package, and Python 3.13 if needed, on first use:

curl -LsSf https://astral.sh/uv/install.sh | sh   # macOS / Linux
# or: brew install uv

uvx opik-mcp reuses the copy uv has cached, so it starts quickly; uv cache clean opik-mcp makes the next start fetch the newest release. uvx opik-mcp@latest asks PyPI on every start, which adds about a second and a half. Codex waits mcp_optional_startup_grace_ms (1 s by default) for servers before it builds the first tool list, so a slower start can leave the tools out of the first turn.

Env vars point the server at your Opik:

Your Opik

Env vars

Open-source Opik on this machine

OPIK_URL=http://localhost:5173/api

Open-source Opik on a server

OPIK_URL=https://<host>/api, plus OPIK_API_KEY only if the deployment adds authentication

Self-hosted Comet platform (the Opik UI is at https://<host>/opik)

COMET_URL_OVERRIDE=https://<host>, OPIK_WORKSPACE, OPIK_API_KEY

  • Open source serves its API at /api and has one workspace, default, so it needs no OPIK_WORKSPACE. COMET_URL_OVERRIDE would point the server at /opik/api, which open source does not serve.

  • A self-hosted Comet has named workspaces. Set OPIK_WORKSPACE to the segment after /opik/ in your Opik URL; left out, reads come from your account's default workspace, with no error to say so.

  • A self-hosted Comet that runs the MCP OAuth server (off by default) can use the hosted flow instead, at https://<host>/opik/api/v1/mcp.

  • Substitute every value. The server refuses a placeholder workspace such as <your-workspace> or ${input:OPIK_WORKSPACE}; a placeholder URL is sent as-is and fails to connect.

  • OPIK_MCP_ANALYTICS_SOURCE="" in the env block opts a self-hosted install out of the cloud-Comet source label on telemetry events.

The examples use a local open-source Opik. For a self-hosted Comet, swap in the three variables from the table. After adding the server, start a new session and ask: "list my Opik projects".

Claude Code

claude mcp add --scope user opik-mcp --env OPIK_URL=http://localhost:5173/api -- uvx opik-mcp

On a self-hosted Comet, with the key in your shell's OPIK_API_KEY:

claude mcp add --scope user opik-mcp \
  --env COMET_URL_OVERRIDE=https://<host> \
  --env OPIK_WORKSPACE=<workspace> \
  --env OPIK_API_KEY="$OPIK_API_KEY" \
  -- uvx opik-mcp

Or edit ~/.claude.json directly:

{
  "mcpServers": {
    "opik-mcp": {
      "type": "stdio",
      "command": "uvx",
      "args": ["opik-mcp"],
      "env": {
        "OPIK_URL": "http://localhost:5173/api"
      }
    }
  }
}

claude mcp get opik-mcp shows ✔ Connected once the server starts. That does not prove the URL or key are right, because only a tool call reaches Opik. It also prints the env block, API key included.

Codex

codex mcp add opik-mcp --env OPIK_URL=http://localhost:5173/api -- uvx opik-mcp

Or edit ~/.codex/config.toml:

[mcp_servers.opik-mcp]
command = "uvx"
args = ["opik-mcp"]
env = { OPIK_URL = "http://localhost:5173/api" }
# On a self-hosted Comet, forward the key from the environment Codex starts in,
# so it stays out of this file:
# env_vars = ["OPIK_API_KEY"]
# Time allowed for the server to start (default 10 s); the first start
# downloads the package.
startup_timeout_sec = 30

Codex gives a local server only the variables in env and the names in env_vars, so a key exported in your shell does not reach it otherwise.

Cursor

Add to ~/.cursor/mcp.json (global) or .cursor/mcp.json (project), or use Cmd+Shift+J → Features → Model Context Protocol:

{
  "mcpServers": {
    "opik-mcp": {
      "type": "stdio",
      "command": "uvx",
      "args": ["opik-mcp"],
      "env": {
        "OPIK_URL": "http://localhost:5173/api"
      }
    }
  }
}

Cursor 60s timeout. Cursor enforces a hard tool-call timeout that doesn't reset on progress notifications. See Known host limits.

VS Code Copilot

Add to .vscode/mcp.json in your workspace, or to your user mcp.json (MCP: Open User Configuration):

{
  "servers": {
    "opik-mcp": {
      "type": "stdio",
      "command": "uvx",
      "args": ["opik-mcp"],
      "env": {
        "OPIK_URL": "http://localhost:5173/api"
      }
    }
  }
}

MCP Inspector (manual testing)

OPIK_URL=http://localhost:5173/api npx @modelcontextprotocol/inspector uvx opik-mcp

Install with a coding agent

For an AI agent asked to install the server. The commands are in the two sections above; these rules pick which one to run. uvx opik-mcp --help prints these rules and the commands for each case.

  1. If Opik runs on this machine — curl -s http://localhost:5173/api/is-alive/ping answers — use the local server with OPIK_URL=http://localhost:5173/api. In a sandboxed shell, such as Codex's, a failed check can mean the shell has no network access rather than that Opik is down; ask the user.

  2. If the user is on Opik Cloud (www.comet.com), use the hosted server.

  3. Otherwise ask the user for their Opik URL, and use the local server with the env vars for that deployment.

  • Don't ask the user to paste an API key into the chat. The hosted server needs none; otherwise take it from the shell environment ("$OPIK_API_KEY"), or let the user run the command.

  • Don't guess the workspace. It is the segment after comet.com/opik/ (or /opik/ on a self-hosted Comet) in the user's Opik URL.

  • An opik-mcp entry may already exist, from the old npx setup or an earlier attempt. Tell the user before replacing it. Claude Code refuses to add over it, so remove it first with claude mcp remove opik-mcp --scope user; Codex's add replaces it.

  • After registering, in Claude Code, read only the status line: claude mcp get opik-mcp | grep Status. The full output prints the env block, API key included. The hosted server shows ! Needs authentication until claude mcp login opik-mcp has run. Codex has nothing that starts the server before a session; codex mcp get opik-mcp shows what was stored, with env values masked.

  • Clients load MCP servers when a session starts. Ask the user to start a new session, then try "list my Opik projects".


Coming from npx opik-mcp?

The TypeScript server (npm opik-mcp@2) is deprecated and stops serving requests on 2026-11-15. On Opik Cloud, switch to the hosted server and drop the API key. Otherwise, in your MCP client config, replace npx -y opik-mcp with uvx opik-mcp. Some env vars were renamed and command-line flags are no longer read: see the migration guide. Support policy: DEPRECATED.md. The TypeScript source is at the git tag legacy-typescript-final.


Tools

opik-mcp exposes a small, outcome-oriented surface that covers the full lifecycle (read → annotate → curate → author → iterate).

Tool

Purpose

read

Universal read by id / name / opik:// URI

list

Universal list with optional name filter + pagination

write

Universal write — log traces/spans, score, comment, save prompts, manage datasets & experiments

schema

Introspect write-operation schemas (used by the LLM to construct valid payloads)

read_skill

Read one of the Opik agent skills bundled with this server

read

One tool for any "show me X" question. Takes an entity_type plus an id (UUID or, for nameable types, a name) or a full opik:// URI. Composite reads (trace, prompt, thread, agent_insights_issue) inline their children so a single call returns the full picture.

The record you name comes back whole. Inlined children do not: their bodies are fetched with the backend's truncate=true, so a field over ~10 KB is cut in ClickHouse and base64 images are replaced with "[image]" — one attachment echoed across 200 spans would otherwise cost more than everything else in the read. The answer says so in spanBodies / messageBodies, and any child is whole again through its own read("span", id) or read("trace", trace_id), which hit endpoints that have no truncate parameter at all.

An inlined collection is also bounded in length: 200 spans, 200 turns, 100 prompt versions. Past that, spansTruncated / messagesTruncated / versionsTruncated is true and a moreSpans / moreMessages / moreVersions line beside it carries the count and the exact list(...) call that continues from where the inlined part stopped.

Supported entities: project, trace, span, dataset, dataset_item, experiment, prompt, thread, agent_insights_issue. Name-based lookup is available for project, experiment, prompt, dataset (slower — two API calls — and may return multiple matches). thread and agent_insights_issue are project-scoped: pass project_id or project_name, or a link/URI that carries the project. dataset and dataset_item were called test_suite and test_suite_item before; the old names still resolve, but they are not advertised and new code should use the new ones.

read(entity_type="trace", id="7f2e3c8a-…")
read(entity_type="project", id="demo")  # name lookup
read(entity_type="trace", id="opik://traces/7f2e3c8a-…")
read(entity_type="agent_insights_issue", id="<issue-uuid>", project_id="<project-uuid>")
read(
    entity_type="agent_insights_issue",
    id="https://www.comet.com/opik/<ws>/projects/<pid>/diagnostics?issue=<id>",
)

A link copied from the Opik UI works as the id: a thread link or a Diagnostics page link carries the project, so no project_id is needed and the entity type is taken from the link.

A project read answers "how is my project doing" in one call. It returns {project, summary, vocabulary, contains, url}: the record, then the four figures the Logs page shows as cards (trace count, error rate, average duration, total cost) for the last 7 days against the 7 before, SDK traffic only, as on screen. since / until move that window; since="30d" is what the UI opens on. A rate or an average over a period with no traces comes back as null, because 0% errors on a week with no traffic reads as a healthy week.

vocabulary is the map you need before you can ask anything else: the project's feedback score names, its token usage keys, and the automation rules scoring its traces. These are the names that go into a filter or into series= below, and guessing them returns an empty page that reads like good news. Score names and rules are capped, always report the true total, and name the call that returns the rest; usage keys are listed in full, since nothing else enumerates them. contains names the freshest experiment, dataset, prompt version and optimization run, so "what has been happening here" does not need four more calls. A part that failed to load says so instead of looking empty, and an empty one is omitted.

An agent_insights_issue read returns {issue, example_trace_ids, details}: the Diagnostics issue record (name, description, cause, suggested fix, severity, status), the deduplicated ids of the traces that exhibit it (the same sample the Diagnostics page shows — open one with read("trace", id)), and the per-day breakdown. Trace bodies are not inlined, so the read stays one backend call. since / until narrow the per-day rows; the default is all-time. When the server knows the Opik URL and the session's workspace, the read also carries url (the issue's Diagnostics page) and trace_url_template (a deep link for any of the example traces), so the assistant can hand you something clickable; under an OAuth session whose workspace could not be resolved the links are omitted rather than guessed.

Traces themselves carry no URL — a link for one is not derivable from the fields a read or list returns, and a guessed shape 404s. The session instructions name a template for it instead, .../v1/session/redirect/projects/?trace_id={trace_id}&path=..., so the assistant fills in an id and hands you a link. It goes through opik-backend's redirect, which resolves the project and the workspace from the trace, so it works where a direct project URL cannot, an OAuth session with an unresolved workspace included. It is the same link the Python SDK prints for a trace.

list

Browse or search a collection with pagination. Project-scoped types (trace, span, thread, agent_insights_issue, dataset_item, prompt_version) need their parent: a project UUID or name, a dataset UUID, or a prompt UUID.

list(entity_type="experiment", page=1, size=25)
list(entity_type="experiment", name="rerank")  # name substring filter
list(entity_type="agent_insights_issue", project_name="demo")  # open Diagnostics issues
list(entity_type="agent_insights_issue", project_id="<uuid>", status="resolved")
list(entity_type="trace", project_name="demo")  # latest traces of one project
list(
    entity_type="trace", project_name="demo", filters="error_info is_not_empty AND duration > 5000"
)
list(
    entity_type="span",
    project_name="demo",  # spans across the whole project
    filters='type = "llm" AND usage.total_tokens > 10000',
)
list(
    entity_type="thread",
    project_name="demo",
    filters="number_of_messages > 20 AND feedback_scores.helpfulness < 0.5",
)
list(entity_type="experiment", filters='dataset_id = "<dataset-uuid>" AND tags contains "baseline"')

Filters. trace, span, thread, experiment and dataset_item take an OQL string, the same grammar as the SDK's search_traces(filter_string=…):

<field>[.<key>] <op> <value> [AND ...]
ops: = != > >= < <= contains not_contains starts_with ends_with is_empty is_not_empty in not_in

Strings go in double quotes, numbers are bare, duration is in milliseconds, dates are ISO-8601 instants with a timezone ("2026-09-08T10:00:00Z"). Scores and dictionaries take a key: feedback_scores.accuracy < 0.5, metadata.environment = "prod". AND is the only connector.

Like the UI's Logs page, trace, span and thread lists add source = "sdk" so evaluator, playground and experiment traces stay out of the way; name source yourself to see them. The first output line echoes the filter that was applied.

A bad filter fails before reaching the backend with what is needed to fix it: the position of a syntax error, the closest field name, the valid operators for the field's type, or the expected value format. Fields with a closed set of values (source, span type, thread status, visibility_mode) are checked against it too, every element of an in list included. source is the one the backend validates itself, and it answers an unknown value with a 500 rather than a 400, so source = "SDK" would otherwise be an opaque server error for a capital letter. The rest are compared as strings and answer with an empty page, which reads as "no matches" when it means "no such value". Ask schema("list.trace") (or list.span, list.thread, list.experiment) for the full field reference, accepted values included.

Finding one case in a dataset. list(entity_type="dataset_item", dataset_id=…) filters on the case itself: data.<key> for the keys the dataset was built with, full_data for a substring of the whole payload (a full scan — name a key when you can), plus id, tags, source, trace_id, span_id and the timestamps. data.<key> takes the six string operators only (=, !=, contains, not_contains, starts_with, ends_with); the backend answers a comparison with a 400, so this one is refused before the call. The endpoint has no sorting and no free-text search — sort is refused rather than dropped. read(entity_type="dataset_item", id=…) returns one case whole, which is how a value the table cut is read back.

list(entity_type="dataset_item", dataset_id="<uuid>", filters='data.question contains "install"')
list(
    entity_type="dataset_item", dataset_id="<uuid>", filters='trace_id = "<trace-uuid>"'
)  # the case made from that trace
read(entity_type="dataset_item", id="<item-uuid>")  # the case, uncut

With experiment_ids the same list is the comparison instead — the cases with each run attached — and it filters on the runs (feedback_scores.<name>, output, duration). The two are different field sets on two backend endpoints: schema("list.dataset_item_case") is the dataset's own cases, schema("list.dataset_item") the comparison.

Sort. trace, span, thread and experiment take sort="<field> [asc|desc]", desc by default and one field only: sort="duration desc", sort="total_estimated_cost", sort="feedback_scores.accuracy asc", sort="usage.total_tokens". The field is checked against the entity's sortable list before the call, because the backend silently ignores fields it cannot sort by. On very large workspaces the backend drops sorting altogether; the header says so when that happens.

dataset_item sorts only as a comparison (with experiment_ids): the items endpoint takes no sorting parameter, so a sort on a plain listing is refused rather than dropped.

Time window and search. trace, span and thread take since and until, each a relative span ("30m", "1h", "7d") or an ISO-8601 instant with a timezone, so "the last hour" needs no clock arithmetic. The window is by record creation time, which is cheap for the backend and agrees with start_time within seconds for live traffic. For an exact bound, put start_time in filters. The same three types take search, free text matched anywhere in id, name, input, output, metadata, tags and thread id. Search scans the whole project on the backend, so the first call on a large project can take tens of seconds. Those calls get a 60-second timeout. Adding since makes them fast again.

Reading the table. Durations are labelled duration_ms / ttft_ms and shown as whole milliseconds; the field stays duration in filters and sort. Timestamps are shown to the second and costs as plain decimals. Project rows carry last_updated_trace_at so you can see which project has live traffic; thread rows carry the first message. An empty page under a time window says when the project's last trace landed, and an empty page under the default source = "sdk" says how to see the other sources. A misspelled project_name comes back with the closest existing name.

list(
    entity_type="trace",
    project_name="demo",
    since="1h",
    filters="error_info is_not_empty",
    sort="duration desc",
)
list(entity_type="trace", project_name="demo", search="order-42")

Diagnostics issues. agent_insights_issue is the Diagnostics page over the MCP: the recurring failures Opik's Diagnostics job grouped for a project, ranked as the UI ranks them (most recently seen first). Columns are severity, status, total_occurrences (all-time sum), latest_count (the most recent report day, the number the issue's own description refers to) and last_seen. Open issues are listed by default; pass status="resolved" or "closed" for the rest. read and list also answer to issue, which is what the UI calls these; the long name is the one in the entity_type enum, so that one entity does not appear there twice. Counts are all-time so they match the UI; the same since / until as for traces narrow the window, truncated to UTC report days because Diagnostics aggregates per day.

An empty list says why it is empty, because "nothing is broken" and "nobody turned Diagnostics on" read the same otherwise. There are five states: Diagnostics is unavailable on this deployment, not enabled for this project, turned off, enabled but not scanned recently, or enabled and clean with the time of the last scan. The ones you can act on name the call to make, and every state links the project's Diagnostics page.

A non-empty list dates itself. The issues are whatever the last scan grouped, so the reply ends with Report covers data through <time>, and when the window you asked about runs past that, it names the uncovered tail and how to close it: a trigger when a rescan reaches back far enough, otherwise raw traces with the since it gives you. Ask for a week on a project scanned nightly and the last day is missing from the grouped answer; this is what says so.

write("agent_insights_job.enable", {"project_name": "demo"}) turns Diagnostics on. It scans daily from then on, and calling it again is safe. write("agent_insights_job.trigger", …) scans the last 24 hours now, without waiting for the nightly run. Both take the permission that reading issues takes, and both refuse where the deployment has no Diagnostics.

An issue moves through its lifecycle with write("agent_insights_issue.resolve", {"issue_id": "<uuid>", "project_name": "demo"}) — dealt with — or …close for one not worth acting on, and …reopen to put either back on the open list. All three take the same permission and answer with a link to the view the issue moved to, since a resolved issue is no longer on the default page. Whether a failure is fixed is a judgment call, so these are for when you ask: the assistant has no business tidying the list while triaging it.

Metrics over time. project_metric charts one metric for a project as a table of time buckets: trace, span and thread counts, durations, error rates, costs, token usage and feedback scores. It answers the question that follows the overview, which is when something changed.

list(entity_type="project_metric", project_name="demo", metric_type="trace_count")
list(
    entity_type="project_metric",
    project_name="demo",
    metric_type="trace_error_rate",
    since="14d",
    interval="daily",
)
list(
    entity_type="project_metric", project_name="demo", metric_type="span_count", breakdown="model"
)  # one column per model
list(
    entity_type="project_metric",
    project_name="demo",
    metric_type="span_duration",
    breakdown="model",
    series="p99",
)  # the p99 of each model

Rows are time buckets, not records, so page, size and sort are refused rather than ignored. interval is hourly, daily, weekly or total; left out, it follows the window the way the Metrics tab does — hourly up to 3 days, daily up to 30, weekly beyond — so a default chart is a few dozen rows whatever the range, and an hourly month (721 rows) is something you ask for. since / until take the same forms as everywhere else and default to the last 7 days. filters uses the fields of whichever entity the metric is about, so a span metric is filtered by span fields.

breakdown splits each bucket by tags, name, error_info, error_type, model, provider, span_type, guardrail_name or metadata.<key>. Not every metric accepts every one of those, and seven accept none at all; the tool knows which and says so before calling the backend, naming a metric that does answer the same question where one exists. Three families come back as several series at once (a duration as p50/p90/p99, a feedback score per name, token usage per key), and the backend charts one of them at a time when grouping, so series= picks it: a percentile, a score name, or a usage key. Duration defaults to p50 and token usage to total_tokens, and whichever was used is echoed on the first line.

Empty buckets are left out and counted underneath, so a quiet month is a few rows instead of a column of zeros, and a rate over a bucket with no traces is absent rather than reported as zero.

Ask schema("list.project_metric") for the metric table, the intervals and the per-metric grouping matrix.

A project's names. score_name lists the feedback score names recorded in a project and online_rule the automation rule evaluators configured on it, which is where most of those names come from. Both are the same lists read("project", …) carries, in full and paginated, for when the capped version in the overview is not enough.

list(entity_type="score_name", project_name="demo")
list(entity_type="online_rule", project_name="demo")

write

Universal write dispatcher. Pass operation + data and the dispatcher validates the payload, applies the right REST verb, and returns the backend response.

Operations:

Operation

What it does

trace.create

Log a single trace (or a batch). Parent for spans / scores / comments.

trace.update

Finalize or amend an existing trace.

span.create

Log a span on an existing trace (or a batch).

score.create

Attach a numeric feedback score to a trace, span, or thread.

comment.create

Attach a free-text comment to a trace, span, or thread.

prompt_version.save

Save a new prompt version (creates the prompt by name if missing).

dataset.create

Create a dataset — type: "test_suite" makes it an evaluation test suite.

dataset_item.upsert

Upsert items into a dataset (always the envelope shape).

experiment.create

Create an experiment scoped to a dataset.

experiment_item.create

Attach trace + dataset_item rows to an experiment.

thread.close

Close a thread (mark it inactive). Pass thread_id and the project.

thread.open

Reopen a closed thread. Pass thread_id and the project.

agent_insights_job.enable

Turn Diagnostics on for a project (daily scans, safe to repeat).

agent_insights_job.trigger

Run a Diagnostics scan now, over the last 24 hours.

agent_insights_issue.resolve

Mark a Diagnostics issue dealt with (ask the user first).

agent_insights_issue.close

Mark a Diagnostics issue not worth acting on (ask the user first).

agent_insights_issue.reopen

Put a resolved or closed Diagnostics issue back on the open list.

write(
    operation="score.create",
    data={
        "target": "trace",
        "target_id": "7f2e3c8a-…",
        "name": "helpfulness",
        "value": 0.9,
        "reason": "great recovery",
    },
)

schema

Inspect the exact JSON shape and required fields of any write operation before you call it — useful when you're not sure what data should look like. Returns the schema, OAuth scope, and one validated example. Pure lookup, no backend call.

schema(operation="score.create")
schema(operation="prompt_version.save")

The same tool answers list.trace, list.span, list.thread and list.experiment with the list tool's reference for that entity: every filterable field with its type and valid operators, the sortable fields, whether a time window and free-text search apply, and two example filters.

schema(operation="list.trace")

Configuration

These configure the local server; every setting is an environment variable. The hosted server on Opik Cloud takes none of them.

Identity / endpoint

Variable

Default

Notes

OPIK_API_KEY

—

API key, for a self-hosted Comet, or for Opik Cloud without the hosted server. Open-source Opik needs none unless the deployment adds authentication.

OPIK_WORKSPACE

unset

Workspace name. On cloud with an API key, unset sends default, which resolves to your account's default workspace — set it explicitly if you work in a different one, or reads come from the wrong workspace silently. Leave unset over OAuth (the token carries it) and on local/OSS (default is the only workspace there).

COMET_WORKSPACE

—

Deprecated alias for OPIK_WORKSPACE (backward compat). OPIK_WORKSPACE wins if both are set.

COMET_WORKSPACE_ID

unset

Optional workspace UUID. Stamped into analytics events when set, and takes precedence over the resolved one. Rarely needed — OAuth installs get the UUID from the token automatically.

COMET_URL_OVERRIDE

https://www.comet.com

Set to your self-hosted Comet host, or https://dev.comet.com for staging.

OPIK_URL

derived from COMET_URL_OVERRIDE + /opik/api

Set it for open-source Opik, which serves its API at /api (http://localhost:5173/api locally). On a Comet platform, override only if Opik lives on a different host/path than the Comet UI.

OPIK_DEFAULT_PROJECT_NAME

unset

When set, the per-session instructions blob tells the LLM to pass this as project_name on every tool call unless the user names a different project.

Server / transport

Variable

Default

Notes

OPIK_MCP_TRANSPORT

stdio

stdio for host-launched, streamable-http to listen on a port.

OPIK_MCP_HOST

127.0.0.1

uvicorn bind host (streamable-http only).

OPIK_MCP_PORT

8080

uvicorn bind port (streamable-http only).

OPIK_MCP_RELOAD

false

true to enable uvicorn --reload (dev only).

OPIK_MCP_AS_URL

unset

OAuth Authorization Server URL, advertised in /.well-known/oauth-protected-resource (RFC 9728) and used as the proxy target for AS-discovery probes. Required for MCP hosts to bootstrap the OAuth dance over HTTP.

OPIK_MCP_RESOURCE_URI

unset

Canonical public URI of this server, advertised as resource in the protected-resource metadata and used to derive the WWW-Authenticate hint.

OPIK_MCP_OAUTH_VALIDATION_CACHE_TTL_S

30

How long a "valid" answer from opik-backend's token introspection is trusted before the next request on the same OAuth token asks again. Bounds the backend load added by per-request validation and the window in which an expired token is still forwarded (that window also ends on the first 401 the backend returns). Capped by the token's own expires_at when the backend reports one.

OPIK_MCP_LOG_LEVEL

INFO

stderr logger threshold.

Choosing a transport

Two bearer shapes, two contracts on HTTP transport. An opik_mcp_at_… OAuth access token is validated on every request against opik-backend's token introspection endpoint (cached, see OPIK_MCP_OAUTH_VALIDATION_CACHE_TTL_S); an expired or revoked token gets an HTTP 401 with WWW-Authenticate: Bearer error="invalid_token", which is what MCP hosts key their silent refresh_token grant on. An Opik API key is not validated locally: it is forwarded verbatim to opik-backend, which is its single point of enforcement. Pick the transport by deployment shape:

Scenario

Transport

Opik Cloud

Nothing to run: the hosted server is this server over HTTP, run by Comet

MCP client and Opik on the same machine (local OSS install)

stdio (recommended — simplest, no port, no OAuth setup)

Local MCP client → self-hosted Opik

stdio with the env vars for the deployment, or HTTP with OAuth (OPIK_MCP_AS_URL pointing at the backend)

opik-mcp served behind the same edge as opik-backend

HTTP — bearers are validated by the backend per request

Note for local OSS installs: the OSS backend does not authenticate requests, so an HTTP opik-mcp in front of it is as open as the OSS REST API itself. Keep the default 127.0.0.1 bind (and prefer stdio) on shared networks.

Telemetry

Anonymous usage events (event type + timing only — no query content). A SHA-256 digest of your API key is included so support can find your account; the raw key never leaves the process. Opt out: OPIK_MCP_ANALYTICS_ENABLED=false.

Variable

Default

Notes

OPIK_MCP_ANALYTICS_ENABLED

true

Set to false to disable all telemetry.

OPIK_MCP_ANALYTICS_URL

https://stats.comet.com/notify/event/

Override for staging.

OPIK_MCP_ANALYTICS_ENVIRONMENT

prod

Tag on every event (prod / staging / dev).

OPIK_MCP_ANALYTICS_SOURCE

comet.com

Receiver uses this to mark on_prem=False. On-prem installs should override to "" or their own domain.

OPIK_MCP_ANALYTICS_CONNECT_TIMEOUT_S

5.0

HTTP connect timeout.

OPIK_MCP_ANALYTICS_TOTAL_TIMEOUT_S

10.0

HTTP total request timeout.


Known host limits

Hosts differ in how long they let a single tool call run:

  • Claude Code — no documented tool-call timeout. Recommended.

  • Cursor — hard 60s timeout that does not reset on progress (upstream bug).

  • MCP Inspector — MAX_TOTAL_TIMEOUT bounds total duration (default 60s). Raise it in the Inspector UI for long operations.

If a call gets stuck, set OPIK_MCP_LOG_LEVEL=DEBUG for the full request log.


Troubleshooting

OPIK_API_KEY isn't picked up — the var isn't reaching the server process. In Claude Code / Cursor / VS Code, env vars only apply when inside the env block of the MCP server config, not your shell; Codex also forwards the names listed in env_vars. Start a new session after editing, since clients read the config when a session starts.

Requests go to /opik/api on an open-source Opik — COMET_URL_OVERRIDE is for a self-hosted Comet platform. Open source serves its API at /api: set OPIK_URL=http://localhost:5173/api (or https://<host>/api) instead.

Cursor call times out at 60s — Cursor's known bug, not opik-mcp. Either narrow the call (smaller size, a tighter window), or run the same operation on Claude Code which has no hard cap.

Server not showing, sign-in not opening, wrong workspace, uvx not found. These are covered in the troubleshooting section of the docs. opik mcp status (from the same uvx opik CLI) lists every client that has the server configured and whether its config has drifted.


Development

git clone git@github.com:comet-ml/opik-mcp.git
cd opik-mcp
make install        # uv sync --locked --extra dev
make check          # lint + typecheck + test
make run-dev        # uvicorn with --reload + DEBUG logs
make inspect        # MCP Inspector against the running server

Common targets:

Target

What it does

make install

uv sync --locked --extra dev

make run

Run the MCP server (stdio by default).

make run-dev

Run with DEBUG logging + uvicorn --reload.

make dev

Run via mcp dev (Inspector dev-mode wrapper).

make inspect

Launch MCP Inspector against a running server.

make test

uv run pytest -q.

make lint

ruff check + format check.

make format

ruff format + ruff check --fix.

make typecheck

mypy.

make check

lint + typecheck + test.

Repo layout:

opik-mcp/
├── src/opik_mcp/        ← server, tools, analytics
├── tests/               ← pytest suites
├── scripts/             ← live-BE smoke + MCP-session smoke
├── legacy/typescript/   ← migration guide for the deprecated v2 TS server (source: tag `legacy-typescript-final`)
├── pyproject.toml
└── Makefile

Get help


License

Apache-2.0.

Available Tools

19 tools
create-projectCInspect

Create a new project/workspace

ParametersJSON Schema
NameRequiredDescriptionDefault
descriptionNoDescription of the project
nameYesName of the project
workspaceNameNoWorkspace name to use instead of the default

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the action ('Create') but doesn't cover critical aspects like required permissions, whether the creation is idempotent, what happens on conflicts, or the expected response format. For a mutation tool with zero annotation coverage, this leaves significant gaps in understanding its behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with zero wasted words—it directly states the tool's purpose without unnecessary elaboration. It's appropriately sized and front-loaded, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that this is a mutation tool ('Create') with no annotations and no output schema, the description is insufficient. It doesn't explain what a 'project' or 'workspace' entails in this system, how to handle errors, or what the tool returns upon success, leaving the agent with incomplete context for effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds no parameter-specific information beyond what's already in the input schema, which has 100% coverage with clear descriptions for all three parameters. This meets the baseline of 3, as the schema adequately documents the parameters, but the description doesn't provide additional context like examples or usage notes.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Create a new project/workspace' clearly states the verb ('Create') and resource ('project/workspace'), making the purpose immediately understandable. However, it doesn't distinguish this tool from its sibling 'create-prompt' or explain what differentiates a 'project' from a 'workspace' in this context, which prevents a perfect score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'update-project' or 'list-projects', nor does it mention prerequisites or constraints. While the verb 'Create' implies it's for new entities, there's no explicit comparison to siblings or context for usage decisions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create-promptCInspect

Create a new prompt

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesName of the prompt

TDQS

C2.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden for behavioral disclosure. 'Create a new prompt' implies a write/mutation operation but provides no details about permissions needed, whether creation is idempotent, what happens on conflicts, or what the response contains. This leaves significant behavioral gaps for a creation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise at just three words. While it's under-specified in content, it's not verbose or poorly structured. Every word earns its place in this minimal description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a creation tool with no annotations and no output schema, the description is inadequate. It doesn't explain what a 'prompt' is in this context, what fields beyond 'name' might be set by default, what the creation response looks like, or how this differs from similar tools. The description fails to provide necessary context for effective tool use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the single 'name' parameter. The description adds no additional parameter context beyond what's in the schema. According to scoring rules, when schema coverage is high (>80%), the baseline is 3 even with no parameter information in the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Create a new prompt' is a tautology that restates the tool name without adding specificity. It doesn't distinguish this tool from sibling tools like 'create-prompt-version' or explain what kind of prompt is being created. The purpose is minimally stated but lacks meaningful differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus alternatives like 'create-prompt-version' or 'update-prompt'. There's no mention of prerequisites, constraints, or appropriate contexts for invoking this creation tool versus other prompt-related operations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create-prompt-versionCInspect

Create a new version of a prompt

ParametersJSON Schema
NameRequiredDescriptionDefault
commit_messageYesCommit message for the prompt version
nameYesName of the original prompt
templateYesTemplate content for the prompt version

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool creates a new version but doesn't explain what that entails—whether it's a write operation, requires specific permissions, affects existing prompts, or has side effects like triggering notifications. For a mutation tool with zero annotation coverage, this lack of detail is a significant gap in transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It's front-loaded with the core action ('Create a new version'), making it easy to parse quickly. Every word earns its place, and there's no redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of a versioning operation (a mutation with potential side effects), no annotations, and no output schema, the description is incomplete. It doesn't cover behavioral aspects like permissions, idempotency, or error handling, nor does it hint at return values. For a tool that modifies data, this leaves critical gaps for an AI agent to understand its full context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents the three parameters (name, template, commit_message). The description adds no additional meaning beyond what's in the schema, such as explaining how 'name' relates to the original prompt or what 'commit_message' is used for. With high schema coverage, a baseline score of 3 is appropriate, as the description doesn't compensate but also doesn't need to.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Create') and resource ('new version of a prompt'), making the purpose immediately understandable. It distinguishes from siblings like 'create-prompt' (which creates a new prompt rather than a version) and 'update-prompt' (which might modify an existing prompt without versioning). However, it doesn't specify what constitutes a 'version' (e.g., whether it's a snapshot, revision, or branch), leaving some ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing an existing prompt), compare to siblings like 'update-prompt' or 'create-prompt', or indicate scenarios where versioning is appropriate (e.g., for tracking changes, experimentation, or deployment). Without this, users must infer usage from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete-projectCInspect

Delete a project

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdYesID of the project to delete
workspaceNameNoWorkspace name to use instead of the default

TDQS

C2.1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden but discloses nothing beyond the basic action. It fails to address critical behavioral traits: whether deletion is permanent or reversible, required permissions, side effects (e.g., on associated data), error conditions, or response format. For a destructive operation, this lack of transparency is severe.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with a single sentence, 'Delete a project', which is front-loaded and wastes no words. While under-specified, it efficiently communicates the core action without redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (destructive operation with 2 parameters), lack of annotations, and no output schema, the description is incomplete. It does not compensate for missing behavioral context, usage guidelines, or output expectations, leaving significant gaps for an AI agent to understand and invoke the tool safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with clear parameter documentation in the schema itself. The description adds no parameter semantics beyond what the schema provides, such as explaining 'projectId' format or 'workspaceName' usage. However, the baseline score of 3 is appropriate since the schema adequately covers parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Delete a project' is a tautology that restates the tool name without adding meaningful context. It specifies the verb 'Delete' and resource 'project', but lacks distinction from sibling tools like 'delete-prompt' or details about what deletion entails. This minimal statement fails to clarify scope or differentiate from alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. It does not mention prerequisites (e.g., project existence), exclusions (e.g., irreversible effects), or comparisons to sibling tools like 'update-project' or 'list-projects'. The description offers no context for decision-making.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete-promptCInspect

Delete a prompt

ParametersJSON Schema
NameRequiredDescriptionDefault
promptIdYesID of the prompt to delete

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states the action is destructive ('Delete') but doesn't elaborate on consequences (e.g., permanent deletion, no undo), permissions required, or error conditions. This is a significant gap for a mutation tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise at three words, front-loading the essential action and resource without any wasted text. Every word earns its place, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's destructive nature, no annotations, and no output schema, the description is incomplete. It lacks critical information about behavioral traits (e.g., irreversibility), error handling, or what happens post-deletion. For a deletion tool, this leaves significant gaps in understanding.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds no parameter semantics beyond what the schema provides. Since schema description coverage is 100% (the 'promptId' parameter is fully documented in the schema), the baseline score is 3. The description doesn't compensate with additional context like format examples or constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Delete') and the resource ('a prompt'), making the tool's purpose immediately understandable. It distinguishes itself from siblings like 'delete-project' by specifying the resource type, though it doesn't explicitly contrast with other deletion tools. The description avoids tautology by not merely restating the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing a prompt ID), exclusions (e.g., cannot delete if in use), or sibling tools like 'delete-project' for different resources. Usage is implied by the action but lacks explicit context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-metricsCInspect

Get metrics data

ParametersJSON Schema
NameRequiredDescriptionDefault
endDateNoEnd date in ISO format (YYYY-MM-DD)
metricNameNoOptional metric name to filter
projectIdNoOptional project ID to filter metrics
projectNameNoOptional project name to filter metrics
startDateNoStart date in ISO format (YYYY-MM-DD)

TDQS

C2/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure but fails completely. 'Get metrics data' reveals nothing about whether this is a read-only operation, whether it requires authentication, what rate limits might apply, what format the data returns in, or any other behavioral characteristics. For a data retrieval tool with zero annotation coverage, this is a critical gap that leaves the agent with no understanding of how the tool behaves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise at just three words. While this represents severe under-specification in terms of content, from a pure conciseness perspective it contains zero wasted words and is front-loaded with the core action. Every word earns its place, even though that place is inadequate for proper tool understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a 5-parameter tool with no annotations and no output schema, the description 'Get metrics data' is completely inadequate. It provides no information about what metrics are available, what system they come from, what the return format looks like, or any behavioral characteristics. For a data retrieval tool of this complexity, the description fails to provide the minimal contextual information needed for an agent to use it effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, meaning all 5 parameters are well-documented in the input schema itself. The description adds absolutely no additional parameter information beyond what's already in the schema. According to the scoring rules, when schema coverage is high (>80%), the baseline score is 3 even with no parameter information in the description, which applies here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get metrics data' is a tautology that essentially restates the tool name 'get-metrics'. It provides no specific information about what kind of metrics, from what system, or what scope. While it includes a verb ('Get') and resource ('metrics data'), it lacks any distinguishing details that would help differentiate it from potential sibling tools or clarify its specific function beyond the obvious.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides absolutely no guidance on when to use this tool versus alternatives. There are no mentions of prerequisites, appropriate contexts, or comparisons to sibling tools like 'get-trace-stats' or 'get-trace-by-id' that might handle similar data. The agent receives no help in determining when this specific metrics retrieval tool is appropriate versus other data-fetching tools in the server.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-opik-examplesCInspect

Get examples of how to use Opik Comet's API for specific tasks

ParametersJSON Schema
NameRequiredDescriptionDefault
taskYesThe task to get examples for (e.g., 'create prompt', 'analyze traces', 'monitor costs')

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool retrieves examples but doesn't describe what the output looks like (e.g., format, structure), whether it's a read-only operation, or any limitations like rate limits or authentication needs. This leaves significant gaps for an AI agent to understand how to handle the tool's behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose without any unnecessary words. It's appropriately sized and front-loaded, making it easy to parse quickly, which is ideal for conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of annotations and output schema, the description is incomplete. It doesn't explain what the tool returns (e.g., code snippets, documentation), how examples are structured, or any behavioral traits. For a tool with no structured data beyond the input schema, the description should provide more context to help an AI agent use it effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, with the 'task' parameter clearly documented. The description doesn't add any additional meaning beyond what the schema provides, such as explaining the semantics of example tasks or providing usage context. With high schema coverage, the baseline score of 3 is appropriate, as the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('Get') and resource ('examples of how to use Opik Comet's API'), making it easy to understand what the tool does. However, it doesn't explicitly differentiate from sibling tools like 'get-opik-help' or 'get-opik-tracing-info', which might provide related but different information.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention when it's appropriate (e.g., for learning API usage) or when not to use it (e.g., for direct API calls), nor does it reference sibling tools like 'get-opik-help' that might serve similar purposes.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-opik-helpBInspect

Get contextual help about Opik Comet's capabilities

ParametersJSON Schema
NameRequiredDescriptionDefault
subtopicNoOptional subtopic for more specific help
topicYesThe topic to get help about (prompts, projects, traces, metrics, or general)

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool retrieves help information, implying a read-only operation, but doesn't mention any behavioral traits like rate limits, authentication needs, or what the output format might be. This is a significant gap for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose without any unnecessary words. It is appropriately sized and front-loaded, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's low complexity (2 parameters, no output schema, no annotations), the description is minimally adequate but incomplete. It explains what the tool does but lacks details on behavioral aspects and usage context, which are important for an agent to operate effectively without annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, clearly documenting both parameters ('topic' and 'subtopic') with their types and purposes. The description adds no additional meaning beyond what the schema provides, so it meets the baseline of 3 where the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get' and the resource 'contextual help about Opik Comet's capabilities', making the purpose understandable. However, it doesn't specifically differentiate from sibling tools like 'get-opik-examples' or 'get-opik-tracing-info', which also provide information about Opik Comet but focus on different aspects.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives such as 'get-opik-examples' or 'get-opik-tracing-info'. It lacks explicit context, exclusions, or prerequisites, leaving the agent to infer usage based on the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-opik-tracing-infoCInspect

Get information about Opik's tracing capabilities and how to use them

ParametersJSON Schema
NameRequiredDescriptionDefault
topicNoOptional specific tracing topic to get information about (e.g., 'spans', 'distributed', 'multimodal', 'annotations')

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden but only states what the tool does without disclosing behavioral traits. It doesn't cover aspects like whether it's a read-only operation, potential rate limits, authentication needs, or output format, which are critical for an agent to use it effectively.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence that efficiently conveys the tool's purpose without unnecessary words. It's front-loaded and appropriately sized, though it could be slightly more structured by including brief usage hints.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's low complexity (one optional parameter, no output schema, no annotations), the description is minimally complete but lacks depth. It explains what the tool does but doesn't provide enough context about behavior or usage relative to siblings, making it adequate but with clear gaps for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 100% description coverage, clearly documenting the optional 'topic' parameter with examples. The description adds no additional parameter semantics beyond what the schema provides, so it meets the baseline of 3 for adequate but not enhanced coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get' and the resource 'information about Opik's tracing capabilities and how to use them', making the purpose specific and understandable. However, it doesn't explicitly differentiate from sibling tools like 'get-trace-by-id' or 'get-trace-stats', which also deal with tracing information but focus on specific data rather than general capabilities.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention scenarios like needing overviews versus detailed data, or how it differs from siblings such as 'get-opik-help' or 'get-opik-examples', leaving the agent to infer usage from context alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-project-by-idCInspect

Get a single project by ID

ParametersJSON Schema
NameRequiredDescriptionDefault
projectIdYesID of the project to fetch
workspaceNameNoWorkspace name to use instead of the default

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden for behavioral disclosure. While 'Get' implies a read operation, it doesn't specify whether this requires authentication, what happens if the project doesn't exist, or any rate limits. The description lacks crucial behavioral context that would help an agent understand how to handle errors or what to expect from the operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise at just 6 words, front-loading the essential information with zero wasted words. Every word earns its place by communicating the core functionality without unnecessary elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read operation with 2 parameters and no output schema, the description is insufficiently complete. It doesn't explain what information the tool returns about projects, how to handle the optional workspaceName parameter, or what format the response takes. With no annotations and no output schema, the agent lacks crucial information about the tool's behavior and outputs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents both parameters (projectId and workspaceName). The description adds no additional parameter information beyond what's in the schema. According to scoring rules, when schema coverage is high (>80%), the baseline is 3 even with no param info in the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get') and resource ('a single project by ID'), making the purpose immediately understandable. It distinguishes from sibling tools like 'list-projects' by specifying retrieval of a single item rather than a collection. However, it doesn't explicitly differentiate from 'get-prompt-by-id' or 'get-trace-by-id' which follow similar patterns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention when to prefer this over 'list-projects' for single-item retrieval, or how it relates to other 'get-by-id' tools for different resource types. There's also no information about prerequisites or contextual constraints.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-prompt-by-idCInspect

Get a single prompt by ID

ParametersJSON Schema
NameRequiredDescriptionDefault
promptIdYesID of the prompt to fetch

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden but only states the basic action without disclosing behavioral traits. It doesn't mention if this is a read-only operation, what happens with invalid IDs (e.g., errors), authentication needs, rate limits, or return format, leaving significant gaps in transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with a single sentence that directly states the tool's purpose, making it front-loaded and free of unnecessary words. Every part of the sentence earns its place by clearly conveying the core functionality.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of a retrieval tool with no annotations and no output schema, the description is incomplete. It doesn't explain what data is returned (e.g., prompt content, metadata), error handling, or how it fits into the broader context of prompt management, failing to compensate for the lack of structured data.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds no meaning beyond the input schema, which has 100% coverage and fully documents the 'promptId' parameter. With high schema coverage, the baseline is 3, as the schema handles parameter documentation adequately without extra detail from the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get' and the resource 'a single prompt by ID', making the purpose specific and understandable. However, it doesn't explicitly differentiate from sibling tools like 'get-project-by-id' or 'list-prompts', which would require mentioning it retrieves individual prompt details rather than projects or lists.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives such as 'list-prompts' for browsing or 'get-project-by-id' for related resources. It lacks context about prerequisites like needing a valid prompt ID or exclusions for non-existent IDs, leaving usage unclear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-server-infoBInspect

Get information about the Opik server configuration

ParametersJSON Schema
NameRequiredDescriptionDefault
random_stringNoDummy parameter for no-parameter tools

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It states it 'gets' information, implying a read-only operation, but doesn't disclose behavioral traits such as authentication needs, rate limits, response format, or potential side effects. This leaves significant gaps for a tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It's appropriately sized and front-loaded, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description is incomplete. It lacks details on what specific server configuration information is returned, how it's structured, or any behavioral context. For a tool in a set with many siblings, more completeness is needed to guide effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 1 parameter with 100% coverage, describing it as a 'Dummy parameter for no-parameter tools'. The description adds no parameter information, which is acceptable since the schema fully documents the parameter, and the tool effectively has 0 functional parameters. Baseline is 3, but the dummy nature justifies a higher score for clarity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get' and resource 'information about the Opik server configuration', making the purpose specific and understandable. However, it doesn't differentiate from sibling tools like 'get-metrics' or 'get-opik-tracing-info', which also retrieve server-related information, so it misses full sibling distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. With siblings like 'get-metrics' and 'get-opik-tracing-info' that might retrieve overlapping or related server data, there's no indication of context, prerequisites, or exclusions for tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-trace-by-idCInspect

Get a single trace by ID

ParametersJSON Schema
NameRequiredDescriptionDefault
traceIdYesID of the trace to fetch
workspaceNameNoWorkspace name to use instead of the default

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states it's a read operation ('get'), but doesn't cover aspects like error handling (e.g., what happens if the trace ID doesn't exist), rate limits, authentication requirements, or response format. For a tool with no annotation coverage, this leaves significant gaps in understanding its behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with zero wasted words. It's front-loaded with the core purpose ('Get a single trace by ID'), making it immediately understandable. Every part of the sentence contributes directly to clarifying the tool's function.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations and no output schema, the description is incomplete for a tool with 2 parameters. It lacks details on behavioral traits (e.g., error cases, permissions), return values, and usage context. While the schema covers parameters well, the overall context for safe and effective use is insufficient, especially for a read operation that might involve workspace-specific data.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with clear descriptions for both parameters: 'traceId' as the ID to fetch and 'workspaceName' as an optional override. The description adds no additional parameter semantics beyond what the schema provides, such as format examples or constraints. With high schema coverage, the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get a single trace by ID' clearly states the action (get) and resource (trace), with specificity about retrieving by ID. It distinguishes from siblings like 'list-traces' (multiple) and 'get-trace-stats' (statistics), though it doesn't explicitly name them. The purpose is unambiguous but lacks explicit sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention when to choose it over 'list-traces' for multiple traces or 'get-trace-stats' for aggregated data, nor does it specify prerequisites like authentication or workspace context. Usage is implied by the name but not explicitly stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-trace-statsCInspect

Get statistics for traces

ParametersJSON Schema
NameRequiredDescriptionDefault
endDateNoEnd date in ISO format (YYYY-MM-DD)
projectIdNoProject ID to filter traces
projectNameNoProject name to filter traces
startDateNoStart date in ISO format (YYYY-MM-DD)
workspaceNameNoWorkspace name to use instead of the default

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden but only states it 'gets' statistics without disclosing behavioral traits like read-only nature, potential rate limits, authentication needs, or what happens if parameters are omitted. It fails to compensate for the lack of annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with zero waste. It's appropriately sized and front-loaded, making it easy to parse quickly without unnecessary elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 parameters, no output schema, no annotations), the description is incomplete. It doesn't explain what statistics are returned, how they're formatted, or error conditions, leaving significant gaps for the agent to infer behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so parameters are fully documented in the schema. The description adds no meaning beyond the schema, such as explaining how parameters interact or default behaviors. Baseline 3 is appropriate as the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get statistics for traces' states a clear verb ('Get') and resource ('statistics for traces'), but it's vague about what specific statistics are retrieved and doesn't distinguish from sibling tools like 'get-metrics' or 'get-trace-by-id'. It provides basic purpose but lacks specificity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like 'get-metrics' or 'list-traces'. The description doesn't mention prerequisites, exclusions, or context for usage, leaving the agent without direction on tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-projectsCInspect

Get a list of projects/workspaces

ParametersJSON Schema
NameRequiredDescriptionDefault
pageYesPage number for pagination
sizeYesNumber of items per page
sortByNoSort projects by this field
sortOrderNoSort order (asc or desc)
workspaceNameNoWorkspace name to use instead of the default

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden for behavioral disclosure. It states it 'gets a list' which implies a read operation, but doesn't mention pagination behavior (implied by parameters), rate limits, authentication requirements, or what the return format looks like. For a tool with 5 parameters and no output schema, this leaves significant behavioral aspects undocumented.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise at just 5 words, front-loading the core purpose with zero wasted words. Every element ('Get', 'list', 'projects/workspaces') earns its place. No structural issues or unnecessary elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 5 parameters, no annotations, and no output schema, the description is insufficiently complete. It doesn't explain what a 'project' or 'workspace' represents in this context, doesn't describe the return format, and provides no behavioral context. The agent would need to infer too much from just the schema parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all parameters are documented in the schema. The description adds no additional parameter information beyond what's already in the schema descriptions. It doesn't explain relationships between parameters (e.g., how 'workspaceName' interacts with the list) or provide usage examples. Baseline 3 is appropriate when schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get a list') and resource ('projects/workspaces'), making the purpose immediately understandable. However, it doesn't distinguish between 'projects' and 'workspaces' or clarify if they're synonymous, and it doesn't differentiate from sibling tools like 'get-project-by-id' or 'list-prompts' which serve different but related purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention when to choose 'list-projects' over 'get-project-by-id' for retrieving specific projects, or 'list-prompts' for different resource types. There are no prerequisites, exclusions, or context for usage decisions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-promptsCInspect

Get a list of Opik prompts

ParametersJSON Schema
NameRequiredDescriptionDefault
pageYesPage number for pagination
sizeYesNumber of items per page

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden for behavioral disclosure. It states the tool retrieves a list but doesn't mention pagination behavior (implied by parameters), rate limits, authentication needs, or what the return format looks like (no output schema). This leaves significant gaps for an agent to understand how the tool behaves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with zero wasted words. It's front-loaded with the core purpose and appropriately sized for a simple list operation, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a list tool with 2 required parameters, no annotations, and no output schema, the description is insufficient. It doesn't explain pagination requirements, return format, or error conditions. Given the complexity (simple but with required params) and lack of structured data, more context is needed for the agent to use this tool effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both parameters ('page' and 'size') clearly documented in the schema. The description adds no additional parameter semantics beyond implying list retrieval, so it meets the baseline score of 3 where the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get a list of Opik prompts' clearly states the verb ('Get') and resource ('Opik prompts'), making the purpose immediately understandable. However, it doesn't distinguish this tool from its sibling 'list-projects' or 'list-traces' beyond the resource type, nor does it specify scope (e.g., all prompts vs. filtered).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like 'get-prompt-by-id' for retrieving a specific prompt, or 'create-prompt' for adding new prompts. There's no mention of prerequisites, typical use cases, or limitations that would help an agent choose appropriately.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-tracesCInspect

Get a list of traces

ParametersJSON Schema
NameRequiredDescriptionDefault
pageYesPage number for pagination
projectIdNoProject ID to filter traces
projectNameNoProject name to filter traces
sizeYesNumber of items per page
workspaceNameNoWorkspace name to use instead of the default

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden but only states the action without disclosing behavioral traits. It doesn't mention if this is a read-only operation, pagination behavior, rate limits, authentication needs, or what the output looks like (e.g., list format, error handling). This leaves significant gaps for a tool with 5 parameters.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence with no wasted words. It's front-loaded and efficiently conveys the core action, making it easy to parse quickly without unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 parameters, no output schema, no annotations), the description is incomplete. It doesn't explain the return values, pagination implications, or how filtering parameters interact, leaving the agent with insufficient context to use the tool effectively beyond basic invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so parameters like 'page', 'size', 'projectId' are well-documented in the schema. The description adds no additional meaning beyond 'list of traces', such as explaining how filtering works or parameter interactions, meeting the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Get a list of traces' states the basic action (get/list) and resource (traces), making the purpose understandable. However, it's vague about what 'traces' are (e.g., execution traces, logging traces) and doesn't distinguish from siblings like 'get-trace-by-id' or 'get-trace-stats', missing specificity for a 4-5 score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives like 'get-trace-by-id' for single traces or 'get-trace-stats' for aggregated data. The description lacks context about filtering capabilities or typical use cases, offering minimal help for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update-projectCInspect

Update a project

ParametersJSON Schema
NameRequiredDescriptionDefault
descriptionNoNew project description
nameNoNew project name
projectIdYesID of the project to update
workspaceNameNoWorkspace name to use instead of the default

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. 'Update a project' implies a mutation operation but doesn't disclose behavioral traits like required permissions, whether changes are reversible, rate limits, or what happens to unspecified fields. It lacks critical context for a write operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with a single sentence, 'Update a project', which is front-loaded and wastes no words. It efficiently communicates the core action, though this brevity contributes to gaps in other dimensions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity as a mutation with 4 parameters and no annotations or output schema, the description is incomplete. It doesn't cover behavioral aspects, usage context, or return values, leaving significant gaps for an AI agent to understand and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters (projectId, name, description, workspaceName). The description adds no meaning beyond what the schema provides, such as explaining interdependencies or default behaviors. Baseline 3 is appropriate when schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Update a project' states the verb and resource but is vague about what aspects can be updated. It distinguishes from siblings like 'create-project' and 'delete-project' by specifying 'update', but doesn't clarify scope compared to 'update-prompt' or differentiate from partial updates vs full replacements.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this tool versus alternatives is provided. The description doesn't mention prerequisites (e.g., needing an existing project), exclusions, or comparisons to siblings like 'get-project-by-id' for read operations or 'list-projects' for discovery.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update-promptCInspect

Update a prompt

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesNew name for the prompt
promptIdYesID of the prompt to update

TDQS

C2.1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure but offers none. It doesn't indicate whether this is a read-only or destructive operation, what permissions are required, whether changes are reversible, what happens on success/failure, or any rate limits. For a mutation tool with zero annotation coverage, this complete lack of behavioral information is inadequate and potentially misleading.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is maximally concise at just three words, with zero wasted language. Every word earns its place by identifying the core action and resource. While this conciseness comes at the expense of completeness, the structure is front-loaded and efficient from a pure brevity perspective.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that this is a mutation tool with no annotations and no output schema, the description is severely incomplete. It doesn't explain what 'updating' entails beyond the name parameter, what happens to other prompt attributes, what the tool returns, or any error conditions. For a tool that modifies data, this minimal description leaves critical gaps in understanding its behavior and outcomes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, with both parameters ('promptId' and 'name') clearly documented in the schema. The description adds no additional parameter information beyond what the schema already provides. According to scoring rules, when schema coverage is high (>80%), the baseline score is 3 even with no parameter information in the description, which applies here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Update a prompt' is a tautology that restates the tool name without adding meaningful context. While it correctly identifies the verb ('update') and resource ('prompt'), it fails to specify what aspects of a prompt can be updated or distinguish this tool from sibling tools like 'update-project'. This minimal statement provides no differentiation or specificity beyond the name itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides absolutely no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing an existing prompt ID), exclusions, or relationships to sibling tools like 'create-prompt', 'delete-prompt', or 'get-prompt-by-id'. Without any context about appropriate usage scenarios, the agent has no basis for making informed decisions about tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 19 tool updatesv1.0.0
    • First observedcreate-project
    • First observedcreate-prompt
    • First observedcreate-prompt-version
    • First observeddelete-project
    • First observeddelete-prompt
    • First observedget-metrics
    • First observedget-opik-examples
    • First observedget-opik-help
    • First observedget-opik-tracing-info
    • First observedget-project-by-id
    • First observedget-prompt-by-id
    • First observedget-server-info
    • First observedget-trace-by-id
    • First observedget-trace-stats
    • First observedlist-projects
    • First observedlist-prompts
    • First observedlist-traces
    • First observedupdate-project
    • First observedupdate-prompt

TDQS

B3.1/5.0

Scored across 19 tools

Disambiguation5/5

Each tool has a clearly distinct purpose targeting specific resources (projects, prompts, traces, metrics, help) and actions (create, get, list, update, delete). There is no overlap or ambiguity between tools, making it easy for an agent to select the right one for any operation.

Naming Consistency5/5

All tools follow a consistent verb_noun pattern with hyphens (e.g., create-project, get-project-by-id, list-projects). The naming is uniform across all 19 tools, using clear verbs like create, get, list, update, and delete paired with specific nouns.

Tool Count4/5

With 19 tools, the count is slightly high but reasonable for a comprehensive server covering projects, prompts, traces, metrics, and help. It feels well-scoped for the Opik domain, though it borders on being heavy compared to typical 3-15 tool ranges.

Completeness5/5

The tool set provides complete CRUD/lifecycle coverage for projects and prompts (create, get, list, update, delete), with additional tools for traces, metrics, and help. There are no obvious gaps; agents can perform all core operations without dead ends in this domain.

Maintenance

ActivityActive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    B
    quality
    D
    maintenance
    Implements the Model Context Protocol (MCP) to provide AI models with a standardized interface for connecting to external data sources and tools like file systems, databases, or APIs.
    1
    153
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    An implementation of the Model Context Protocol (MCP) that enables interaction with debug adapters, allowing language models to control debuggers, set breakpoints, evaluate expressions, and navigate source code during debugging sessions.
    40
    AGPL 3.0
  • F
    license
    Not graded
    quality
    D
    maintenance
    A Python-based implementation of the Model Context Protocol that enables communication between a model context management server and client through a request-response architecture.
    -
  • -
    license
    Not graded
    quality
    D
    maintenance
    A standardized foundation for building Model Context Protocol servers that integrate with VS Code, using Python with stdio transport for seamless AI tool integration.
    -