Skip to main content
Glama

Soma MCP Server

为你的 AI 智能体提供它自己无法做到的一件事:真正针对测试运行代码,并证明它通过了。

Soma 是一个执行验证的代码服务。这个 MCP 服务器暴露两个工具:

  • soma_verify_code — 在隔离沙箱中针对测试运行候选代码;获得 PASS/FAIL 判定,外加一个签名的、可离线验证的证书(Ed25519)。在信任代码之前,用它来独立确认代码是否有效。

  • soma_generate_verified_code — 让 Soma 为任务编写代码;当任务可验证时,返回的代码已经针对派生测试执行过,并附有证书。

支持 15+ 种语言的验证(Python、JavaScript/TypeScript、Go、C/C++、Java、Rust、Ruby、PHP、Bash 等)。

安装

需要 Node.js 18+。通过 stdio 运行。

添加到你的 MCP 客户端配置(Claude Desktop、Cursor 等):

{
  "mcpServers": {
    "soma": {
      "command": "npx",
      "args": ["-y", "soma-verify-mcp"],
      "env": {
        "SOMA_API_KEY": "YOUR_SOMA_KEY"
      }
    }
  }
}
  • Claude Desktop:Settings → Developer → Edit Config,添加上面的块,然后重启。

  • Cursor:Settings → MCP → Add,或者将同样的块放入 ~/.cursor/mcp.json。

Related MCP server: OmniBridge MCP Server

配置

环境变量

必需

默认值

用途

SOMA_API_KEY

是

—

你的 Soma API 密钥。

SOMA_BASE_URL

否

https://170-9-236-56.sslip.io

Soma API 基础 URL。

SOMA_TIMEOUT_MS

否

300000

每次请求的超时时间。

获取免费预览密钥:联系 centrum.arvind@gmail.com(预览期间免费)。

工具

soma_verify_code

针对测试运行代码并返回签名判定。

  • language(字符串)— 例如 python、javascript、go、rust。

  • code(字符串)— 要验证的完整源代码。

  • tests(数组)— 以下之一:

    • 函数模式(默认):[{ "input": [arg1, arg2], "expected": value }] 加上 entrypoint(函数名)。

    • stdio 模式:设置 mode: "stdio" 和 [{ "stdin": "...", "expected_stdout": "..." }];不需要 entrypoint。

  • entrypoint(字符串,可选)— 函数模式下的函数名。

  • mode("function" | "stdio",可选)。

返回:verdict、tests_passed、tests_total,以及你可以离线检查的 signature / public_key / sig_alg。

soma_generate_verified_code

获取任务的代码,在返回之前已针对派生测试执行过。

  • prompt(字符串)— 编码任务。包含具体的输入/输出示例(例如 >>> f(2) == 4),这样结果就是可验证的,而不是尽力而为。

  • max_tokens(整数,可选,默认 1500)。

返回:代码、certified(布尔值),以及验证通过时的 certificate(verdict、tests_passed、tests_total)。如果任务不可验证,输出会以未认证的方式返回并明确标注 — 绝不会出现虚假的“已验证”。

证书的含义

证书证明列出的测试在生成时于隔离沙箱中通过了。它是签名的(Ed25519),并且可以针对返回的公钥离线检查。它不是任何用途适用性的保证 — 在生产使用前请审查输出。

隐私

不会对你的提示进行训练。请参阅 ${SOMA_BASE_URL}/privacy 处的 Soma 隐私与数据政策。

许可证

MIT。# Soma MCP Server

给你的 AI 代理一个它自己做不到的东西:真正运行代码来测试,并证明它通过了。

Soma 是一个执行验证的代码服务。这个 MCP 服务器暴露两个工具:

  • soma_verify_code — 在隔离沙箱中运行候选代码并针对测试进行验证;获得 PASS/FAIL 判定,外加一个签名的、可离线验证的证书(Ed25519)。在信任代码之前,用它来独立确认代码是否有效。

  • soma_generate_verified_code — 让 Soma 为任务编写代码;当任务可验证时,返回的代码已经针对派生测试执行过,并附有证书。

支持 15+ 种语言的验证(Python、JavaScript/TypeScript、Go、C/C++、Java、Rust、Ruby、PHP、Bash 等)。

安装

需要 Node.js 18+。通过 stdio 运行。

添加到你的 MCP 客户端配置(Claude Desktop、Cursor 等):

{
  "mcpServers": {
    "soma": {
      "command": "npx",
      "args": ["-y", "soma-verify-mcp"],
      "env": {
        "SOMA_API_KEY": "YOUR_SOMA_KEY"
      }
    }
  }
}
  • Claude Desktop:Settings → Developer → Edit Config,添加上面的块,然后重启。

  • Cursor:Settings → MCP → Add,或者将同样的块放入 ~/.cursor/mcp.json。

配置

环境变量

必需

默认值

用途

SOMA_API_KEY

是

—

你的 Soma API 密钥。

SOMA_BASE_URL

否

https://170-9-236-56.sslip.io

Soma API 基础 URL。

SOMA_TIMEOUT_MS

否

300000

每次请求的超时。

获取免费预览密钥:联系 centrum.arvind@gmail.com(预览期间免费)。

工具

soma_verify_code

运行代码并返回签名的判定结果。

  • language(字符串)— 例如 python、javascript、go、rust。

  • code(字符串)— 要验证的完整源代码。

  • tests(数组)— 以下之一:

    • 函数模式(默认):[{ "input": [arg1, arg2], "expected": value }] 加上 entrypoint(函数名)。

    • stdio 模式:设置 mode: "stdio" 和 [{ "stdin": "...", "expected_stdout": "..." }];不需要 entrypoint。

  • entrypoint(字符串,可选)— 函数模式下的函数名。

  • mode("function" | "stdio",可选)。

返回:verdict、tests_passed、tests_total,以及一个你可以离线检查的 signature / public_key / sig_alg。

soma_generate_verified_code

获取任务的代码,在返回之前已针对派生测试执行过。

  • prompt(字符串)— 编码任务。包含具体的输入/输出示例(例如 >>> f(2) == 4),这样结果就是可验证的,而不是尽力而为。

  • max_tokens(整数,可选,默认 1500)。

返回:代码、certified(布尔值),以及验证通过时的 certificate(verdict、tests_passed、tests_total)。如果任务不可验证,输出会以未认证的方式返回,并明确标注 — 绝不会出现虚假的“已验证”。

证书的含义

证书证明列出的测试在生成时于隔离沙箱中通过了。它是签名的(Ed25519),并且可以针对返回的公钥离线检查。它不是任何用途的适用性保证 — 在生产使用前请审查输出。

隐私

不会对你的提示进行训练。请参阅 ${SOMA_BASE_URL}/privacy 处的 Soma 隐私与数据政策。

许可证

MIT。

Available Tools

2 tools
soma_generate_verified_codeGenerate execution-verified codeA
Read-only

Ask Soma to write code for a task. When the task is verifiable, the returned code has been executed against derived tests in an isolated sandbox before it is returned, and a certificate (verdict + tests passed) is attached. Include concrete examples in the prompt (e.g. doctest-style '>>> f(2) == 4') to make the result verifiable rather than best-effort.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesThe coding task. Include example input/output pairs to enable verification.
max_tokensNoMaximum output tokens (default 1500).

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and destructiveHint=false, indicating a safe, non-destructive operation. The description adds valuable behavioral context: the code is executed in an isolated sandbox, a certificate with verdict and tests passed is attached, and the verification is conditional on task verifiability. This goes beyond the annotations and provides important expectations about the output and process.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured. It front-loads the core purpose, then explains the verification process and provides actionable advice on prompt construction. Every sentence adds value, and there is no fluff or repetition. The structure is logical: what it does, how it works, and how to use it effectively.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (verification, sandbox, certificate) and the absence of an output schema, the description covers the essential aspects: what the tool does, how verification works, and how to improve results. It does not detail the certificate format or the exact conditions for verifiability, but these are not critical for an agent to invoke the tool correctly. The description is complete enough for effective use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (prompt and max_tokens) are already documented in the schema. The description adds guidance on how to use the prompt parameter (include concrete examples) and mentions the default for max_tokens, but this is largely redundant with the schema. The description does not add significant new meaning beyond the schema, so a baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Ask Soma to write code for a task.' It specifies the key differentiator—execution-verified code with a certificate—and distinguishes it from a best-effort generation. The verb 'write code' and resource 'task' are specific, and the mention of verification sets it apart from the sibling tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear guidance on when to use this tool: when the task is verifiable, and it advises including concrete examples in the prompt to enable verification. It implies that for non-verifiable tasks, this tool may not be ideal, but it does not explicitly name the sibling tool or state when to use soma_verify_code instead. The guidance is strong but lacks an explicit exclusion or alternative mention.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

soma_verify_codeVerify code against tests (signed certificate)A
Read-onlyIdempotent

Run candidate code against tests inside an isolated sandbox and return a PASS/FAIL verdict with a signed, offline-checkable certificate (Ed25519). Use this to independently confirm that code actually works before trusting it.

Two test shapes:

  • function mode (default): tests = [{"input": [arg1, arg2], "expected": value}] and provide 'entrypoint' (the function name).

  • stdio mode: set mode='stdio' and tests = [{"stdin": "...", "expected_stdout": "..."}]; no entrypoint needed.

Supported languages include python, javascript, typescript, go, c, cpp, java, rust, ruby, php, bash and more.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesThe complete source code to verify.
modeNo'function' (default) or 'stdio'.
testsYesfunction mode: [{input:[args], expected: value}]. stdio mode: [{stdin:'...', expected_stdout:'...'}].
languageYesProgramming language, e.g. 'python', 'javascript', 'go', 'rust'.
entrypointNoFunction name to call (required for function mode; omit for stdio mode).

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description reveals the isolated sandbox behavior and the signed certificate output, adding value beyond the annotations. It does not specify sandbox limits, timeouts, or what happens on ambiguous test results, but the key isolation and non-destructive nature are explicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well-organized with clear separation of modes, languages, and purpose. Slightly verbose but every line adds value. Could be tightened, but readability is strong.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers when to use (independent confirmation), the two modes, supported languages, and the output certificate nature. Missing: no example usage, no details on error handling or timeouts, and output format isn't specified, but the description is otherwise informative enough for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema documents all five parameters with types and required status, and the description adds concrete examples for tests structures in both modes. Entry point requirement is clarified. Coverage is 100% and the semantics are unambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool verifies code against tests in a sandbox and returns a signed certificateat. Uses specific verbs (verify, return PASS/FAIL) and names two precise modes. Though siblings weren't supplied for comparison, it self-defines scope well.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear usage context: 'Use this to independently confirm that code actually works before trusting it.' Describes two modes (function and stdio) with examples. Does not explicitly list when-not-to-use or alternatives, but the context is enough for most agents.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 2 tool updatesv0.1.0
    • First observedsoma_generate_verified_code
    • First observedsoma_verify_code

TDQS

A4.3/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have clearly distinct purposes: one verifies provided code, the other generates code with verification. There is no overlap or confusion between them.

Naming Consistency5/5

Both tools follow the identical pattern of 'soma_' prefix plus a verb_noun structure: verify_code and generate_verified_code. Naming is uniform and predictable.

Tool Count3/5

With only 2 tools, the server is minimal but focused on a narrow purpose. While the count feels thin for a general toolkit, it is reasonable for a specialized verification service. It sits at the borderline.

Completeness4/5

The server covers the core workflows of code verification and generation with verification. Minor gaps exist, such as no ability to fetch or check certificate details after generation, but the primary lifecycle is covered.

Maintenance

ActivitySlowing
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables secure cloud-based execution of code across 14+ programming languages within a sandboxed environment. It supports file management, standard input/output handling, and automatic generation of visual artifacts like plots and charts.
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides isolated sandbox environments for AI agents to execute code securely, generating signed receipts for every execution to ensure auditability and trust.
    24 npm
    4
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Provides an isolated workspace for testing candidate code, runs tests, and returns deterministic pass/fail verdicts. Enables automated grading of software engineering solutions by ensuring reproducible test runs.
    5
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Provides structured, sandboxed test and lint feedback for coding agents, returning compact typed verdicts with failure fingerprints instead of raw runner output. It enables impact-selected test execution in isolated containers, distinguishing pre-existing failures from regressions.
    4
    Apache 2.0