CARBON Studio Pro
Integrates with the OpenAI plugin framework (ChatGPT/Codex) by implementing @openai/mcp-extensions and the MCP Apps UI resource contract, enabling native visual quality workspace entrypoints, structured settings, live activity, host message handoff, and model-context selection when supported.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@CARBON Studio Prorun my test suite and show the live evidence workspace"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
CARBON Studio Pro · Codex edition preview
This separate full-featured Codex edition's first-party source and bundled CARBON runtime use PolyForm Perimeter 1.0.0. See NOTICE.md. Internal business use is permitted; providing competing products is restricted. This is source-available, not unrestricted open source. Dependency licenses remain unchanged. No private benchmark dataset or hosted-service account is granted.
CARBON Lite is the smaller, skills-only edition intended for submission first. This repository does not replace, hide, withdraw, or update the original CARBON Studio marketplace submission.
A native visual quality workspace for the new OpenAI plugin framework. Built by testers.ai. Separate from CARBON Light and Pro; neither is modified or replaced.
What works
Global app and thread-panel entrypoints, plus a dedicated settings entrypoint.
Native structured settings and an interactive custom settings screen, backed by the same local store.
Live activity, elapsed time, evidence-qualified domain charts, screenshot maps, input/behavior stacks, findings and reproduction steps.
Human priorities and comments persisted against findings and available to the agent.
Host message handoff and model-context selection when supported; an honest copy-prompt fallback otherwise.
Dark/light/system themes, reduced motion, keyboard-accessible dialogs and compact phone layout.
Strict evidence requirements, atomic storage, conflict detection, private screenshot containment, isolated loopback preview.
Studio now bundles 54 maintained CARBON command workflows and their local MCP runtime, in addition to the Studio entrypoint. Security, privacy, accessibility, issues, map, confidence and the other specialized skills open Studio first and retain their detailed CARBON reports. Command availability does not mean every workflow has been end-to-end acceptance-tested in Studio. Proprietary benchmarks and the separate hosted feedback service are not included.
Related MCP server: unfour
Install locally
Requires Node.js 22+. The distributable already bundles its runtime and UI; no npm install is needed to use it.
With the personal marketplace entry created for this workspace:
codex plugin add carbon-studio-pro@personalStart a new Codex conversation after installation. Ask: “Use CARBON Studio Pro to test this project and show the live workspace.” The skill coordinates the host agent's actual testing tools; the MCP server records and visualizes evidence rather than pretending to execute tests itself.
Open Quality workspace from the plugin's global/thread entrypoint on supported hosts, or invoke studio_open. Open Testing preferences or invoke studio_preferences for settings. Every open call also returns a private live localhost URL: the skill opens this beside chat if the native panel cannot be confirmed. Opening a workspace alone does not start testing.
Connection health is a small focusable dot. Hover, focus or tap it for details, including the last received update. Startup, background pause and lost connection are distinct; failed refreshes keep the last good evidence, retry with backoff, and never change a test result.
Host boundaries
Implements the official @openai/mcp-extensions 0.1.0 and MCP Apps UI resource contract. The initial SDK declares @modelcontextprotocol/ext-apps ^1.7.5 compatibility; this package pins 1.7.5 rather than forcing 2.x. Zod is deduplicated to 4.4.3 so settings metadata survives schema conversion.
Native entrypoint availability depends on the installed host version and negotiated extensions. Protocol validation is not proof of native-host visual acceptance. Classic ChatGPT web is not assumed to support desktop local stdio plugins. A public ChatGPT App deployment would be a separate hosted, authenticated product with its own review and data-boundary design—not a reason to expose this local store publicly.
Reference: https://developers.openai.com/plugins/build/extensions and https://developers.openai.com/plugins/build/chatgpt-ui
Data and safety
State is stored locally in ~/.carbon-studio-pro/workspace.json (override with CARBON_STUDIO_HOME). Files are written with owner-only permissions. The Studio visual workspace makes no analytics or model calls; the bundled CARBON server is configured with analytics disabled. Specialized workflows can access the selected test target or explicitly configured services under their permission gates. The coding agent still sends the context it uses to its configured provider. This is not an air-gap guarantee or compliance certification.
Only a project explicitly selected by the agent is attached. Screenshot reads require a real path inside that project, a supported image signature, and a 5 MB limit. Screenshots are supplied to the UI in _meta, not in the model's structured text. All evidence text is escaped before rendering. The MCP surface is local stdio; the HTTP fallback starts lazily on a workspace/settings open call (or with --preview), binds loopback, and requires a private capability token plus Origin/Host checks for data access.
Budgets guide the agent; they do not enforce a host token quota. Studio never authorizes destructive testing or remediation by itself. Closing a run preserves its record. Evidence is retained until the owner removes the local workspace; there is no cloud retention service. Back up the state file before manual cleanup.
Develop and preview
Ad hoc browser checks and control panel
Ad hoc exploration uses the coding agent's built-in browser by default. Automatic does not launch Selenium, Playwright, CDP, or Vibium as a fallback. A missing host capability is reported before asking for an alternative. Explicit permitted controller preferences and requested existing framework suites remain supported. This is a host-agent routing contract, not a machine-level enforcement sandbox.
The overview separates executed checks from blocked/deferred work, exposes a compact test queue, and ranks demonstrated findings before suspicions. Empty capture panels are omitted. Evidence-map inputs share the same status semantics. Preferences guide future runs; saved guidance is available on the agent's next read, not proof of immediate execution. Mobile navigation retains text labels.
npm ci
npm run build
npm test
node scripts/seed-demo.mjs output/my-demo-state
CARBON_STUDIO_HOME="$PWD/output/my-demo-state" npm run previewOpen the private URL printed by the preview command. Keep the fragment; it is the access capability. The demo is intentionally labeled synthetic and isolated from the default workspace. It is not a benchmark or a real quality assessment. To attach a screenshot, use studio_screenshot in a real run, or call the same Store method in a local fixture script.
Current limitations
Snapshot polling currently reloads all stored runs. Large histories and many large images need pagination/lazy resources before production-scale use.
An abruptly terminated writer may leave
write.lock; do not remove it until all Studio processes using that store have stopped.Native app rendering, message sending, and model context must be acceptance-tested in each target host build. Browser preview tests do not establish ChatGPT/Codex host parity.
The user interface is a functional foundation, not an automatic migration of all historic CARBON reports.
Packaging
node scripts/package.mjs builds a distribution-only staging directory and ZIP under output/release/. Source, lockfile, notices, assets, skills and bundled runtime are included. User state, screenshots, tests, node_modules and private CARBON benchmark data are excluded.
Available Tools
11 toolssettings.readDRead-only
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| layout | No | |
| schema | Yes | |
| values | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Tool has no description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tool has no description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tool has no description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tool has no description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
settings.updateD
| Name | Required | Description | Default |
|---|---|---|---|
| set | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| values | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Tool has no description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Tool has no description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Tool has no description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Tool has no description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Tool has no description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
studio_feedbackB
Read or save personal report feedback for this OS account and project. Hiding never alters test evidence or confidence. Read returned preferences before further investigations.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| operation | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and destructiveHint=false, and the description adds real value by clarifying that hiding 'never alters test evidence or confidence' — important given 'hide' sounds destructive. It still omits the write-side mechanics: the required revision field implies optimistic concurrency, and nothing explains what happens if the operation object is omitted or if a revision is stale.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with no filler, and the read-or-save scope is front-loaded. It is efficiently sized, though the final sentence blends usage guidance with the behavioral note in a slightly ambiguous way.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with a nested object parameter, enum-driven actions, and no output schema, the description is thin. It says nothing about return shape, how votes/rules differ from hides, or how revision conflicts are handled — all things an agent needs before invoking a write path.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the schema is a nested operation object with five required sub-fields (key, action, value, reason, revision) plus an enum of vote/hide/rule. The description only loosely gestures at these through the words 'feedback,' 'save,' and 'hiding'; it never explains runId, the action types, the value semantics, or the revision counter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a concrete verb pair and resource: 'Read or save personal report feedback,' scoped to 'this OS account and project.' It is understandable on its own, but it never distinguishes itself from the close sibling studio_preferences, so an agent may hesitate between the two.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Read returned preferences before further investigations' gives one sequencing hint (read first), which is genuinely useful. However there is no guidance on when to save versus read, when the operation object is required, or why this tool should be chosen over studio_preferences.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
studio_openQuality workspaceARead-only
Open the CARBON Studio Pro visual quality workspace. Does not execute tests.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already tells the agent this is a non-mutating operation, so the description's main added value is the scope note that no tests are executed. It says nothing about whether opening a workspace has UI/session side effects or what state it leaves behind.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler; the core action is front-loaded and the scope qualifier follows immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters, no output schema, and annotations covering the safety profile, the description is nearly sufficient. It could still state what opening the workspace yields or requires, but little is missing for a zero-argument tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is no parameter semantics to document and the baseline of 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb (Open) and a specific resource (CARBON Studio Pro visual quality workspace), so the action is unambiguous. It does not, however, differentiate itself from siblings like studio_start or studio_snapshot beyond the negative statement about tests.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Does not execute tests" implies a boundary against studio_start, giving an implicit when-not-to-use cue. There is no positive guidance on when to open the workspace or how it relates to studio_preferences/studio_save_settings, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
studio_preferencesTesting preferencesCRead-only
Open the CARBON Studio Pro visual quality workspace. Does not execute tests.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes this is a safe, non-mutating operation, and the description is consistent with that. It adds one useful behavioral fact ('Does not execute tests'), but says nothing about whether it changes workspace state, requires auth, or what it returns. Some added value over the annotation, but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and no filler. However, the second sentence's value is undermined by the tool's 'Testing preferences' title, which makes the 'does not execute tests' disclaimer feel disconnected rather than clarifying.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters, no output schema, and a read-only annotation, the structured fields carry little, yet the description still omits what 'preferences' means, whether any state changes, and how it differs from studio_open. The name/title/description mismatch leaves an agent unable to confidently select the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies. Schema coverage is trivially 100% and the empty input schema is self-explanatory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a verb and resource ('Open the CARBON Studio Pro visual quality workspace'), but it conflicts with the tool name (studio_preferences) and title ('Testing preferences'), which reference preferences/testing rather than visual quality. It also fails to distinguish itself from the sibling studio_open, which appears to do a similar 'open' action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage signal is a negative constraint ('Does not execute tests'), with no statement of when to use this tool, prerequisites, or how it relates to the overlapping studio_open and settings.read siblings. An agent has no basis for choosing it over alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
studio_save_settingsC
Save testing preferences with optimistic concurrency protection. Budget values guide the agent; they are not hard host token limits.
| Name | Required | Description | Default |
|---|---|---|---|
| patch | Yes | ||
| revision | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this as a non-destructive write in a closed world, so the description's main added value is the revision-based concurrency model and the clarification that budget values are advisory rather than hard token limits. It still does not say whether 'patch' merges with or replaces existing settings, which matters for a write tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, and the concurrency guarantee is front-loaded alongside the core action. Efficient, though the second sentence addresses only one of the seven patch fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a mutation tool with a nested 7-field object at 0% schema coverage and no output schema. The description should carry the burden of explaining merge semantics, revision conflict behavior, and the individual preference fields, but it covers only concurrency and budgets.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the patch object holds seven undocumented fields, so the description must compensate. It clarifies only 'budget values' (budgetMinutes) and implies revision's role; theme, motion, askHuman, maxChecks, businessWeight, and explorationPercent remain unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Save testing preferences') plus a distinctive mechanism ('optimistic concurrency protection'), so the operation is clear. However, it never distinguishes itself from nearby siblings such as studio_preferences or settings.update, leaving overlap ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this versus studio_preferences or settings.update, nor any prerequisite (e.g., fetch current revision first). The only guidance is a caveat about budget semantics, which is not usage routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
studio_screenshotC
Attach an actual PNG, JPEG, or WebP screenshot inside the selected project. No arbitrary file access.
| Name | Required | Description | Default |
|---|---|---|---|
| file | Yes | ||
| runId | Yes | ||
| pageId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint=false, openWorldHint=false, and destructiveHint=false, indicating a safe, local write operation. The description adds the constraint 'No arbitrary file access,' which is useful context beyond annotations. However, it doesn't disclose details like whether the attachment is reversible, required permissions, or any rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no waste; the main action is front-loaded. However, the second sentence is somewhat terse and could be integrated more smoothly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter mutation tool with no output schema and 0% schema description coverage, the description is incomplete. It omits why parameters are needed, prerequisites, and any failure modes, which an agent would need to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for all three required parameters (runId, pageId, file). The description only names formats (PNG, JPEG, WebP) but does not explain the meaning of runId, pageId, or the expected format for file, leaving parameters entirely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (attach) and resource (PNG/JPEG/WebP screenshot inside the selected project), which is clear. It doesn't explicitly differentiate from siblings like studio_snapshot or studio_save_settings, so it lacks sibling differentiation for a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives such as studio_snapshot or studio_preferences. The usage context is implied only by 'inside the selected project' and 'No arbitrary file access,' which is insufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
studio_snapshotARead-only
Read current runs, evidence, settings, and human steering. Poll this data tool rather than reopening the UI. sinceRevision avoids retransmitting unchanged snapshots.
| Name | Required | Description | Default |
|---|---|---|---|
| sinceRevision | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds genuinely useful context beyond them: this is a polling data source whose sinceRevision argument enables incremental updates instead of full retransmission. It does not describe error behavior or the full-snapshot case when sinceRevision is omitted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences that lead with what is read, then how to use it, then the key parameter. The final clause ("sinceRevision avoids retransmitting unchanged snapshots") is a fragment, but it is dense and earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must convey what comes back and it does by enumerating runs, evidence, settings, and human steering. Combined with the sinceRevision polling note, an agent has enough to call it correctly; only the default/omitted-revision behavior is unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only one integer parameter and 0% schema description coverage, the description carries the burden and does explain it: sinceRevision prevents retransmitting unchanged snapshots, which conveys the incremental-pull semantics. It omits where to obtain the revision value (presumably from a prior response) and what happens when it is omitted, so it compensates only partially.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Read") and enumerates the resources returned: current runs, evidence, settings, and human steering. An agent can identify it as the snapshot/state reader. It does not, however, differentiate itself from the overlapping sibling settings.read, which also implies reading settings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Poll this data tool rather than reopening the UI" gives an explicit usage directive and names the alternative behavior (reopening the UI). It stops short of telling the agent when to prefer this over the sibling settings.read, so no exclusions are given, matching a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
studio_startC
Create a durable testing run for an explicitly selected local project. The coding agent must perform checks and report observations; this tool does not test the app itself.
| Name | Required | Description | Default |
|---|---|---|---|
| root | Yes | ||
| title | Yes | ||
| target | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds that the run is 'durable' and that the tool itself performs no testing, which is genuinely useful behavioral framing, but it says nothing about side effects, persistence scope, or whether the call blocks.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the main action front-loaded and the scope caveat immediately after. No filler, though the second sentence is more workflow advice than tool definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-required-parameter mutation tool with no output schema and zero parameter documentation, the description leaves too much unspecified — the meaning of target, what a run produces, and how it relates to sibling studio tools are all absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for all three required parameters (root, title, target), and the description does not compensate — only 'explicitly selected local project' loosely gestures at root, while title and target are entirely unexplained. The agent cannot determine what target should contain.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a clear verb+resource — 'Create a durable testing run' — and adds a useful negative qualifier ('this tool does not test the app itself'). However, it never distinguishes itself from siblings like studio_open, studio_snapshot, or studio_update, so the agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies a precondition ('an explicitly selected local project') and a workflow expectation (the coding agent must perform checks and report observations), but gives no explicit when-to-use or when-not-to-use guidance and never names an alternative tool among the many studio_* siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
studio_steerB
Persist user guidance on a finding. The agent reads it before selecting the next check.
| Name | Required | Description | Default |
|---|---|---|---|
| note | Yes | ||
| runId | Yes | ||
| priority | Yes | ||
| findingId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish that this is a non-read-only, non-destructive, closed-world write. The description adds that the persisted guidance is later read by the agent before choosing the next check, which is useful behavioral context beyond the annotations, but it does not cover permissions, durability, or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, both front-loaded and free of filler. The purpose is stated first, followed immediately by the key behavioral consequence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with four required parameters, no schema descriptions, and no output schema, the description is too sparse. It communicates the core purpose but omits parameter meaning, priority semantics, and invocation context needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and all four parameters are required, so the description carries the burden of explaining them. It only vaguely hints that the guidance concerns a finding, leaving runId, priority, note limits, and the relationship between parameters unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Persist user guidance on a finding.' It distinguishes the tool from generic update/save siblings by describing the artifact persisted and its role, but it does not explicitly name an alternative tool or scope boundaries.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains that the agent reads the guidance before selecting the next check, which is behavioral context rather than invocation guidance. It gives no explicit when-to-use, when-not-to-use, or alternative sibling comparison, leaving the agent to infer the trigger from the purpose alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
studio_updateB
Record actual test observations, findings, progress, and concise decision rationale. Do not provide private chain of thought. Requires current run revision.
| Name | Required | Description | Default |
|---|---|---|---|
| why | No | ||
| pages | No | ||
| runId | Yes | ||
| checks | No | ||
| status | No | ||
| current | No | ||
| summary | No | ||
| blockers | No | ||
| findings | No | ||
| journeys | No | ||
| revision | Yes | ||
| confidence | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false (a write) and destructiveHint=false, so the safety profile is covered. The description adds useful behavioral context beyond that: the revision prerequisite hints at optimistic-concurrency control, and the 'no private chain of thought' rule is a real content constraint. It still omits what happens on a stale revision or whether writes append vs replace.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences: purpose first, then a content prohibition, then the prerequisite. No wasted wording. It is efficient, though the terseness is partly a symptom of under-specification rather than disciplined brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a complex write tool with 12 parameters, deeply nested objects (checks, findings, journeys, pages, confidence), enums, and no output schema. The description is far too thin to guide correct construction of these structures; it does not explain any of the nested object shapes or the enum-driven status fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 12 parameters at 0% schema description coverage, the description carries the full burden, but it only touches a few concepts (findings, progress, decision rationale) and the revision prerequisite. Core parameters like runId, runId format, pages, checks, journeys, status, blockers, and confidence are entirely undocumented anywhere.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resources: 'Record actual test observations, findings, progress, and concise decision rationale.' This is far clearer than the generic name 'studio_update' and helps distinguish it from siblings like studio_snapshot or studio_steer. However, it never explicitly names or contrasts a sibling, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'Record actual test observations, findings, progress' and a prerequisite is given ('Requires current run revision'), plus a content restriction ('Do not provide private chain of thought'). But there is no explicit statement of when to use this versus alternatives like studio_feedback or studio_steer, leaving routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
11 tool updates
v0.1.0- First observed
settings.read - First observed
settings.update - First observed
studio_feedback - First observed
studio_open - First observed
studio_preferences - First observed
studio_save_settings - First observed
studio_screenshot - First observed
studio_snapshot - First observed
studio_start - First observed
studio_steer - First observed
studio_update
TDQS
Scored across 11 tools
studio_open and studio_preferences share an identical description, making them indistinguishable, and studio_save_settings, settings.update, and settings.read all appear to write the same preferences. studio_feedback vs studio_save_settings also overlap. Several tools have unclear boundaries despite good individual descriptions elsewhere.
Eight tools use a consistent studio_verb/noun snake_case pattern, but settings.read and settings.update break the convention with dot notation and no studio prefix. Mixing two naming schemes on a single server is a clear inconsistency.
11 tools is a reasonable scope for a testing workspace lifecycle. However, the presence of redundant pairs (open/preferences, save_settings/settings.update) inflates the count slightly beyond what is earned.
The surface covers the full run lifecycle: opening the workspace, starting a run, snapshotting state, recording observations, attaching screenshots, saving settings, feedback, and steering. Main gap is that settings read/update duplicate existing functionality rather than filling a real hole.
Maintenance
Related MCP Connectors
Cross-agent artifact workspace with provenance across Claude Code, Codex, Cursor, LangGraph.
- HydrantOAuthdev.hydrant
Track issues, projects and dependencies with your agent in a workspace you control.
Machine-native research commons for agent evidence, discovery, rooms, and bounded research quests.
Coding agents build full-stack apps in persistent workspaces and share them by link.
Related MCP Servers
- FlicenseBqualityDmaintenanceStructured workspace runtime for long-running coding agents, providing controlled workspace capabilities with task state, snapshots, checkpoints, drift detection, verification evidence, audit logs, and structured handoff.20-
- AlicenseBqualityAmaintenanceLocal-first backend developer workspace exposing API debugging, SSH, database, workspace, and diagnostics tools to AI agents through a local MCP server722Apache 2.0
- AlicenseNot gradedqualityCmaintenanceProvides AI coding agents with low-latency project intelligence, including workspace setup, encrypted backups, diagnostics, evolution digests, and cross-project pattern scoring from local SQLite data.MIT
- AlicenseNot gradedqualityBmaintenanceProvides coding agents with local tools to inspect project UI inventories, propose and compare visual direction boards, compile versioned design contracts and DTCG tokens, retrieve section-specific blueprints, and audit running interfaces with browser evidence including screenshots, accessibility findings, and overflow measurements.MIT