CARBON Studio Pro
# CARBON Studio Pro · Codex edition preview
This separate full-featured Codex edition's first-party source and bundled CARBON
runtime use [PolyForm Perimeter 1.0.0](LICENSE). See [NOTICE.md](NOTICE.md).
Internal business use is permitted; providing competing products is restricted.
This is source-available, not unrestricted open source. Dependency licenses remain
unchanged. No private benchmark dataset or hosted-service account is granted.
[CARBON Lite](https://github.com/jarbon/carbon-lite) is the smaller, skills-only
edition intended for submission first. This repository does not replace, hide,
withdraw, or update the original CARBON Studio marketplace submission.
A native visual quality workspace for the new OpenAI plugin framework. Built by testers.ai. Separate from CARBON Light and Pro; neither is modified or replaced.
## What works
- Global app and thread-panel entrypoints, plus a dedicated settings entrypoint.
- Native structured settings and an interactive custom settings screen, backed by the same local store.
- Live activity, elapsed time, evidence-qualified domain charts, screenshot maps, input/behavior stacks, findings and reproduction steps.
- Human priorities and comments persisted against findings and available to the agent.
- Host message handoff and model-context selection when supported; an honest copy-prompt fallback otherwise.
- Dark/light/system themes, reduced motion, keyboard-accessible dialogs and compact phone layout.
- Strict evidence requirements, atomic storage, conflict detection, private screenshot containment, isolated loopback preview.
Studio now bundles 54 maintained CARBON command workflows and their local MCP runtime, in addition to the Studio entrypoint. Security, privacy, accessibility, issues, map, confidence and the other specialized skills open Studio first and retain their detailed CARBON reports. Command availability does not mean every workflow has been end-to-end acceptance-tested in Studio. Proprietary benchmarks and the separate hosted feedback service are not included.
## Install locally
Requires Node.js 22+. The distributable already bundles its runtime and UI; no npm install is needed to use it.
With the personal marketplace entry created for this workspace:
```sh
codex plugin add carbon-studio-pro@personal
```
Start a new Codex conversation after installation. Ask: **“Use CARBON Studio Pro to test this project and show the live workspace.”** The skill coordinates the host agent's actual testing tools; the MCP server records and visualizes evidence rather than pretending to execute tests itself.
Open **Quality workspace** from the plugin's global/thread entrypoint on supported hosts, or invoke `studio_open`. Open **Testing preferences** or invoke `studio_preferences` for settings. Every open call also returns a private live localhost URL: the skill opens this beside chat if the native panel cannot be confirmed. Opening a workspace alone does not start testing.
Connection health is a small focusable dot. Hover, focus or tap it for details, including the last received update. Startup, background pause and lost connection are distinct; failed refreshes keep the last good evidence, retry with backoff, and never change a test result.
## Host boundaries
Implements the official `@openai/mcp-extensions` 0.1.0 and MCP Apps UI resource contract. The initial SDK declares `@modelcontextprotocol/ext-apps ^1.7.5` compatibility; this package pins 1.7.5 rather than forcing 2.x. Zod is deduplicated to 4.4.3 so settings metadata survives schema conversion.
Native entrypoint availability depends on the installed host version and negotiated extensions. Protocol validation is not proof of native-host visual acceptance. Classic ChatGPT web is not assumed to support desktop local stdio plugins. A public ChatGPT App deployment would be a separate hosted, authenticated product with its own review and data-boundary design—not a reason to expose this local store publicly.
Reference: https://developers.openai.com/plugins/build/extensions and https://developers.openai.com/plugins/build/chatgpt-ui
## Data and safety
State is stored locally in `~/.carbon-studio-pro/workspace.json` (override with `CARBON_STUDIO_HOME`). Files are written with owner-only permissions. The Studio visual workspace makes no analytics or model calls; the bundled CARBON server is configured with analytics disabled. Specialized workflows can access the selected test target or explicitly configured services under their permission gates. The coding agent still sends the context it uses to its configured provider. This is not an air-gap guarantee or compliance certification.
Only a project explicitly selected by the agent is attached. Screenshot reads require a real path inside that project, a supported image signature, and a 5 MB limit. Screenshots are supplied to the UI in `_meta`, not in the model's structured text. All evidence text is escaped before rendering. The MCP surface is local stdio; the HTTP fallback starts lazily on a workspace/settings open call (or with `--preview`), binds loopback, and requires a private capability token plus Origin/Host checks for data access.
Budgets guide the agent; they do not enforce a host token quota. Studio never authorizes destructive testing or remediation by itself. Closing a run preserves its record. Evidence is retained until the owner removes the local workspace; there is no cloud retention service. Back up the state file before manual cleanup.
## Develop and preview
### Ad hoc browser checks and control panel
Ad hoc exploration uses the coding agent's built-in browser by default. Automatic
does not launch Selenium, Playwright, CDP, or Vibium as a fallback. A missing host
capability is reported before asking for an alternative. Explicit permitted
controller preferences and requested existing framework suites remain supported.
This is a host-agent routing contract, not a machine-level enforcement sandbox.
The overview separates executed checks from blocked/deferred work, exposes a
compact test queue, and ranks demonstrated findings before suspicions. Empty
capture panels are omitted. Evidence-map inputs share the same status semantics.
Preferences guide future runs; saved guidance is available on the agent's next
read, not proof of immediate execution. Mobile navigation retains text labels.
```sh
npm ci
npm run build
npm test
node scripts/seed-demo.mjs output/my-demo-state
CARBON_STUDIO_HOME="$PWD/output/my-demo-state" npm run preview
```
Open the private URL printed by the preview command. Keep the fragment; it is the access capability. The demo is intentionally labeled synthetic and isolated from the default workspace. It is not a benchmark or a real quality assessment. To attach a screenshot, use `studio_screenshot` in a real run, or call the same Store method in a local fixture script.
## Current limitations
- Snapshot polling currently reloads all stored runs. Large histories and many large images need pagination/lazy resources before production-scale use.
- An abruptly terminated writer may leave `write.lock`; do not remove it until all Studio processes using that store have stopped.
- Native app rendering, message sending, and model context must be acceptance-tested in each target host build. Browser preview tests do not establish ChatGPT/Codex host parity.
- The user interface is a functional foundation, not an automatic migration of all historic CARBON reports.
## Packaging
`node scripts/package.mjs` builds a distribution-only staging directory and ZIP under `output/release/`. Source, lockfile, notices, assets, skills and bundled runtime are included. User state, screenshots, tests, node_modules and private CARBON benchmark data are excluded.
TDQS
Scored across 11 tools
studio_open and studio_preferences share an identical description, making them indistinguishable, and studio_save_settings, settings.update, and settings.read all appear to write the same preferences. studio_feedback vs studio_save_settings also overlap. Several tools have unclear boundaries despite good individual descriptions elsewhere.
Eight tools use a consistent studio_verb/noun snake_case pattern, but settings.read and settings.update break the convention with dot notation and no studio prefix. Mixing two naming schemes on a single server is a clear inconsistency.
11 tools is a reasonable scope for a testing workspace lifecycle. However, the presence of redundant pairs (open/preferences, save_settings/settings.update) inflates the count slightly beyond what is earned.
The surface covers the full run lifecycle: opening the workspace, starting a run, snapshotting state, recording observations, attaching screenshots, saving settings, feedback, and steering. Main gap is that settings read/update duplicate existing functionality rather than filling a real hole.