Flutter Lamp
Provides tooling for Android device transports, listing available transports, recommending a wireless one, and optionally promoting a USB-connected device to TCP so a Flutter app's VM Service tunnel survives cable disconnection.
Uses the Dart VM Service Protocol (WebSocket or HTTP) as the transport for observing app state: subscribing to VM timeline events (build/paint/layout/GC), collecting VM exceptions and logs, reading Dart heap and external memory, and recording HTTP timeline activity for network capture.
Connects an AI agent to a running Flutter app's Dart VM Service to stream live runtime data โ exceptions with reconstructed stack traces, console/dart:developer logs, network requests, frame timings and jank, memory, route/navigation history, widget rebuild hotspots resolved to file and line, and widget-tree/Inspector snapshots โ then correlates the evidence into evidence-first root-cause diagnoses for runtime errors and performance problems.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Flutter Lampmy app just crashed on login, what went wrong?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
๐ก Flutter Lamp
Give your AI live eyes on a running Flutter app โ no more pasting logs.
An MCP server that connects Claude Code, Claude Desktop, Cursor, Codex & Gemini directly to a running Flutter app through the Dart VM Service Protocol โ streaming exceptions, logs, network calls, frame timings and memory as structured data, plus an evidence-first root-cause diagnosis engine and a live browser dashboard.
Why
Today you debug Flutter with your AI by copy-pasting stack traces, flutter run
output and DevTools screenshots. The AI is blind between messages.
Flutter Lamp makes the AI runtime-aware. It reads the app's live state over official Flutter/Dart APIs (never scraping DevTools), so instead of "paste the error" the AI can ask the app "what just happened, and why?"
โ You: *pastes 40 lines of red stack trace*
โ
AI: connect_vm โ get_exceptions โ diagnose_runtime
โ "RenderFlex overflow in Column at home.dart:42, triggered right after
GET /api/user returned 500. Confidence 85%. Fix: โฆ"Related MCP server: mcp-console-hub
Features
๐ One-line connect to any running Flutter app's VM Service
๐ฅ Realtime exceptions with reconstructed stack traces (framework + unhandled)
๐ Console & structured logs (Stdout / Stderr /
dart:developer)๐ Network capture via
dart:ioprofiling โ covers Dio &package:http, no interceptor needed๐๏ธ Frame timings with jank detection, percentiles, and a build-vs-raster verdict
๐งญ Route awareness โ current screen, transitions, and failures attributed to the screen they happened on
๐ง State-management activity โ Riverpod, Provider and Bloc churn correlated with rebuild storms
๐ Widget rebuild attribution โ which widget rebuilt, how often, at which file and line, with your code ranked apart from package code
๐งฌ Widget tree & selected-widget snapshots from the Inspector
๐งฎ Memory (Dart heap / external) and VM timeline events
๐ฉบ
diagnose_runtimeโ correlates evidence into summary ยท root cause ยท evidence ยท confidence ยท fixes; says "Unknown" below 70% instead of hallucinating๐ Live browser dashboard at
http://127.0.0.1:7373โ streams everything over WebSocket, independent of the AI connection๐งฉ Ships a reusable
flutter-runtime-diagnosisClaude Code skill
Tools
Tool | Purpose | Safety |
| Connect to a running app's Dart VM Service ( | mutates |
| Android transports; recommend wireless, optionally promote USB. | mutates |
| One-call triage โ verdict plus exception/network/frame/log/memory summary. | read-only |
| Evidence from the window before a failure, with a timeline. | read-only |
| Current route and recent transitions, with per-route failures. | read-only |
| Widget rebuild hotspots resolved to widget, file and line. | read-only |
| Riverpod, Provider and Bloc activity over time, and how often it coincides with build-heavy frames. | read-only |
| The whole session as versioned JSON โ | read-only |
| Why a diagnosis was reached: resolved evidence, alternatives, gaps. | read-only |
| Active collectors, tool safety classes, what cannot be observed. | read-only |
| Connection, session, reconnect state, event counts, retention window, what the session has cost in responses, and the VM-to-host clock offset. | read-only |
| Console + | read-only |
| Framework & unhandled exceptions with stack traces. | read-only |
| Frame build/raster timings; | read-only |
| HTTP requests (Dio & | read-only |
| Widget-tree snapshot from the Flutter Inspector. | read-only |
| Widget currently selected in the Inspector. | read-only |
| Dart heap + external memory (MB). | read-only |
| Recent VM timeline events (build/paint/layout/GC). | mutates |
| Root cause with a stable | read-only |
| Why the app is janky โ percentiles, phase split, rebuild attribution. | read-only |
| URL of the live browser dashboard. | read-only |
mutates means the tool changes app or VM state, not your project โ nothing
here writes code or files. connect_vm enables dart:io HTTP timeline logging
on the app so network capture works; get_timeline with recordFrom: true
changes the VM's recording flags. Every tool carries an MCP readOnlyHint
annotation so a client can enforce this itself.
Agents should not call all of them. The recommended flow is
runtime_health โ what_changed โ a targeted get_* โ diagnose_runtime;
see AI Agent Integration.
Install
Requires Node โฅ 20. Nothing to clone โ npx fetches it on first run:
npx -y flutter-lampOr install it globally:
npm install -g flutter-lampgit clone https://github.com/itsonu/flutter-lamp.git
cd flutter-lamp
npm install
npm run build
node dist/index.jsConnect your AI client
Claude Code plugin (recommended) โ installs the MCP server and the
flutter-runtime-diagnosis skill in one step:
/plugin marketplace add itsonu/flutter-lamp
/plugin install flutter-lamp@flutter-lampThe plugin pins an exact, published server version. Flutter Lamp is also
listed in the MCP Registry as
io.github.itsonu/flutter-lamp.
Or add the server to your MCP client config yourself, then restart the client.
Claude Code โ .mcp.json in your project:
{
"mcpServers": {
"flutter-lamp": {
"command": "npx",
"args": ["-y", "flutter-lamp"]
}
}
}Or from the CLI:
claude mcp add --scope user flutter-lamp -- npx -y flutter-lamp--scope user makes it available in every project; drop it to register
for the current project only.
Cursor (~/.cursor/mcp.json) and Claude Desktop
(claude_desktop_config.json) use the same mcpServers shape.
Running from a clone instead? Point command at node and args at
/absolute/path/to/flutter-lamp/dist/index.js.
Optional environment variables:
Var | Default | Meaning |
|
| Dashboard HTTP/WS port. |
|
| Bind address (localhost only by default). |
| โ | Set to |
| on | Set to |
| โ | Comma-separated extra header-name patterns to redact. |
Usage
Run your Flutter app in debug/profile mode:
flutter runCopy the line it prints:
A Dart VM Service on <device> is available at: http://127.0.0.1:PORT/TOKEN=/Tip:
flutter run --vm-service-port=8181gives a stable URI across restarts.Ask your AI to connect and diagnose โ e.g. "connect to my Flutter app at
<uri>and tell me why it's throwing." With the bundled skill, Claude Code runs the whole flow (connect โ gather โdiagnose_runtime) automatically and never asks you to paste logs when the VM Service is reachable.Open http://127.0.0.1:7373 in a browser for the live dashboard โ it runs alongside the AI, not instead of it.
Debugging without a cable (Android)
A flutter run started over USB loses its VM Service tunnel the moment the
cable moves; one started over a TCP transport does not. Ask the server which
transports exist:
ensure_tcp_device # read-only: lists transports, recommends one
ensure_tcp_device { promote: true } # puts a USB-only device on TCP (needs the cable once)Then launch against the wireless serial it recommends:
flutter run -d 192.168.88.3:5555The cable is only needed for the one-time promotion. Reverse it any time with
adb usb.
Live dashboard
A zero-dependency dark UI (native HTTP + WebSocket, no build step) that streams runtime data as it happens:
Overview โ a session report: health, FPS, memory, worst frame, MCP activity, and findings that link through to the events behind them (the worst-frame finding opens the timeline filtered to that frame) ยท Logs (search / filter / auto-scroll) ยท Network (expandable headers & timing) ยท Exceptions (expandable stack traces) ยท Timeline ยท Performance (live canvas charts) ยท MCP โ which agent connected, every tool it called, with latency, errors and response size ยท Inspector โ which inspection links exist and which are pull-only ยท Diagnostics โ the measured topology, per-collector health, and what this target cannot observe at all.
Controls: pause/resume ยท clear view ยท export JSON ยท per-tab search ยท auto-reconnect.
The UI distinguishes states that look alike and are not: 0 MB is a
measurement, not sampled yet is the absence of one, and an empty tab says
whether a filter is hiding rows, the collector is blind on this target, or the
collector is watching and has seen nothing. Readings taken before a disconnect
are labelled as such rather than left looking live.
How it works
Running Flutter app
โ Dart VM Service Protocol (JSON-RPC over WebSocket)
โผ
VmService client โโโถ Collectors (log ยท exception ยท frame ยท network ยท โฆ)
โ
โผ
RuntimeStore (one centralized, capped event stream;
every event: timestamp ยท source ยท severity ยท category)
โ โ
โโโโโโโโโโโโ โโโโโโโโโโโโ
โผ โผ
MCP tools (stdio) Dashboard (HTTP + WebSocket)
โ Claude Code / Cursor / โฆ โ your browserEverything flows through one event store. Adding a new runtime source =
implement the Collector interface and register it โ no other layer changes.
Design principles: official Flutter/Dart APIs only, structured JSON over text,
never scrape DevTools, and never claim a cause the evidence doesn't support.
Limitations
Debug/profile builds only. The VM Service, Inspector and
dart:ioHTTP profiling are not available in release builds.Network is pull-on-demand. Dart exposes no push stream for
dart:ioHTTP, so requests are fetched whenget_networkordiagnose_runtimeruns โ not streamed continuously.Exceptions are not always observable.
Debug.PauseExceptionneeds pause-on-exception enabled in the app to fire.Flutter.Erroris posted by the widget inspector only while structured error reporting is on, which the framework disables in profile mode and on the web โ so on those targets a thrown framework error leaves no trace at all. The collector asks the app and reportsdegradedwhen that is the case, so an empty exception list is never silently mistaken for a quiet one.Dart-side HTTP only. Calls made from platform (Kotlin/Swift) code or from a WebView don't appear in
get_network.Retention is bounded. Each category keeps its own fixed window (3,000 logs, 2,000 state, 1,000 exceptions, 1,000 network, 1,000 frames, 1,000 rebuilds, 500 navigation, 500 system). Frames roll over fastest โ about 17 seconds at 60fps, and measured at 4 minutes under a 10fps workload.
runtime_statusreports what is retained and what was evicted, and once eviction starts every diagnosis carries a note saying so, so a truncated window is never presented as a whole session.The dashboard binds to
127.0.0.1by default. ChangeDASHBOARD_HOSTonly on a network you trust โ runtime data is served unauthenticated.
Security
Runtime data is sensitive. HTTP headers carry bearer tokens and cookies, URIs carry API keys, and developers print credentials into logs โ and everything captured is handed to an AI model and streamed to any browser watching the dashboard.
Secrets are redacted at capture, so they never enter the event store and no
consumer can leak what was never stored. Redacted by default: Authorization,
Proxy-Authorization, Cookie, Set-Cookie, WWW-Authenticate, any header
whose name contains token, secret, password, credential, api-key or
session, sensitive query-string parameters, and JWT- or Bearer-shaped
strings in log lines and error text. Header names that were hit are listed in
data.redactedHeaders so you can see that something was withheld rather than
getting a silently partial picture. Add patterns with
FLUTTER_LAMP_REDACT_EXTRA, or disable entirely with FLUTTER_LAMP_REDACT=off
for a local-only session.
The VM Service credential is scrubbed unconditionally. The path segment of a
VM Service URI authorises evaluate โ arbitrary Dart execution in the app being
observed โ and an app can print its own URI to its console, which a web target
does on startup. That token is removed from captured text wherever it appears,
and unlike everything above it is not governed by FLUTTER_LAMP_REDACT=off:
that switch exists so you can read your own app's headers and log text, not so
the server will hand out the key to the app it is attached to.
The dashboard is not exposed to other pages. Binding to loopback does not
protect a WebSocket โ browsers exempt WebSocket from the same-origin policy, so
without a check any page you have open could connect to
ws://127.0.0.1:7373/ws and read your whole runtime stream. The handshake
requires a per-process token that is inlined into the served page, which
cross-origin script cannot read, and a present Origin header must be loopback.
The page is served X-Frame-Options: DENY, and /health returns liveness only
โ never the VM Service URI, which embeds the VM's own auth token.
Setting DASHBOARD_HOST to a non-loopback address logs a warning and puts
runtime evidence on your network. Only do that on a network you trust.
Found a vulnerability? See SECURITY.md.
Roadmap
Shipped: VM connect, logs, exceptions (with stacks), network, frames, widget tree, memory, timeline, route awareness, rebuild attribution, state-management activity, diagnosis, live dashboard, and an evaluation suite that replays recorded sessions to check the diagnosis against apps that actually misbehaved.
Next: deeper correlation (memory/timeline into diagnose_runtime), CPU sampling
& leak heuristics, knowledge graph, auto-fixes. See
docs/Phases.md and
docs/Observability-Roadmap.md.
Development
npm run build # compile TypeScript โ dist/
npm run watch # incremental compile
npm test # node:test suite (engine + collectors + dashboard)npm test runs the compiled output, so build first. It relies on glob support in
node --test, which needs Node โฅ 21 โ the server itself runs on Node โฅ 20.
Docs live in docs/:
Doc | Contents |
Problem, users, principles, non-goals, constraints | |
Data flow, components, event model, how to add a collector | |
Non-negotiable constraints every change is checked against | |
Roadmap and status | |
The investigation protocol agents should follow | |
Current audit and prioritized backlog | |
Architecture target, phases, and explicit rejections | |
Where runtime events actually flow, and what is observable | |
How versions are staged, tagged and published | |
Dashboard design system โ tokens, status vocabulary, component rules | |
Non-obvious things learned building it |
Contributing
Issues and PRs welcome. Keep it modular, prefer official APIs over hacks, and
add a node:test check for any non-trivial logic.
License
MIT ยฉ Chandrabhushan Prakash
Available Tools
22 toolsconnect_vmConnect to Flutter VM ServiceAIdempotent
Connect to a running Flutter app's Dart VM Service and start collecting runtime data (logs, exceptions, frames, network). Pass the ws:// or http:// URI printed by flutter run (line: 'A Dart VM Service ... is available at:'). NOT purely read-only: enables dart:io HTTP timeline logging on the app so network capture works, and adds the GC stream to the VM timeline recorder so garbage-collection pauses can be correlated with jank. Existing recorded streams are preserved, never replaced.
| Name | Required | Description | Default |
|---|---|---|---|
| uri | Yes | VM Service URI, e.g. http://127.0.0.1:52719/abcdef=/ or ws://... |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Explicitly discloses side effects beyond the annotations: it enables dart:io HTTP timeline logging on the app and adds the GC stream to the VM timeline recorder, and it states existing recorded streams are preserved rather than replaced. This clarifies the non-read-only (readOnlyHint=false) and idempotent (idempotentHint=true) hints with concrete mechanism.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and payoff, then the parameter source, then the mutation caveat. Information-dense and mostly waste-free, though the parenthetical stream-logging detail is slightly wordy for an agent that mainly needs the URI.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with no output schema, it covers what the tool does, what it changes on the target app, and how to obtain the URI. It omits session lifecycle details (whether repeated connects stack sessions, how to disconnect), which are minor but would complete the picture.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so the baseline is 3, but the description adds real value by telling the agent where the URI comes from (the `flutter run` output line) and that both ws:// and http:// forms are accepted. That sourcing guidance is not present in the schema example.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (connect) plus resource (Flutter app's Dart VM Service) and the immediate consequence (start collecting runtime data). Clearly distinguishable from sibling readers like get_logs, get_timeline, or runtime_status, which presume an already-connected session.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives the precondition for the required parameter precisely: pass the ws:// or http:// URI printed by `flutter run`, quoting the exact line to look for. It does not, however, name alternatives or state when *not* to call it (e.g., when a session is already connected), so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
diagnose_performanceDiagnose performanceARead-onlyIdempotent
Why the app is janky, not just how much. Returns frame percentiles, the build-vs-raster split, and findings correlating jank against in-flight requests, route transitions and heap growth โ each with its own evidence ids, strength and fix. Reports 'healthy' when jank is within normal range and 'unknown' when there are too few frames to tell a pattern from noise. States what it cannot see: no CPU sampling, no widget rebuild counts. GC pauses are captured and correlated, but the strength of that correlation is reported against how much of the window frames actually covered.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent/non-destructive annotations, the description discloses a great deal: the shape of findings (evidence ids, strength, fix), the 'healthy' and 'unknown' sentinel states, explicit blind spots (no CPU sampling, no widget rebuild counts), and the caveat that GC correlation strength is bounded by frame coverage of the window. This is unusually rich behavioral disclosure that annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a long description, but it is front-loaded with the core proposition ('why... not just how much') and every subsequent clause adds distinct information about outputs, states, or limits. Density is high enough that the length is mostly justified, though it could be trimmed slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description bears full responsibility for describing returns, and it does: percentiles, build-vs-raster split, correlated findings with evidence ids/strength/fix, plus healthy/unknown semantics and explicit limitations. An agent has everything it needs to interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline of 4 applies and there is no parameter semantics to clarify. Nothing in the description contradicts or undermines this.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource and sharpens the scope with the contrast 'why the app is janky, not just how much,' which implicitly separates it from raw measurement tools like get_frames. It does not name any sibling explicitly, so an agent must infer the boundary from the 'why vs how much' framing rather than a direct reference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is clear: reach for this when you need causes rather than magnitudes, and the 'healthy'/'unknown' states tell the caller when the output is trustworthy versus noise. There are no explicit exclusions or named alternatives, so routing against diagnose_runtime or get_frames is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
diagnose_runtimeDiagnose runtimeARead-onlyIdempotent
Correlate captured runtime evidence into a root-cause diagnosis. Returns status (diagnosed|unknown), summary, rootCause, evidence (each with a citable eventId), a chronological timeline around the root cause, alternativeCauses that also fit, limitations describing what could not be seen, confidence (0-1) with a breakdown of evidence strength / data completeness / alternative strength, and recommended fixes. Status is 'unknown' below 70% rather than a guess.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare this as read-only, idempotent, non-destructive, and closed-world, so safety behavior is covered. The description adds substantial behavioral context beyond annotations: returned fields, diagnosed/unknown status, the 70% confidence cutoff, confidence breakdown, limitations, alternative causes, timeline, and recommended fixes. There is no contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first sentence, followed by a dense but useful inventory of return behavior and the unknown-status threshold. No sentence is filler, and the length is justified because no output schema exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Because there is no output schema, the description's detailed return-value explanation is essential and largely complete for a zero-parameter diagnostic tool. It still omits routing guidance versus sibling diagnostic tools and any precondition about how evidence must be captured, keeping it short of 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so there are no parameter semantics for the schema or description to document. The baseline for a zero-parameter tool is 4, and the description appropriately avoids inventing parameter guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (correlate captured runtime evidence into a root-cause diagnosis) and resource, going beyond merely restating the title. However, it does not distinguish itself from sibling diagnostic tools such as diagnose_performance or explain_diagnosis, so it is clear but not sibling-differentiating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit when-to-use, prerequisite, or alternative guidance. It implies the tool is used after runtime evidence has been captured, but does not say when to choose it over diagnose_performance, explain_diagnosis, or the lower-level evidence tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ensure_tcp_devicePrefer a wireless device transportAIdempotent
Report Android device transports and recommend one, preferring wireless. A flutter run started on a USB transport loses its VM Service tunnel when the cable moves; one started on a TCP transport does not. Read-only by default. With promote:true it runs adb tcpip and adb connect to put a USB-attached device on a TCP transport โ that restarts adbd on the device, needs the cable once, and is reversible with adb usb. Android-only: reports adbAvailable:false and changes nothing on iOS, desktop or web targets.
| Name | Required | Description | Default |
|---|---|---|---|
| port | No | Device-side TCP port. | |
| serial | No | Which USB device to promote. Defaults to the first promotable one. | |
| promote | No | Put a USB-only device onto a TCP transport. Changes device state. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=false and idempotentHint=true but leave readOnlyHint=false unexplained; the description resolves that by stating 'Read-only by default' and disclosing the mutation's side effects โ adbd restarts on the device, the cable is needed once, and `adb usb` reverses it. That is material behavior beyond what the annotations convey, and it does not contradict them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the read-only report behavior before the promote path and platform caveats. Dense but essentially waste-free; the clause about the VM Service tunnel is justification, not filler, though the sentence count could be trimmed slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and an optional state-changing mode, the description covers the return signal (adbAvailable:false on unsupported platforms), the mutation's preconditions and reversibility, and the default safety posture. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it explains that promote performs `adb tcpip`/`adb connect`, and that serial selects which USB device to promote. Port is left to the schema, which fully documents it, so this is a modest but genuine uplift.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Report Android device transports and recommend one, preferring wireless') and then scopes the mutation mode ('With promote:true it runs adb tcpip and adb connect'). No sibling tool overlaps this behavior, so an agent can route to it unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames when the tool matters (a `flutter run` on USB loses its VM Service tunnel when the cable moves), what the default is (read-only), when to escalate (promote:true), when it does NOT apply (iOS, desktop, web), and how to undo it ('reversible with `adb usb`'). This is when/when-not/alternative guidance in full.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explain_diagnosisExplain a diagnosisARead-onlyIdempotent
Why diagnose_runtime reached its conclusion: the claim, every cited event resolved back to its full record, the timeline around the root cause, competing explanations, what evidence is missing, and the confidence breakdown. Use to answer 'why do you think that' without inventing reasoning.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds meaningful behavioral content beyond that โ the return surface (claim, resolved events, timeline, competing explanations, missing evidence, confidence breakdown) and the constraint against fabricating reasoning. It doesn't address latency or cost, but for a read-only explainer this is solid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose ('Why diagnose_runtime reached its conclusion'), followed by a dense but purposeful enumeration of returned content and one usage sentence. The list is long but every element earns its place given there is no output schema; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must carry the return-value burden, and it does so thoroughly by enumerating the claim, resolved evidence, timeline, competing explanations, gaps, and confidence breakdown. Combined with zero parameters and clear annotations, nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to disambiguate at the parameter level.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (explain) and resource (the diagnosis produced by diagnose_runtime), and it explicitly frames the tool as the 'why' counterpart to diagnose_runtime. An agent can distinguish it from siblings like diagnose_runtime or get_timeline without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear triggering condition: 'Use to answer "why do you think that" without inventing reasoning,' which implies invocation after a diagnose_runtime conclusion. It doesn't explicitly name alternatives or exclusions, but the usage context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_sessionExport the debugging sessionARead-onlyIdempotent
The whole session as one versioned JSON artifact: metadata, per-collector health, retention, the captured events, and every diagnosis (runtime, performance, navigation, rebuilds). Use mode 'brief' for the smallest sufficient context โ the diagnoses plus only the events their evidence cites โ and 'full' to archive everything retained, for a bug report, offline analysis or a regression fixture. Sizes are not close: measured on a ~1,700-event session, 'brief' returned 36kB and 'full' returned 247kB โ roughly 62,000 tokens, about a third of a 200k context window, so treat 'full' as something to write to a file rather than read inline. Credentials are already redacted at capture, so nothing here was ever stored raw.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | 'brief': diagnoses plus only the cited events. 'full': everything retained. | brief |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, and the description goes further with measured payload sizes (36kB vs 247kB, ~62,000 tokens) and the fact that credentials are redacted at capture. These are behavioral traits an agent could not derive from the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the artifact definition, then mode semantics, then the cost rationale. Every sentence earns its place; the size/token figures and redaction note are load-bearing, not padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the burden of describing the return value and does so thoroughly (contents of the artifact, both mode shapes, size implications). Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the enum already defines both modes, so the baseline is 3. The description nonetheless adds real decision-relevant meaning beyond the schema: the measured size gap between modes and the practical implication (inline vs file), enabling a smarter mode choice.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('export the whole session as one versioned JSON artifact') and enumerates exactly what the artifact contains: metadata, per-collector health, retention, captured events, and every diagnosis category. This distinguishes it from the sibling tools that each return one slice (get_logs, get_frames, diagnose_runtime, etc.).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes between the two modes with concrete conditions: 'brief' for the smallest sufficient context, 'full' for bug reports, offline analysis, or a regression fixture. It even prescribes behavior ('treat full as something to write to a file rather than read inline'), which is exactly the when-to-use guidance an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_capabilitiesWhat this server can observeARead-onlyIdempotent
Machine-readable capability report: active collectors, every tool with its safety class, what can and cannot be observed on this target, and the current redaction, dashboard and retention configuration. Read this before attempting an operation that may not be supported.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, non-open-world), so the bar is lowered. The description adds genuine value beyond that by disclosing what information the report exposes โ including per-tool safety classes and current redaction/retention settings โ which tells the agent this is a discovery/authorization-check step. It does not describe output format or size, but the content disclosure is substantive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the noun phrase 'Machine-readable capability report' and the contents listed before the call-to-action. Dense but every clause carries information; the mid-sentence inventory is slightly list-heavy but not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must convey what comes back โ and it does, specifying collectors, tool safety classes, observability boundaries, and configuration. With zero parameters and annotations covering safety, nothing an agent needs in order to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, which is the baseline-4 case. The description correctly implies a parameterless, whole-server query and adds no misleading parameter guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific artifact (a machine-readable capability report) and enumerates its contents: active collectors, per-tool safety class, observability limits, and redaction/dashboard/retention configuration. No sibling tool in the list returns this meta-level inventory, so an agent can distinguish it immediately from the diagnostic getters (get_logs, get_frames, diagnose_runtime, etc.).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger condition: 'Read this before attempting an operation that may not be supported.' That is clear when-to-use guidance. It stops short of naming a specific alternative tool or a when-not-to-use case, but for a self-describing capability tool no meaningful alternative exists.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_dashboard_urlGet dashboard URLARead-onlyIdempotent
Return the URL of the live Realtime Runtime Dashboard (a browser UI streaming logs, network, exceptions, frames & memory). Open it in a browser to watch the app alongside the AI.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and non-destructive, so safety is covered. The description usefully adds what the dashboard displays (logs, network, exceptions, frames, memory), but says nothing beyond that โ no note on URL lifetime, whether it requires a running session, or what to do if it's unavailable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the return value front-loaded and the parenthetical describing the dashboard placed immediately after the noun. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only URL getter with full annotation coverage and no output schema, the description is essentially complete. Minor omissions (URL validity window, any prerequisite session state) keep it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so there is no parameter semantics to convey; the baseline for a no-arg tool is 4. The description neither needs to nor does add parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("Return the URL of the live Realtime Runtime Dashboard") and characterizes the resource as a browser UI streaming logs, network, exceptions, frames & memory, which implicitly differentiates it from the individual get_logs/get_network/get_memory siblings. It stops short of explicitly routing the agent to those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Open it in a browser to watch the app alongside the AI" gives a clear intended action, so usage is implied rather than absent. However, it never states when to prefer this aggregate view over the granular siblings, nor any when-not condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_exceptionsGet runtime exceptionsARead-onlyIdempotent
Flutter framework errors and unhandled VM exceptions, most recent first. Each includes the error summary, offending widget, library, and a reconstructed stack trace (data.stackTrace) when available.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive, closed-world behavior, so safety is covered. The description adds real behavioral value beyond that: result ordering (most recent first) and the specific content of each entry, including the data.stackTrace path where a stack trace may be absent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the first front-loads scope and ordering, the second enumerates returned fields. No filler, though the second sentence packs several clauses together.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the returned fields and the optional stack trace, and annotations cover the safety profile. Still missing is any mention of the limit parameter's effect or whether results are truncated, which would matter for a list-returning tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter (limit, default 20, max 200) with 0% schema description coverage, and the description never mentions it or what capping the list does. The name and bounds are largely self-explanatory, so this is adequate but adds no meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (Flutter framework errors and unhandled VM exceptions) and ordering (most recent first), which lets an agent distinguish it from generic log tools. It does not name a sibling such as get_logs to sharpen the boundary, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The scope wording ('framework errors and unhandled VM exceptions') implies when this tool is the right one, but there is no explicit when-to-use, when-not, or alternative named among the many siblings (get_logs, runtime_health, diagnose_runtime). Usage is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_framesGet frame timingsARead-onlyIdempotent
Frame build/raster timings. Set onlyJanky to focus on frames over the frame budget โ 16.67ms (60fps) by default, which is an assumption: the VM Service does not report the display refresh rate. Override with FLUTTER_LAMP_FRAME_BUDGET_MS when the target's rate is known.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| onlyJanky | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only, idempotent safety profile, so the description adds valuable behavioral context: the default frame budget of 16.67ms is an assumption because the VM Service does not report refresh rate, and it can be overridden via FLUTTER_LAMP_FRAME_BUDGET_MS. It does not describe return format, but that is a minor gap against otherwise strong disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, front-loaded with the core purpose and then the onlyJanky nuance and override instruction. Every sentence adds useful information with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple query tool with annotation coverage, the description is adequate but leaves gaps: it does not explain the limit parameter, what fields the timings contain, or whether results are paginated. It also gives no tool-selection context among related performance tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It fully explains onlyJanky (filters to frames over the frame budget, default budget, env override) but completely omits the limit parameter, leaving half the parameters undocumented beyond their types and defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the resource ('Frame build/raster timings') with an implied retrieval verb, so an agent can tell what is returned. It does not differentiate from siblings like get_timeline or diagnose_performance, but the purpose is clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives guidance only for the onlyJanky parameter, not for when to use this tool versus alternatives such as get_timeline or diagnose_performance. No explicit context or exclusions for tool selection are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_logsGet console & structured logsBRead-onlyIdempotent
Console output (Stdout/Stderr) and dart:developer logging, most recent first.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| source | No | Restrict to one log source. | |
| sinceMs | No | Only events at/after this epoch-ms timestamp. | |
| contains | No | Case-insensitive substring filter. | |
| minSeverity | No | Minimum severity to include. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds the useful ordering behavior ('most recent first') and the source scope, but says nothing about the default limit (50), truncation, or whether results are capped/paginated. Adds some value beyond annotations but leaves behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tight sentence that front-loads the content sources and ends with the ordering guarantee. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter read-only tool with no output schema, the description omits return shape, default limit behavior, and pagination/truncation semantics that an agent needs to call it correctly. The annotations cover safety, but the operational gaps remain, leaving it only minimally adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 80%, so the schema already documents source, sinceMs, contains, and minSeverity. The description's mention of Stdout/Stderr/Logging loosely maps to the source enum but adds no syntax or format detail for limit or the filters. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('get_logs') and names the exact content it returns: Console output (Stdout/Stderr) and dart:developer logging. The 'most recent first' clause pins down ordering. It doesn't explicitly contrast with siblings like get_exceptions or get_network, but the log-source enumeration makes the domain clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance, no prerequisites, and no alternatives. An agent must infer that this is for raw console/developer output versus get_exceptions for crashes or get_network for traffic. There is no explicit routing signal.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_memoryGet memory usageARead-onlyIdempotent
Current Dart heap usage for the main isolate (Dart heap in use, capacity, and external/native memory), in MB. Also records a snapshot into runtime history.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), so the bar is lower; the description adds genuinely new context by disclosing a side effect the annotations do not surface โ that each call records a snapshot into runtime history. It does not explain retention or snapshot cost, but this is meaningful added behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, and the primary purpose is front-loaded ahead of the less-obvious snapshot side effect. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters and no output schema, the description must stand in for the return value, and it does so by listing the metric components and their unit. The snapshot-to-history behavior is stated but its semantics (how much history, whether it affects later queries) are left open.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description correctly adds no parameter discussion, and the schema (100% coverage) is trivially complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names the precise resource (Dart heap usage for the main isolate) and enumerates the sub-metrics returned (heap in use, capacity, external/native memory) with units. An agent can distinguish this from siblings like get_logs or runtime_status without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this versus alternatives such as runtime_status, runtime_health, or diagnose_performance. The description explains what comes back but leaves the selection decision entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_networkGet network requestsARead-onlyIdempotent
HTTP requests/responses captured via dart:io profiling (covers Dio & package:http). Fetches the latest profile on demand, then returns completed requests most recent first.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds genuine behavioral context beyond that: it fetches a fresh profile on demand (a live capture, not a passive read of cached state) and returns only completed requests, most recent first. The main gap is that it never mentions the default page size or that only a bounded window of requests is returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with what the data is and where it comes from, then the fetch/return behavior. No filler, no repetition of the title, every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only inspector with annotations covering safety and no output schema, the description supplies the essentials: data source, on-demand capture, and result ordering/filtering. It stops short of sketching what a returned request record looks like or how the limit bounds the result set, which an agent would want given the absent output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one parameter (limit) and schema description coverage is 0% โ the schema supplies type, default 30, and max 200 but no explanation of what is being limited. The description never mentions limit at all, so it fails to compensate for the coverage gap. The name is largely self-evident, which keeps this above 1, but the description contributes nothing here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (HTTP requests/responses) and discloses the capture mechanism (dart:io profiling, covering Dio & package:http), which makes the tool's data source unambiguous. It does not explicitly distinguish itself from siblings like get_logs or get_timeline, but the domain is clear enough to select it confidently.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: an agent can infer this is the tool to call when it needs network traffic, but there is no 'use this when / instead of X' guidance and no mention of prerequisites or timing relative to other diagnostics. Adequate but leaves routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_rebuildsWidget rebuild hotspotsARead-onlyIdempotent
Which widgets are rebuilding and how often, resolved to widget name, file and line, with your own code ranked above package code. Use for 'why is this screen slow to build' and to find needless rebuilds. Requires a debug build with widget creation tracking; reports why it is empty otherwise.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many hotspots to return. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), so the bar is lower. The description nonetheless adds real context beyond them: a debug build with widget creation tracking is required, and the tool reports why it is empty otherwise โ a genuine prerequisite and failure-mode disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core purpose, then usage, then the prerequisite. Every sentence carries information; nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter diagnostic tool with no output schema, the description explains what is returned (widget names, file/line, ranked), the scenario it answers, and the empty-result condition. Nothing an agent needs to call it correctly appears to be missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With a single parameter at 100% schema description coverage, the schema already documents 'limit' and its bounds. The description adds no meaning about the parameter, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('which widgets are rebuilding and how often') and pins down the output resolution (widget name, file, line) plus the ranking rule (own code above package code). This is clearly distinguishable from siblings like get_widget_tree or diagnose_performance, which cover different concerns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete usage scenarios ('why is this screen slow to build', finding needless rebuilds), which tells the agent when to reach for it. It does not, however, name alternative tools (e.g. diagnose_performance, get_timeline) or state when this is the wrong choice, so routing is not fully closed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_selected_widgetGet selected widgetARead-onlyIdempotent
The widget currently selected in the Flutter Inspector (via 'select widget mode' in the app/DevTools). Returns null if nothing is selected.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds genuinely useful behavior beyond that: the selection must come from 'select widget mode', and the result is null when nothing is selected.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero waste. The primary identity ('what is returned') is front-loaded, and the null fallback follows immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly carries the return-value burden by documenting the null case. It slightly under-specifies the shape of a returned widget (id, node, properties), but for a zero-parameter read tool this is close to complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate. Baseline of 4 applies; no parameter semantics are needed or missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb+resource ('the widget currently selected') and identifies its source (Flutter Inspector select widget mode), so the agent knows exactly what it retrieves. It does not explicitly contrast with the sibling get_widget_tree, though the distinction is inferable from 'selected' vs. tree-wide.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the 'select widget mode' reference and the null case, telling the agent this only works when the user has an active selection. There is no explicit guidance on when to prefer this over get_widget_tree.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_state_activityState-management activityARead-onlyIdempotent
How much state-management activity the app is doing and when, plus how often build-heavy frames coincide with it โ use with get_rebuilds to answer 'is this rebuild storm driven by state churn'. Counts and timing only: Riverpod sends an app-side buffer offset and provider an element id, neither resolvable to a provider name or value. Counts are NOT transition counts โ a provider event means dependents were notified, so one state change in a widget-heavy tree produces many events. Stock Bloc posts nothing itself; a flutter_bloc app appears here only as the provider activity its notifications cause.
| Name | Required | Description | Default |
|---|---|---|---|
| buckets | No | How many of the busiest one-second buckets to return. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive behavior, but the description adds substantial non-obvious context: Riverpod emits an opaque buffer offset and provider emits an element id, neither resolvable to a provider name; counts are event notifications rather than transition counts, so one state change in a heavy tree yields many events; and stock Bloc posts nothing. These caveats materially change how an agent should interpret the numbers.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, each carrying distinct information, with the primary purpose front-loaded before the get_rebuilds pairing and the interpretation caveats. The later sentences are long but every clause earns its place; nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must convey what comes back, and it does describe the shape at a high level (counts and timing, per-bucket busiest seconds) plus the key interpretive traps. It could be slightly more explicit about the return structure, but it covers what an agent needs to interpret results correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'buckets' parameter is already documented in the schema with its default, maximum, and meaning. The description adds no further syntax, units, or semantics for it, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource and measurement ('how much state-management activity the app is doing and when, plus how often build-heavy frames coincide with it'), which is far more precise than the tool name alone. It distinguishes itself from siblings by naming get_rebuilds as the complementary tool rather than a substitute.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit usage context: pair with get_rebuilds to answer 'is this rebuild storm driven by state churn.' That is a clear condition that routes the agent to the right combination of tools. It stops short of stating when not to use this tool or what other diagnostic siblings might supersede it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_timelineGet VM timeline eventsAIdempotent
Recent VM timeline trace events (build/paint/layout/GC/etc.), most recent first. Requires timeline recording โ enable with recordFrom=true (sets Dart, GC, Compiler & Embedder streams) then reproduce the activity. recordFrom=true is NOT read-only: it changes the VM's recording configuration. Check recorderLagMs/stalled in the result: the VM recorder can stall permanently once its buffer fills while still reporting its streams as recorded, so events may be historical rather than current.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| recordFrom | No | Turn on timeline recording streams before reading. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say readOnlyHint=false; the description goes further and explains that recordFrom=true mutates the VM's recording configuration and enables specific Dart/GC/Compiler/Embedder streams. It also discloses the recorderLagMs/stalled failure mode where the recorder can stall permanently and still report streams as recorded, which is behavior no annotation conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the returned resource and ordering, then the recording prerequisite, then the result caveat. Every sentence carries needed information, though the final sentence is dense and would benefit from being split.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does name specific result fields (recorderLagMs, stalled) and warns events may be historical. It stops short of stating what is returned when recording was never enabled (empty vs. error), which is the main remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50%. recordFrom is richly explained (which streams it turns on, and that it is not read-only), adding value beyond the schema, but the `limit` parameter is never mentioned, so its semantics (cap of 200, default 50) rest entirely on the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a precise resource (VM timeline trace events) and enumerates the event classes returned (build/paint/layout/GC/etc.) plus ordering (most recent first). This clearly distinguishes it from siblings like get_logs, get_frames, get_memory, and get_rebuilds.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the prerequisite explicitly: recording must be enabled via recordFrom=true and the activity reproduced before events appear. It does not name alternative tools (e.g., get_logs or get_frames) for adjacent diagnostic needs, so the when-not-to-use side is absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_widget_treeGet widget treeARead-onlyIdempotent
Snapshot of the running app's widget tree (summary tree from the Flutter Inspector). Use to understand structure, find a widget, or see what is mounted.
| Name | Required | Description | Default |
|---|---|---|---|
| maxDepth | No | How deep to traverse before truncating. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and openWorldHint=false, so the safety profile is covered. The description adds that this is a 'summary' tree of a 'running app'โimplying an attached session and truncated detailโbut does not say what is omitted, whether it requires a live connection, or how large trees are handled. Useful but modest added context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with what the tool returns, followed by when to reach for it. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the return-value burden; 'summary tree from the Flutter Inspector' does convey the shape and the summary/truncated nature. It could say more about output scale or the need for a live app session, but it is adequate for a simple read-only diagnostic.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter (maxDepth) and schema description coverage is 100%, with a clear default and maximum already documented. The description adds no traversal or truncation semantics beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (snapshot of the running app's widget tree) and names the source (Flutter Inspector summary tree). An agent can distinguish it from get_selected_widget, though the description does not explicitly contrast with that sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives implied usage contexts ('understand structure, find a widget, or see what is mounted'), which is genuine guidance, but names no alternative and gives no when-not conditions (e.g., when to prefer get_selected_widget or get_rebuilds).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
runtime_healthRuntime health snapshotARead-onlyIdempotent
One compact answer to 'is this app healthy right now'. Returns a verdict (healthy/degraded/failing/no-data) plus exception, network, frame, log and memory summaries with citable event ids, the retention window, and notes about anything that qualifies the numbers. Call this FIRST instead of calling six get_* tools.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry the safety profile (readOnly, idempotent, non-destructive), so the bar is lower. The description adds real behavioral value by disclosing the verdict enum values (healthy/degraded/failing/no-data), the retention window, and that some numbers carry qualifying notes. It does not describe latency, cost, or staleness, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the one-line value proposition and then the return contents plus the routing directive. Every clause earns its place with no repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must convey the return shape, and it does so well (verdict, five summaries, event ids, retention, qualifiers). With annotations covering safety and a zero-param contract, this is nearly complete; it could note whether data is app-scoped versus session-scoped but that is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the baseline is 4. There is nothing for the description to clarify beyond a no-arg call, though the description could note that no input is required at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific purpose ('is this app healthy right now') and enumerates exactly what it returns (verdict, exception/network/frame/log/memory summaries, retention window, qualifiers). It explicitly distinguishes itself from six sibling get_* tools, so an agent can route without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Call this FIRST instead of calling six get_* tools,' which is a direct when-to-use directive with named alternatives. That is exactly the routing guidance the sibling-heavy list (get_logs, get_exceptions, get_frames, get_network, get_memory, runtime_status) needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
runtime_statusRuntime health checkARead-onlyIdempotent
Report connection health, the current debugging session, reconnection state, how many runtime events have been captured by category, and the retention window (per-category capacity, how many events were evicted, and the oldest event still held). Use to confirm the MCP is receiving live data and to know how far back the evidence goes.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so safety is covered. The description still adds substantive behavioral context the annotations lack: event categories, per-category capacity, eviction counts, and the oldest retained event โ i.e., that results are bounded by a retention window. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences: the first front-loads the returned content, the second states the use case. Dense but every clause carries information; no filler or restatement of the title. Slightly list-heavy in the first sentence but still efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must carry return-value explanation โ and it does, detailing connection/session state, per-category event counts, capacity, evictions, and the retention window. Combined with read-only annotations, an agent has enough to call and interpret this correctly; only sibling disambiguation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to disambiguate and no undocumented inputs to compensate for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('report connection health, current debugging session, reconnection state, runtime event counts, retention window') and enumerates the reported fields, so an agent knows exactly what comes back. It does not, however, differentiate itself from close siblings like runtime_health or diagnose_runtime, which is the main clarity gap.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a positive use case ('confirm the MCP is receiving live data', 'know how far back the evidence goes'), which is real guidance. But it names no alternative and no exclusion condition, leaving the agent to guess when runtime_status is preferable to runtime_health or diagnose_runtime in the sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
what_changedWhat changed before an incidentARead-onlyIdempotent
Evidence from the window leading up to a failure, plus a baseline comparison: the incident window measured against the equal window before it, per dimension (exceptions, network volume/failures/latency p50/p95, jank ratio, log errors, memory) with directions new/spiked/increased/decreased and citable evidence. Anchors on the given eventId, or the most recent exception, or the current time. Network uses interval matching, so a request that started before the window but failed inside it still counts. When the baseline predates observation, directions are unknown rather than fabricated.
| Name | Required | Description | Default |
|---|---|---|---|
| eventId | No | Anchor on this event (e.g. 'exc_00142'). Defaults to the most recent exception. | |
| windowMs | No | How far back to look, in milliseconds. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive, closed-world behavior, but the description adds substantive semantics beyond them: network uses interval matching so a request failing inside the window counts even if it started earlier, and baseline directions are reported as unknown rather than fabricated when observation predates the window. These edge-case disclosures are exactly the kind of context annotations cannot carry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is one dense, front-loaded block that leads with what is returned before the mechanics. Every clause contributes information (dimensions, directions, anchoring, interval matching, baseline honesty) with no filler, though the nested parentheticals make it heavier than it needs to be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the burden of describing returns falls on the description, and it does so thoroughly: dimensions, direction vocabulary, and citable evidence are all covered, as are anchoring fallbacks and the baseline edge case. Safety behavior is covered by annotations, so the only mild shortfall is the absence of any usage routing against siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already well documented, including the eventId default and windowMs bounds. The description largely restates the anchoring default and adds no format or constraint detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific analytical operation โ measuring the incident window against an equal preceding baseline, per dimension, with directions and citable evidence. That baseline-comparison framing distinguishes it from plain sibling readers like get_timeline or get_exceptions, and the listed dimensions make the resource concrete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: it is for reconstructing what preceded a failure, and the anchoring rules (eventId, else most recent exception, else now) are stated. There is no explicit when-to-use/when-not or routing to alternatives such as diagnose_performance or explain_diagnosis, which is a real gap given how many overlapping diagnostic siblings exist.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
22 tool updates
v0.21.0- First observed
connect_vm - First observed
diagnose_performance - First observed
diagnose_runtime - First observed
ensure_tcp_device - First observed
explain_diagnosis - First observed
export_session - First observed
get_capabilities - First observed
get_dashboard_url - First observed
get_exceptions - First observed
get_frames - First observed
get_logs - First observed
get_memory - First observed
get_navigation - First observed
get_network - First observed
get_rebuilds - First observed
get_selected_widget - First observed
get_state_activity - First observed
get_timeline - First observed
get_widget_tree - First observed
runtime_health - First observed
runtime_status - First observed
what_changed
TDQS
Scored across 22 tools
Each tool targets a distinct aspect of runtime debugging (logs, exceptions, frames, network, memory, etc.) with minimal overlap. While runtime_health aggregates several get_* tools, its purpose is clearly differentiated from runtime_status and the raw data collectors. Descriptions are detailed enough to guide correct selection.
All tools use lower_snake_case, but the verb_noun pattern is not fully consistent: many use get_ prefix, while others are noun phrases (runtime_status, runtime_health) or phrases (what_changed). Still, the casing and readability are consistent throughout.
At 22 tools, the set is on the heavy side for the stated scope, though each tool has a clear role. This falls into the borderline '16-25 feels heavy' range, as some raw data tools could potentially be consolidated.
The server covers connection, raw data collection, multi-dimensional diagnostics, export, and capabilities, forming a nearly complete lifecycle for runtime observation. Missing only advanced interactions like hot reload or CPU profiling, which are noted as intentional limitations.
Maintenance
Related MCP Connectors
Drive real devices from your AI Coding tool. Embed a client SDK (Unity, Godot, Flutter, iOS/macOS, Android, React Native, Web) in your app, then capture screenshots, traverse the UI tree, inject taps and key events, and run automated test tasks on the physical device over a secure relay.
Live browser debugging for AI assistants โ DOM, console, network via MCP.
Create, edit, preview, and build Flutter apps via FlutterGo.AI cloud MCP (OAuth).
remote debug iOS/Android/Unity/Godot/Flutter/RN/Web on real-device.ui-tree/screenshots/taps,tests.
Related MCP Servers
- AlicenseCqualityCmaintenanceEnables AI assistants to connect to browser DevTools and backend debuggers for full-stack debugging, including frontend console, network, performance, and backend log analysis.1638 npm7MIT
- AlicenseNot gradedqualityDmaintenanceStreams browser DevTools console, network, storage, and performance data to your IDE's AI agent for real-time debugging and analysis.8 npm1MIT
- FlicenseAqualityBmaintenanceConnects AI agents to live observability stacks including Sentry, GitHub, Vercel, Better Stack, and Cloudflare, enabling end-to-end incident investigation, deployment correlation, root cause analysis, and regression triage.14-
- AlicenseNot gradedqualityAmaintenanceExposes 35 stdio JSON-RPC tools that let AI agents snapshot the Flutter Semantics tree, drive gestures, capture screenshots, and observe live app state against running Flutter apps over VM Service.3MIT