Charlotte
Charlotte
La web, legible.
Tu agente de IA consume ~50.000 caracteres del árbol de accesibilidad solo para mirar la portada de Hacker News. Charlotte lo hace en 364.
Charlotte es un servidor MCP que ofrece a los agentes de IA un acceso estructurado y eficiente en tokens a la web. En lugar de volcar el árbol de accesibilidad completo en cada llamada, Charlotte devuelve solo lo que el agente necesita: un resumen compacto de la página al llegar, consultas específicas para elementos concretos y el detalle completo solo cuando se solicita explícitamente. En páginas con mucho contenido, esa orientación es hasta ~140x más pequeña que una instantánea completa del árbol de accesibilidad de Playwright MCP; en páginas trivialmente pequeñas, ambos tienen un tamaño aproximadamente similar.
¿Por qué Charlotte?
La mayoría de los servidores MCP de navegador vuelcan el árbol de accesibilidad completo en cada llamada: un bloque de texto plano que puede superar el millón de caracteres en páginas con mucho contenido. Los agentes pagan por todo ello, lo necesiten o no.
Charlotte descompone cada página en una representación estructurada y tipada: puntos de referencia, encabezados, elementos interactivos, formularios, resúmenes de contenido — y permite a los agentes controlar cuánto reciben con tres niveles de detalle. Cuando un agente navega a una nueva página, recibe una orientación compacta (364 caracteres para Hacker News) en lugar del volcado completo de elementos (~50.000 caracteres). Cuando necesita detalles, los pide.
Puntos de referencia
Medido en Charlotte v0.8.0 contra Playwright MCP v0.0.79, por caracteres devueltos por llamada de herramienta en sitios web reales (npx tsx benchmarks/run-benchmarks.ts --suite comparison), 2026-08-08. Esta sección es un resumen: la página canónica de puntos de referencia (incluido el coste por tarea y la deriva de versiones) está en charlotte.mintlify.site/benchmarks; metodología, instrumentos y resultados brutos: benchmarks/.
Coste de orientación (lo que paga un agente por "ver" una página al llegar):
Un navigate de Charlotte devuelve una orientación utilizable por defecto: puntos de referencia, encabezados y recuentos de elementos interactivos agrupados por región de la página. Para obtener lo equivalente con Playwright MCP, un agente llama a browser_snapshot, que devuelve el árbol de accesibilidad completo. (El browser_navigate de Playwright por sí solo devuelve solo una confirmación breve, no el contenido de la página, por lo que no es una comparación equivalente.)
Sitio |
|
| Más pequeño por |
example.com | 415 | 465 | 1.1x |
formulario httpbin | 619 | 1,847 | 3.0x |
repositorio de GitHub | 3,778 | 38,983 | 10x |
Wikipedia (artículo de IA) | 22,134 | 1,137,928 | 51x |
Hacker News | 364 | 50,706 | 139x |
La ventaja escala con la complejidad de la página: en páginas con mucho contenido, la orientación estructurada es ~10–140x más pequeña que la instantánea completa, mientras que en una página trivialmente pequeña como example.com ambas están dentro de ~20% la una de la otra (y en una página tan pequeña, la representación estructurada puede ser la más grande de las dos — simplemente no hay nada que resumir). El valor de Charlotte se muestra precisamente donde el volcado plano de Playwright duele más. Cuando un agente necesita más que la orientación, llama a observe o find para obtener exactamente la parte que quiere, en lugar de pagar por todo el árbol por adelantado.
Sobrecarga de definición de herramientas (coste invisible por llamada de API):
Perfil | Herramientas | Tokens de definición/llamada | Ahorro vs. completo |
completo | 43 | 8,500 | — |
navegación (por defecto) | 23 | 4,372 | ~49% |
núcleo | 7 | 2,186 | ~75% |
Las definiciones de herramientas se envían en cada ida y vuelta de API. Con el perfil browse por defecto, Charlotte lleva ~49% menos sobrecarga de definición que cargar las 43 herramientas; el perfil mínimo core la reduce en ~75%. Consulta el informe de puntos de referencia de perfiles para ver los resultados completos.
La diferencia en el flujo de trabajo: Un agente de Playwright que lee la instantánea completa recibe ~50.000 caracteres cada vez que mira Hacker News, ya sea leyendo titulares o buscando un botón de inicio de sesión. Un agente de Charlotte recibe 364 caracteres al llegar, llama a find({ type: "link", text: "login" }) para obtener exactamente lo que necesita y nunca paga por el resto.
Related MCP server: krwl3r
Cómo funciona
Charlotte mantiene una sesión persistente de Chromium sin interfaz gráfica y actúa como una capa de traducción entre la web visual y el razonamiento nativo de texto del agente. Cada página se descompone en una representación estructurada:
┌─────────────┐ MCP Protocol ┌──────────────────┐
│ AI Agent │<────────────────────>│ Charlotte │
└─────────────┘ │ │
│ ┌────────────┐ │
│ │ Renderer │ │
│ │ Pipeline │ │
│ └─────┬──────┘ │
│ │ │
│ ┌─────▼──────┐ │
│ │ Headless │ │
│ │ Chromium │ │
│ └────────────┘ │
└──────────────────┘Los agentes reciben puntos de referencia, encabezados, elementos interactivos con metadatos tipados, cuadros delimitadores, estructuras de formularios y resúmenes de contenido, todo derivado de lo que el navegador ya sabe sobre cada página.
Características
Navegación — navigate, back, forward, reload
Observación — observe (3 niveles de detalle, vista de árbol estructural), find (búsqueda espacial y semántica, modo selector CSS, output_file para conjuntos de resultados grandes), screenshot (con gestión persistente de artefactos), screenshots, screenshot_get, screenshot_delete, diff (comparación estructural contra instantáneas)
Interacción (consciente de iframes) — click, click_at (basado en coordenadas), type (con soporte de escritura lenta), select, toggle, submit, scroll, hover, drag, key (única/secuencia con selección de elemento), wait_for (sondeo de condiciones asíncronas), upload (entrada de archivos), fill_form (relleno de formularios por lotes), dialog (aceptar/descartar diálogos JS)
Monitorización — console (todos los niveles de gravedad, filtrado, marcas de tiempo), requests (historial HTTP completo, filtrado por método/estado/tipo de recurso)
Gestión de sesión — tabs, tab_open, tab_switch, tab_close, viewport (preajustes genéricos o dispositivos con nombre como "iPhone 15" con emulación de DPR, táctil y agente de usuario), network (limitación de ancho de banda, bloqueo de URL), set_cookies, get_cookies, clear_cookies, set_headers, configure
Modo de desarrollo — dev_serve (servidor estático + observación de archivos con recarga automática), dev_inject (inyección de CSS/JS), dev_audit (a11y, rendimiento, SEO, contraste, enlaces rotos)
Utilidades — evaluate (ejecución de JS arbitrario en el contexto de la página)
Perfiles de herramientas
Charlotte incluye 43 herramientas (42 registradas + la meta-herramienta charlotte_tools), pero la mayoría de los flujos de trabajo solo necesitan un subconjunto. Los perfiles de inicio controlan qué herramientas se cargan en el contexto del agente, reduciendo la sobrecarga de definición hasta en ~75%.
charlotte --profile browse # 23 tools (default) — navigate, observe, interact, tabs
charlotte --profile core # 7 tools — navigate, observe, find, click, type, submit
charlotte --profile full # 43 tools — everything
charlotte --profile interact # 31 tools — full interaction + dialog + evaluate
charlotte --profile develop # 34 tools — interact + dev_serve, dev_inject, dev_audit
charlotte --profile audit # 14 tools — navigation + observation + dev_audit + viewportLos agentes pueden activar más herramientas a mitad de sesión sin reiniciar:
charlotte_tools enable dev_mode → activates dev_serve, dev_audit, dev_inject
charlotte_tools disable dev_mode → deactivates them
charlotte_tools list → see what's loadedAutoalojamiento (Charlotte Remoto)
Ejecuta Charlotte como un servidor MCP remoto y conéctalo a claude.ai — un solo comando:
docker run --cap-add SYS_ADMIN --shm-size 2g -p 3737:3737 ghcr.io/ticktockbent/charlotteImprime una URL de conector pública y un token de operador. En claude.ai: Configuración → Conectores → Añadir conector personalizado — pega la URL, deja los campos de ID de cliente OAuth/Secreto en blanco e introduce el token en la página de consentimiento de Charlotte cuando aparezca. Eso es todo; ya estás navegando.
La URL de demostración y el token son efímeros (ambos rotan al reiniciar). Para ejecutarlo de verdad — dominio estable, tu propio túnel o proxy inverso, docker compose: Autoalojamiento. Modelo de confianza y protecciones de red: Seguridad. Internals del contenedor y sandbox: Docker.
Inicio rápido
Requisitos previos
Node.js >= 20
npm
Instalación
Charlotte está listado en el Registro MCP como io.github.TickTockBent/charlotte y publicado en npm como @ticktockbent/charlotte:
npm install -g @ticktockbent/charlotteLas imágenes de Docker están disponibles en Docker Hub y en GitHub Container Registry:
# Alpine (default, smaller)
docker pull ticktockbent/charlotte:alpine
# Debian (if you need glibc compatibility)
docker pull ticktockbent/charlotte:debian
# Or from GHCR
docker pull ghcr.io/ticktockbent/charlotte:latestO instala desde el código fuente:
git clone https://github.com/ticktockbent/charlotte.git
cd charlotte
npm install
npm run buildEjecutar
Charlotte se comunica a través de stdio usando el protocolo MCP:
# If installed globally (default browse profile)
charlotte
# With a specific profile
charlotte --profile core
# If installed from source
npm startConfiguración del cliente MCP
Claude Code
Crea .mcp.json en la raíz de tu proyecto:
{
"mcpServers": {
"charlotte": {
"type": "stdio",
"command": "npx",
"args": ["@ticktockbent/charlotte"],
"env": {}
}
}
}Claude Desktop
Añade a claude_desktop_config.json:
{
"mcpServers": {
"charlotte": {
"command": "npx",
"args": ["@ticktockbent/charlotte"]
}
}
}Cursor
Añade a .cursor/mcp.json:
{
"mcpServers": {
"charlotte": {
"command": "npx",
"args": ["@ticktockbent/charlotte"]
}
}
}Windsurf
Añade a ~/.codeium/windsurf/mcp_config.json:
{
"mcpServers": {
"charlotte": {
"command": "npx",
"args": ["@ticktockbent/charlotte"]
}
}
}VS Code (Copilot)
Añade a .vscode/mcp.json:
{
"servers": {
"charlotte": {
"type": "stdio",
"command": "npx",
"args": ["@ticktockbent/charlotte"]
}
}
}Cline
Añade a la configuración de MCP de Cline (a través de la barra lateral de Cline > Servidores MCP > Configurar):
{
"mcpServers": {
"charlotte": {
"command": "npx",
"args": ["@ticktockbent/charlotte"]
}
}
}Amp
Añade a ~/.amp/settings.json:
{
"mcpServers": {
"charlotte": {
"command": "npx",
"args": ["@ticktockbent/charlotte"]
}
}
}Consulta docs-internal/mcp-setup.md para la guía de configuración completa, incluido el modo de desarrollo, clientes MCP genéricos, pasos de verificación y solución de problemas.
Configuración
Charlotte resuelve la configuración de cuatro fuentes, con la precedencia más alta primero: argumentos de CLI → variables de entorno → archivo de configuración → valores predeterminados integrados. Consulta docs/configuration.md para la referencia completa.
Archivo de configuración
Pasa un archivo de configuración JSON con --config, o coloca un charlotte.config.json en el directorio de trabajo y Charlotte lo carga automáticamente:
charlotte --config charlotte.config.json{
"browser": { "headless": true, "noSandbox": false },
"tools": { "profile": "browse" },
"rendering": { "includeIframes": false, "iframeDepth": 3 },
"output": { "dir": "./charlotte-output" },
"limits": {
"maxInteractiveElements": 2000,
"maxFullContentChars": 200000,
"maxResponseBytes": 1000000,
"maxEvaluateBytes": 256000
}
}Cada sección es opcional; un {} vacío es válido. El archivo se valida con zod: claves desconocidas, tipos incorrectos o valores de enumeración no válidos producen un error de inicio claro en stderr y Charlotte sale con un código distinto de cero. Cuatro ajustes también tienen variables de entorno: CHARLOTTE_NO_SANDBOX, CHARLOTTE_OUTPUT_DIR, CHARLOTTE_CDP_ENDPOINT y CHARLOTTE_INIT_SCRIPT. Los scripts que deben ejecutarse en cada documento nuevo antes del JS de la página van en browser.initScripts o --init-script <ruta> (repetible); consulta Scripts de inicio.
El sandbox de Chromium está activado por defecto
Cambio de comportamiento en v0.7.0: Las versiones anteriores incluían
--no-sandboxen cada lanzamiento de Chromium. A partir de v0.7.0, el sandbox de Chromium está habilitado por defecto — la principal defensa entre una página no confiable y la cuenta con la que se ejecuta Charlotte. Debes optar por desactivarlo explícitamente donde el sandbox del kernel no esté disponible.
charlotte --no-sandbox # CLI flag
CHARLOTTE_NO_SANDBOX=1 charlotte # environment variable
# or "browser": { "noSandbox": true } in the config fileNota de migración (Docker / metal desnudo): Los contenedores normalmente no pueden configurar el sandbox del kernel, por lo que los Dockerfiles proporcionados establecen CHARLOTTE_NO_SANDBOX=1 por ti, y docker-compose.yml ahora mantiene el filtro seccomp predeterminado de Docker (ya no ejecuta seccomp=unconfined). Si ejecutas Charlotte en metal desnudo como root, Chromium se niega a iniciarse con el sandbox habilitado — ejecútalo como un usuario no root (recomendado) o pasa --no-sandbox. Las configuraciones existentes que antes dependían del --no-sandbox implícito y se ejecutan en un entorno donde el sandbox no puede inicializarse ahora deben establecer CHARLOTTE_NO_SANDBOX=1 (o el indicador/equivalente de configuración) para seguir funcionando.
Ejecutar Charlotte Remoto (modo HTTP) a través de la red plantea preguntas adicionales sobre límites de confianza y protecciones de red más allá del sandbox — consulta Seguridad.
Límites de tamaño de salida
Las claves limits.* limitan cuánto puede devolver una sola respuesta de herramienta para que una página patológica (100k enlaces, un feed de desplazamiento infinito, un cuerpo de documento gigante) no pueda desbordar la ventana de contexto del agente. Cuando una respuesta de página supera maxResponseBytes, se degrada a un resumen compacto y sugiere escribir el resultado completo en disco mediante output_file; los resultados de charlotte_evaluate se limitan de forma independiente con maxEvaluateBytes. Las respuestas truncadas llevan un marcador truncation. Consulta docs/configuration.md para las claves y los valores predeterminados.
Recuperación de fallos
Un fallo de Chromium ya no bloquea el servidor. La siguiente llamada a una herramienta relanza automáticamente el navegador, limpia la pestaña muerta y las cachés de sesión CDP, y abre una nueva pestaña en blanco — de modo que un agente puede seguir trabajando después de un fallo del renderizador sin reiniciar el servidor MCP.
Ejemplos de uso
Una vez conectado, un agente puede usar las herramientas de Charlotte:
Navegar por un sitio web
navigate({ url: "https://example.com" })
// → 612 chars: landmarks, headings, interactive element counts
find({ type: "link", text: "More information" })
// → just the matching element with its ID
click({ element_id: "lnk-a3f1c2" })Rellenar un formulario
navigate({ url: "https://httpbin.org/forms/post" })
find({ type: "text_input" })
type({ element_id: "inp-c7e29b", text: "hello@example.com" })
select({ element_id: "sel-e8a3f5", value: "option-2" })
submit({ form_id: "frm-b1d4e7" })Bucle de retroalimentación para desarrollo local
dev_serve({ path: "./my-site", watch: true })
observe({ detail: "full" })
dev_audit({ checks: ["a11y", "contrast"] })
dev_inject({ css: "body { font-size: 18px; }" })Representación de la página
Charlotte devuelve representaciones estructuradas con tres niveles de detalle que permiten a los agentes controlar cuánto contexto consumen:
Mínimo (predeterminado para navigate)
Hitos (landmarks), encabezados y recuentos de elementos interactivos agrupados por región de la página. Diseñado para la orientación — «¿qué hay en esta página?» — sin enumerar cada elemento.
{
"url": "https://news.ycombinator.com",
"title": "Hacker News",
"viewport": { "width": 1280, "height": 720 },
"structure": {
"headings": [{ "level": 1, "text": "Hacker News", "id": "hdg-a1b2c3" }]
},
"interactive_summary": {
"total": 93,
"by_landmark": {
"(page root)": { "link": 91, "text_input": 1, "button": 1 }
}
}
}Resumen (predeterminado para observe)
Lista completa de elementos interactivos con metadatos tipados, estructuras de formularios y resúmenes de contenido.
{
"url": "https://example.com/dashboard",
"title": "Dashboard",
"viewport": { "width": 1280, "height": 720 },
"structure": {
"landmarks": [
{ "id": "rgn-b2c1d0", "role": "banner", "label": "Site header", "bounds": { "x": 0, "y": 0, "w": 1280, "h": 64 } },
{ "id": "rgn-d4e5f6", "role": "main", "label": "Content", "bounds": { "x": 240, "y": 64, "w": 1040, "h": 656 } }
],
"headings": [{ "level": 1, "text": "Dashboard", "id": "hdg-1a2b3c" }],
"content_summary": "main: 2 headings, 5 links, 1 form"
},
"interactive": [
{
"id": "btn-a3f1c2",
"type": "button",
"label": "Create Project",
"bounds": { "x": 960, "y": 80, "w": 160, "h": 40 },
"state": {}
}
],
"forms": []
}Completo
Todo lo del resumen, más todo el contenido de texto visible en la página.
Niveles de detalle
Nivel | Tokens | Caso de uso |
| ~50-200 | Orientación tras la navegación. ¿Qué regiones existen? ¿Cuántos elementos interactivos hay? |
| ~500-5000 | Trabajo con la página. Lista completa de elementos, estructuras de formularios, resúmenes de contenido. |
| variable | Lectura del contenido de la página. Todo el texto visible incluido. |
Las herramientas de navegación usan minimal por defecto. La herramienta observe usa summary por defecto. Ambas aceptan un parámetro opcional detail para anularlo.
IDs de elementos
Los IDs de elementos son estables ante mutaciones menores del DOM. Se generan aplicando un hash a una clave compuesta por el tipo de elemento, el rol ARIA, el nombre accesible y la firma de la ruta DOM:
btn-a3f1c2 (button) inp-c7e29b (text input)
lnk-d4b910 (link) sel-e8a3f5 (select)
chk-f1a204 (checkbox) frm-b1d4e7 (form)
rgn-e0d2a8 (landmark) hdg-0f4063 (heading)
dom-b2c3d9 (DOM element, from CSS selector queries)Cambio de formato de ID en v0.7.0: los hashes de IDs de elementos ahora son de 6 caracteres hexadecimales (p. ej.
btn-a3f1c2), frente a los 4 de versiones anteriores. Esto reduce drásticamente las colisiones de hash entre elementos en páginas grandes. Los agentes que hayan codificado o hecho coincidencia de patrones con IDs de 4 caracteres deberían volver afindlos elementos en lugar de reutilizar IDs en caché tras la actualización.
Los IDs sobreviven a cambios no relacionados del DOM y a la reordenación de elementos dentro del mismo contenedor. Cuando un agente navega con detalle mínimo (sin IDs de elementos individuales), usa find para localizar elementos por texto, tipo o proximidad espacial; los elementos devueltos incluyen IDs listos para la interacción.
Desarrollo
# Run in watch mode
npm run dev
# Run all tests
npm test
# Run only unit tests
npm run test:unit
# Run only integration tests
npm run test:integration
# Type check
npx tsc --noEmitEstructura del proyecto
src/
browser/ # Puppeteer lifecycle, tab management, CDP sessions
renderer/ # Accessibility tree extraction, layout, content, element IDs
state/ # Snapshot store, structural differ
tools/ # MCP tool definitions (navigation, observation, interaction, session, dev-mode)
dev/ # Static server, file watcher, auditor
types/ # TypeScript interfaces
utils/ # Logger, hash, wait utilities
tests/
unit/ # Fast tests with mocks
integration/ # Full Puppeteer tests against fixture HTML
fixtures/pages/ # Test HTML filesArquitectura
El Renderer Pipeline es el núcleo: llama a los extractores en orden y ensambla una PageRepresentation:
Extracción del árbol de accesibilidad (CDP
Accessibility.getFullAXTree)Extracción del diseño (CDP
DOM.getBoxModel)Extracción de hitos, encabezados, elementos interactivos y contenido
Generación de IDs de elementos (basada en hash, estable entre re-renderizados)
Todas las herramientas pasan por renderActivePage(), que gestiona instantáneas, eventos de recarga, detección de diálogos y el formato de las respuestas.
Sandbox
Charlotte incluye un sitio web de prueba en tests/sandbox/ que ejercita todas las herramientas sin tocar la internet pública. Sírvelo localmente con:
dev_serve({ path: "tests/sandbox" })Cinco páginas cubren navegación, formularios, elementos interactivos, ventanas emergentes, contenido retardado, contenedores de desplazamiento y más. Consulta docs-internal/sandbox.md para la referencia completa de páginas y una lista de ejercicios herramienta por herramienta.
Problemas conocidos
Shadow DOM — El shadow DOM abierto funciona de forma transparente. El árbol de accesibilidad de Chromium atraviesa los límites del shadow DOM abierto, por lo que los componentes web (p. ej., <relative-time> y <tool-tip> de GitHub) renderizan su contenido en la representación de Charlotte sin tratamiento especial. Las raíces shadow cerradas son opacas para el árbol de accesibilidad y no se capturarán.
Hoja de ruta
Sesión y configuración
Hoja de ruta de funciones
Grabación de vídeo — Graba las interacciones como vídeo, capturando la secuencia completa de navegación y manipulación dirigida por el agente para depuración, documentación y revisión.
Consulta docs-internal/playwright-mcp-gap-analysis.md para el análisis completo de brechas frente a Playwright MCP, incluidos los elementos de menor prioridad (herramientas de visión, pruebas/verificación, trazado, transporte, seguridad) y las áreas donde Charlotte tiene ventajas.
Especificación completa
Consulta docs-internal/CHARLOTTE_SPEC.md para la especificación completa, incluidos todos los parámetros de las herramientas, el formato de representación de la página, la estrategia de identidad de los elementos y los detalles de arquitectura.
Licencia
Comunidad
Abre un informe de error para defectos reproducibles, regresiones o problemas específicos de clientes MCP.
Abre una solicitud de función para mejoras de flujo de trabajo o nuevas capacidades.
Abre una solicitud de herramienta si quieres proponer una nueva herramienta, superficie de parámetros o ubicación de perfil.
Explora los problemas abiertos para encontrar trabajo actual y discusiones.
Consulta el filtro de buenos primeros problemas planificado, ya que los mantenedores etiquetan tareas aptas para principiantes.
Contribuciones
Consulta CONTRIBUTING.md para las directrices.
Parte de un conjunto creciente de servidores MCP con nombres literarios. Ver más en github.com/TickTockBent.
Available Tools
23 toolscharlotte_backA
Navigate back in browser history. Returns page representation after navigation.
| Name | Required | Description | Default |
|---|---|---|---|
| detail | No | "minimal" (default), "summary" (includes content context), "full" (includes all text content) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the main behavior (navigating back) and the return value (page representation), which provides some context. However, it omits important edge-case details such as behavior when there is no history, whether it waits for page load, and the exact nature of the 'page representation,' leaving gaps in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences, front-loading the action and then adding the return type. There is no filler, redundancy, or unnecessary detail; every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with one well-documented parameter and no annotations or output schema. The description covers the core purpose and return, which is largely sufficient. It could be improved by mentioning how the 'detail' parameter affects the returned representation or failure behavior with an empty history, but overall it is complete for a simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides 100% coverage for the single optional 'detail' parameter, including its enum and description. The tool description adds no additional parameter information, but the baseline of 3 applies given the high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Navigate back') and resource ('browser history'), clearly distinguishing this from sibling tools like charlotte_forward and charlotte_navigate. It also states the return value ('page representation after navigation'), leaving no ambiguity about its function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by stating the action ('navigate back'), but it does not explicitly mention when to prefer this over alternatives (e.g., charlotte_forward, charlotte_navigate) or provide exclusions. There is no guidance on scenarios like empty history, but the action itself is a clear directive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_clickA
Click an interactive element on the page. Returns full page representation after the click.
| Name | Required | Description | Default |
|---|---|---|---|
| modifiers | No | Modifier keys to hold during click: ["ctrl"], ["shift"], ["alt"], ["meta"], or combinations like ["ctrl", "shift"] | |
| click_type | No | Click type: "left" (default), "right", "double" | |
| element_id | Yes | Target element ID from page representation |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full transparency burden. It does disclose that the tool returns a full page representation after the click, which is useful. However, it does not mention potential side effects like navigation, form submission, or waiting behavior, which are important for a click action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, concise and front-loaded with the action and return value. Every word earns its place, with no fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a basic click tool, covering the action and return value. However, it lacks any mention of the sibling tool charlotte_click_at, does not explain what 'full page representation' entails, and omits potential side-effect warnings. More context is needed for full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides descriptions for all three parameters (element_id, click_type, modifiers) with 100% coverage. The description itself adds no extra parameter context, so it relies on the schema, matching the baseline score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool clicks an interactive element on the page, which distinguishes it from scrolling, typing, and navigation. However, it does not explicitly differentiate from charlotte_click_at, which is a sibling tool likely for coordinate-based clicks, leaving the distinction implied through the word 'element'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'interactive element' implies element-based clicking, which suggests using this tool over charlotte_click_at, but it does not explicitly state when to use this tool versus alternatives or provide exclusions. The guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_click_atA
Click at specific page coordinates. Use when target elements are not in the accessibility tree (custom widgets, canvas, non-semantic interactive divs). Dispatches real CDP-level mouse events. Returns full page representation after the click.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | X coordinate in page pixels | |
| y | Yes | Y coordinate in page pixels | |
| modifiers | No | Modifier keys to hold during click: ["ctrl"], ["shift"], ["alt"], ["meta"], or combinations | |
| click_type | No | Click type: "left" (default), "right", "double" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full responsibility for behavioral disclosure. It reveals that the click dispatches 'real CDP-level mouse events' and returns a 'full page representation' after the click. While it doesn't cover every edge case (e.g., out-of-viewport coordinates, waiting for navigation), it provides the core behavioral traits relevant to a coordinate click tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: action statement, usage context, and return behavior. No filler or redundancy. The structure is front-loaded with the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (4 params, no output schema), and the description covers the return value ('full page representation') and the type of events dispatched. It doesn't mention prerequisites like page load state, but such details are likely unnecessary for a click tool with clear semantics. Overall, it is sufficiently complete for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents all four parameters with clear descriptions, including enum values for modifiers and click_type (100% coverage). The description adds no additional parameter semantics beyond labeling x/y as 'page coordinates', so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Click at specific page coordinates' with a specific verb and resource. It also distinguishes itself from the sibling element-based click tool by noting it's for when elements are not in the accessibility tree (custom widgets, canvas, non-semantic interactive divs).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly provides when-to-use guidance: 'Use when target elements are not in the accessibility tree...' This implies not to use it when elements are accessible, effectively differentiating it from alternatives like charlotte_click. The mention of CDP-level events preempts expectations about event orchestration.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_diffA
Compare current page state to a previous snapshot. Returns structural diff showing added, removed, moved, and changed elements.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | No | "all" (default), "structure" (landmarks/headings), "interactive" (elements/forms), "content" (text/url/title) | |
| snapshot_id | No | Compare against a specific snapshot ID (default: previous snapshot) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the transparency burden. It discloses the output type (structural diff with added/removed/moved/changed elements) and implies a read-only comparison, but it doesn't explicitly state that it is non-destructive or mention prerequisites like the existence of a previous snapshot. Some behavioral context is added, but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is composed of two short sentences that front-load the primary action and the return value. There is no redundant phrasing or filler, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two optional, well-documented parameters and no output schema, the description provides the core purpose and a high-level summary of return categories. It does not explain snapshot prerequisites or error behavior, but the combination of description and schema is sufficient for the tool's apparent simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers both parameters 100%, with scope enum descriptions and snapshot_id details. The tool description itself adds no additional parameter semantics beyond what the schema already provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Compare' with a clear resource ('current page state to a previous snapshot') and states the return type ('structural diff showing added, removed, moved, and changed elements'). This distinguishes it from sibling tools that operate on the live page (e.g., click, type, observe) and makes the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear context: use this tool when you need to see differences between the current page state and a previous snapshot. It doesn't explicitly list exclusions or alternatives, but the diff focus is obvious among the siblings, which are mostly navigation and interaction tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_findA
Search for elements matching criteria. Filters interactive elements by text, role, type, or spatial proximity. Use the selector parameter to find DOM elements by CSS selector — this reaches elements not in the accessibility tree (custom widgets, non-semantic divs). Selector results return Charlotte element IDs usable with click, hover, drag, etc.
| Name | Required | Description | Default |
|---|---|---|---|
| near | No | Element ID — find elements spatially near this one (within ~200px) | |
| role | No | ARIA role filter | |
| text | No | Text content to search for (case-insensitive substring match) | |
| type | No | Interactive element type filter (button, link, text_input, select, checkbox, etc.) | |
| within | No | Element ID — find elements geometrically contained within this one's bounds | |
| selector | No | CSS selector to query the DOM directly. Returns elements that may not be in the accessibility tree. Results include durable Charlotte element IDs (dom-…) that remain valid across subsequent renders and interactions, and work with fill_form; they are re-resolved against the live DOM by re-running the selector. | |
| output_file | No | Write the full match results to this file path instead of returning them inline. Relative paths resolve against output_dir (see charlotte_configure). Returns only a confirmation with the file path and size. Use for broad selectors (e.g. 'div', '*') that match many elements. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden and discloses important behaviors: selector reaches non-accessibility-tree elements, results return durable Charlotte element IDs re-resolved against the live DOM, and output_file writes to a file with a confirmation. It omits details on behavior when no filters are provided or whether hidden elements are included, leaving minor gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: lead purpose statement, filter list, selector explanation, then output_file behavior. Every clause adds functional detail without redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main search use case, the unique selector capability, and output_file output, but does not specify the return format for inline results or behavior when no criteria are supplied. This is a minor gap given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds meaning beyond the schema for selector (durable IDs, re-resolution, works with fill_form) and output_file (confirmation with path and size, use for broad selectors), justifying a score above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Search for elements matching criteria' with specific filters (text, role, type, spatial proximity) and highlights the selector parameter for DOM access. This distinguishes it from sibling action tools like click, type, and navigate by focusing on search/find.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly recommends using the selector parameter to find elements not in the accessibility tree, providing clear contextual guidance. However, it does not explicitly state when not to use the tool or compare it to alternatives, so it stops short of full exclusionary usage advice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_forwardA
Navigate forward in browser history. Returns page representation after navigation.
| Name | Required | Description | Default |
|---|---|---|---|
| detail | No | "minimal" (default), "summary" (includes content context), "full" (includes all text content) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the return behavior ('Returns page representation after navigation'), but does not specify edge cases like what happens when there is no forward history, or how 'page representation' is structured. This is adequate but minimal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with no wasted words. It front-loads the primary action and then describes the return value, making it easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (one optional parameter, no output schema), and the description covers the essential purpose and return behavior. It could mention edge cases, but given the low complexity, it is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single optional 'detail' parameter, with enum values and descriptions. The tool description does not add parameter information, but the schema fully documents it, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Navigate forward in browser history', using a specific verb and resource. It distinguishes itself from sibling tools like charlotte_back (backward navigation) and charlotte_navigate (URL navigation).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use this tool (when the user wants to go forward in browser history). It doesn't explicitly mention alternatives or exclusions, but the context is clear enough given the sibling tool names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_observeA
Get current page state without performing any action. Use detail levels to control verbosity: "minimal" for landmarks, headings, and interactive element counts by landmark (use charlotte_find to get specific elements with actionable IDs, or observe({ detail: "summary" }) to see all elements), "summary" (default) for content summaries and full element list, "full" for all text content. Use view: "tree" for a compact structural outline (cheapest orientation tool), or view: "tree-labeled" to include labels on interactive elements (still much cheaper than minimal JSON, and shows which button/link/input is which).
| Name | Required | Description | Default |
|---|---|---|---|
| view | No | "default" (structured JSON), "tree" (compact structural outline — element types only, cheapest), or "tree-labeled" (structural outline with interactive element labels — shows which button/link/input is which, still ~70% cheaper than minimal JSON) | |
| detail | No | "summary" (default), "full" (includes all text content), "minimal" (landmarks + interactive only) | |
| selector | No | CSS selector to scope observation to a subtree | |
| output_file | No | Write observation data to this file path instead of returning inline. Relative paths resolve against output_dir (see charlotte_configure). Returns only a confirmation with the file path and size. | |
| include_styles | No | Include computed styles for visible elements (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It states 'without performing any action' (non-mutating), explains cost trade-offs ('cheapest', 'still much cheaper than minimal JSON'), and reveals output behavior for output_file ('Returns only a confirmation with the file path and size'). This provides meaningful context beyond the schema, though it omits potential error conditions or page-load requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured: it leads with the main purpose, then explains optional parameters in a logical flow. Each sentence adds operational value without filler. Slightly longer than ideal, but the complexity of 5 parameters and absence of annotations justify the length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (5 params, no output schema, no annotations), the description covers the essential aspects: purpose, parameter behaviors, alternatives, and cost considerations. It lacks explicit return format details for the default view, but for a read-only observation tool, the provided information is sufficient for an agent to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds significant value by explaining detail level semantics (minimal/summary/full), view trade-offs (tree vs tree-labeled with cost estimates), and output_file resolution ('relative paths resolve against output_dir'). This goes beyond the bare enum names and provides actionable guidance for parameter selection.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Get current page state without performing any action.' The verb 'get' and resource 'page state' precisely convey the read-only nature, distinguishing it from action-oriented siblings like charlotte_click and charlotte_type. Mentioning charlotte_find and observe variants further differentiates it from element-finding tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly directs when to use alternatives: 'use charlotte_find to get specific elements with actionable IDs' and 'or observe({ detail: "summary" }) to see all elements.' Recommends view: 'tree' as the 'cheapest orientation tool,' giving clear context on when to choose this tool over others.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_reloadA
Reload the current page. Returns page representation after reload.
| Name | Required | Description | Default |
|---|---|---|---|
| hard | No | Bypass cache (default: false) | |
| detail | No | "minimal" (default), "summary" (includes content context), "full" (includes all text content) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It does state that the tool returns a page representation after reload, which is valuable. However, it does not mention potential side effects such as losing unsaved form state, and the 'hard' parameter's cache-bypass behavior is only in the schema, not the description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One concise sentence that immediately states the action and outcome. Every word earns its place with no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple reload action with two fully documented optional parameters and no output schema, the description provides adequate context: it says what the tool does and what it returns. Slightly more detail on how 'detail' affects the returned representation could improve it, but the schema already covers that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: both 'hard' and 'detail' already have meaningful descriptions in the input schema. The tool description adds no additional parameter semantics, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Reload') with a clear resource ('current page') and explicitly states the return value ('Returns page representation after reload'). It is distinct from sibling navigation tools like charlotte_navigate, charlotte_back, and charlotte_forward.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage context is implied: reload the current page when a refresh is needed. However, there are no explicit guidelines on when to choose this over alternatives like navigate or back/forward, nor any exclusionary conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_screenshotA
Capture a visual screenshot. Fallback for when structured representation isn't sufficient (complex visualizations, canvas elements, images). Use save: true to persist as a file artifact that can be referenced later.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | Save as a persistent file artifact (default: false). When true, the screenshot is written to disk and artifact metadata is returned alongside the image. | |
| format | No | "png" (default), "jpeg", "webp" | |
| quality | No | 1-100 for jpeg/webp quality | |
| selector | No | CSS selector to capture specific element (default: full page) | |
| full_page | No | Capture the entire scrollable page (default: true). Set false to capture only the current viewport — much smaller output for long pages. Ignored when 'selector' is provided. | |
| output_file | No | Write screenshot to this file path instead of returning base64 inline. Relative paths resolve against output_dir (see charlotte_configure). Returns only a confirmation with the file path and size. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full transparency burden. It adds useful context about the persistent behavior of save:true and the tool's fallback role. However, it does not disclose the default return format (e.g., inline base64), whether the operation has side effects beyond saving, or that it is a read-only action. These are notable omissions for an unannotated tool, but the provided behavior hints (persistence, fallback) prevent a lower score.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long, front-loaded with the primary verb behavior ('Capture a visual screenshot'), then a use-case clause, then a targeted parameter tip. Every sentence earns its place; no filler or redundant restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 6 parameters, no output schema, and no annotations, the description does more than a minimal effort by establishing purpose and usage guidance. However, it leaves important contextual gaps: the default return format (base64) is not mentioned, and there is no hint about how to later retrieve saved artifacts (e.g., charlotte_screenshot_get). This is adequate but not fully complete for an agent making an uninformed invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all six parameters with meaningful descriptions. The description adds a bit of extra context for 'save' (persisting as a file artifact) but largely repeats what the schema explains. It neither contradicts nor significantly enhances the schema's parameter documentation, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource phrase ('Capture a visual screenshot'), clearly stating the tool's core function. It further distinguishes the tool from siblings by framing it as a 'fallback for when structured representation isn't sufficient,' referencing complex visualizations, canvas elements, and images—which implicitly contrasts with observation-oriented siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use context: 'Fallback for when structured representation isn't sufficient' and gives concrete example cases (complex visualizations, canvas elements, images). It also offers conditional parameter guidance ('Use save: true to persist as a file artifact'). It stops short of naming a specific alternative sibling tool or outright saying 'do not use for simple text/structured content,' so it earns a 4 rather than 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_screenshot_deleteA
Delete a saved screenshot artifact by its ID. Removes the file from disk.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Screenshot artifact ID to delete |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of disclosing side effects. It explicitly states 'Removes the file from disk,' which conveys the destructive, persistent nature of the operation. It does not detail error handling or permissions, but for a simple delete operation, this is substantial disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences, front-loads the core action, and contains no filler or repetition. Every word contributes meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple deletion tool with one parameter and no output schema, the description adequately covers the action and its effect. It does not mention how to obtain the ID or that the deletion is permanent, but the essentials are present, and the simplicity of the tool reduces the need for further context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already has 100% coverage, documenting the single 'id' parameter as 'Screenshot artifact ID to delete.' The description adds little beyond restating 'by its ID,' so the schema carries the semantic weight. This matches the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Delete') and clearly identifies the resource ('saved screenshot artifact') and the required identifier ('by its ID'). It is unambiguous and distinguishes this tool from sibling tools like charlotte_screenshot_get, which retrieves rather than deletes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool—when a saved screenshot artifact needs to be deleted—but gives no explicit guidance about when not to use it or alternatives. It does not mention that this is the only tool for deletion or that retrieval tools should be used for other purposes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_screenshot_getA
Retrieve a previously saved screenshot artifact by its ID. Returns the image data and metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Screenshot artifact ID (e.g. ss-20260224103000-a1b2c3) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. The verb 'Retrieve' clearly implies a non-destructive read-only operation, and 'Returns the image data and metadata' sets expectations for the output. No side effects or special requirements are mentioned, but for a simple retrieval tool this is sufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that states the action, the target resource, and the return value without any redundant wording or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has only one parameter and no output schema, so the description needs to explain the outcome, which it does ('Returns the image data and metadata'). For a simple getter, this is fully complete and requires no additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers the only parameter 'id' with a clear description and format example. The tool description adds only 'by its ID', which reinforces the schema without adding new semantic meaning. Given 100% schema coverage, the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Retrieve') and resource ('saved screenshot artifact by its ID'), making the operation clear. It also distinguishes from sibling tools like charlotte_screenshot (which likely captures a new screenshot) and charlotte_screenshot_delete (which deletes) by focusing on retrieval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context that this tool is for previously saved screenshots, implying it should be used when an artifact ID is already available. It does not explicitly name alternatives or exclusions, but the sibling tool names and the phrase 'previously saved' give practical usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_screenshotsA
List all saved screenshot artifacts. Returns metadata for each saved screenshot including ID, filename, page URL, and timestamp.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses that the tool returns metadata (not image contents), which is useful. However, it doesn't mention pagination, limiting, ordering, or whether the list spans multiple sessions. The read-only nature is implied by 'List' but not explicitly stated. There is no contradiction with annotations since none exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description consists of two short sentences that are direct and information-dense. Every word contributes: 'List', 'all saved screenshot artifacts', and the specific metadata fields. No fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (no parameters, no output schema), the description is largely complete. It clearly states the tool lists all saved artifacts and specifies the metadata fields returned. It does not explain edge cases like empty lists or error scenarios, but for a list operation with no inputs, the provided information is sufficient for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema coverage is 100% (empty properties). The description adds no parameter details, but none are needed. By the rubric, zero params gives a baseline of 4, and the description does provide additional context about what the returned metadata includes (ID, filename, page URL, timestamp) which is helpful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('List all saved screenshot artifacts') and the resource (saved screenshots). It distinguishes from sibling tools like charlotte_screenshot (capture) and charlotte_screenshot_get (fetch one) by focusing on listing all saved artifacts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly indicates when to use the tool: when you need to list all saved screenshot artifacts. It doesn't explicitly mention exclusions or alternatives, but the context is unambiguous given the sibling tool names. A minor gap is lack of guidance on when to use this vs. charlotte_screenshot_get or charlotte_screenshot_delete, but the primary use case is evident.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_scrollA
Scroll the page or a specific container. Returns full page representation after scrolling.
| Name | Required | Description | Default |
|---|---|---|---|
| amount | No | Scroll amount: "page" (default), "half", or pixel value (e.g. "200") | |
| direction | Yes | Scroll direction | |
| element_id | No | Scroll within a specific container element |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It does disclose that the tool returns a full page representation after scrolling, which is useful behavioral context. However, it does not mention any potential side effects, prerequisites, or the nature of the scroll (e.g., instant, smooth). The safety/read-only nature is inferred but not stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences: one for the action and one for the return value. There is no redundant or filler content, and the key information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is relatively simple with 3 parameters, all well-documented in the schema. The description covers the action and explicitly mentions the return value, which is important since there is no output schema. However, 'full page representation' is a bit vague, and the lack of annotations leaves safety assumptions implicit. Overall, it is reasonably complete for a scroll action.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with each parameter (amount, direction, element_id) already described. The description adds little beyond the schema, only hinting at element_id via 'specific container'. Since the schema does the heavy lifting, a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Scroll') and a resource ('page or a specific container'), which clearly differentiates it from sibling tools like charlotte_navigate or charlotte_observe. It also states the outcome ('Returns full page representation after scrolling'), making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance is given about when to use this tool versus alternatives. The usage is implied by the action itself (scrolling), but there is no mention of exclusions or alternative tools, such as using navigation for moving between pages. This is a minimum viable level of guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_selectA
Select an option in a select/dropdown element. Returns full page representation after selection.
| Name | Required | Description | Default |
|---|---|---|---|
| value | Yes | Value or text of the option to select | |
| element_id | Yes | Target select element ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that the tool returns a full page representation after selection, which is useful, but it does not mention side effects, prerequisites (e.g., element visibility), or event triggering. This is a basic but not comprehensive behavioral description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences that immediately state the core function and the return behavior. It is front-loaded with the action, contains no fluff, and is appropriately sized for the tool's simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema, the description covers the essential points: what it does and what it returns. It could be marginally improved by noting that it is specifically for dropdowns, but the description and schema together provide sufficient context for an agent to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already covers 100% of parameters with clear descriptions ('Value or text of the option to select' and 'Target select element ID'). The tool description adds no additional semantics beyond restating the purpose, so the schema is the primary source of parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Select an option') and the target resource ('select/dropdown element'), making it distinct from sibling tools like charlotte_click or charlotte_type. It also notes the return behavior, further clarifying its function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for interacting with select/dropdown elements, which provides context on when to use it. However, it does not explicitly contrast it with alternatives like charlotte_click or charlotte_toggle, nor does it mention scenarios where it should not be used.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_submitB
Submit a form. Can submit by form ID or by clicking its submit button. Returns full page representation after submission.
| Name | Required | Description | Default |
|---|---|---|---|
| form_id | Yes | Form ID from page representation |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It does state the return value ('Returns full page representation after submission'), which is useful. However, it does not disclose that submitting a form is a mutating action with potential side effects (e.g., data changes, navigation, or irreversible submissions). The description mentions a 'clicking' method without clarifying whether it simulates a user click or requires the button to be visible, which is a behavioral gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise at two sentences, with the primary action front-loaded. The second sentence adds return information without excessive detail. The 'Can submit by form ID or by clicking its submit button' clause is somewhat ambiguous but does not significantly bloat the description. Overall, it is appropriately sized for a simple tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one parameter and no output schema, the description covers the core action and return value, which is fairly complete. However, it omits prerequisites (e.g., needing to be on a page with a form, ensuring the form_id is valid) and does not clarify the 'clicking' method. The mention of an alternative submission method without explaining how to invoke it via the schema reduces completeness. Given the tool's simplicity, a score of 3 is appropriate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers the single parameter 'form_id' with a description ('Form ID from page representation'), giving a baseline of 3. However, the description introduces an alternative submission method ('or by clicking its submit button') that is not represented in the schema, making the parameter semantics confusing. It doesn't add meaningful detail about how form_id is used or obtained, and the alternative method could mislead the agent into expecting an additional parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Submit a form.' This specifies the action (submit) and the resource (form), distinguishing it from sibling tools like charlotte_click or charlotte_type. However, the added 'Can submit by form ID or by clicking its submit button' introduces ambiguity about how submission is performed, detracting from full clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no explicit guidance on when to use this tool versus alternatives like charlotte_click. It mentions two submission methods (by form ID or clicking the submit button) but does not explain when one should be preferred, nor does it contrast with sibling tools that could also perform similar actions. This leaves the agent without clear selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_tab_closeA
Close a browser tab by its ID. If the closed tab was active, switches to the first remaining tab.
| Name | Required | Description | Default |
|---|---|---|---|
| tab_id | Yes | ID of the tab to close |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral disclosure burden. It adds useful context beyond the action itself: if the closed tab was active, it switches to the first remaining tab. This discloses a side effect that an agent would not otherwise know.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no redundant words. It front-loads the core action and then adds one key behavioral detail, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with no output schema, the description is largely complete. It could benefit from mentioning where the tab_id comes from (e.g., from charlotte_tabs), but the sibling context and clear action make this a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% because the only parameter is tab_id with a clear description. The description's 'by its ID' merely restates the schema, adding no additional semantic detail about the parameter's format or source.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the specific verb 'Close' and resource 'browser tab', clearly distinguishing this tool from siblings like charlotte_tab_open and charlotte_tab_switch. The addition 'by its ID' specifies the exact input needed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the usage context obvious: use this when you want to close a browser tab. It doesn't explicitly mention alternatives, but sibling tool names (tab_open, tab_switch) provide enough differentiation to infer when this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_tab_openA
Open a new browser tab. Optionally navigate to a URL. The new tab becomes the active tab.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL to navigate to (default: blank page) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It effectively discloses the key behavioral traits: a new tab is created, optional navigation occurs, and the new tab becomes active. For a simple tool, this is sufficient, though it could mention that the previous tab remains open.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences, with key information front-loaded. Every word earns its place, avoiding unnecessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with no output schema, the description fully covers the purpose and behavior. It is complete enough for an agent to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with the url parameter already described as 'URL to navigate to (default: blank page)'. The description does not add significant new semantics beyond repeating the optionality, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Open') and resource ('browser tab'), clearly stating the action. It also distinguishes itself from sibling tools like charlotte_tab_switch and charlotte_tab_close by focusing on opening a new tab.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: opens a new tab and optionally navigates to a URL, with the new tab becoming active. However, it does not explicitly mention when to use this tool instead of charlotte_navigate (which likely navigates the current tab), leaving a slight gap in alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_tabsA
List all open browser tabs with their URLs, titles, and active status.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It clearly indicates a read-only listing operation and the content of the results. It does not mention pagination or ordering, but for a zero-parameter tool with straightforward behavior, this is adequately transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the action and resource. Every word adds value, with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema tool, the description fully specifies what the tool does and what information is returned. There are no missing details that would prevent an agent from using it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema coverage is 100% (vacuously). Baseline for zero parameters is 4. The description adds no parameter-specific detail because none exists, which is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description uses the specific verb 'List' and clearly identifies the resource as 'all open browser tabs', while also specifying the fields returned (URLs, titles, active status). This distinguishes it from sibling tools such as charlotte_tab_open, charlotte_tab_switch, and charlotte_tab_close.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for obtaining an overview of tabs, but it does not explicitly contrast with alternatives or state when to use this tool versus opening, switching, or closing tabs. Usage context is present but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_tab_switchA
Switch to a different browser tab by its tab ID. Returns the page representation of the activated tab.
| Name | Required | Description | Default |
|---|---|---|---|
| tab_id | Yes | ID of the tab to switch to |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of disclosure. It does reveal an important behavioral aspect: the tool returns the page representation of the activated tab. However, it does not mention what happens if the tab ID is invalid, whether focus changes, or any side effects beyond the switch. This is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences, no filler, and front-loads the action. Every word contributes to understanding the tool's purpose and return value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one parameter and no output schema, the description provides the essential information: what it does and what it returns. Some details like error cases are not covered, but the tool's simplicity makes the description sufficiently complete for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides 100% coverage for the single parameter tab_id with the description 'ID of the tab to switch to'. The tool description adds minimal extra meaning beyond restating 'by its tab ID', so it does not improve on the schema. Baseline 3 is appropriate given high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Switch') and resource ('different browser tab') with the input identifier ('by its tab ID'). It clearly distinguishes itself from sibling tools like charlotte_tab_open and charlotte_tab_close by focusing on switching to an existing tab, and it also specifies the return value.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the basic usage clear (switch to a tab by ID) but does not explicitly state when to prefer this over alternatives like charlotte_tabs or charlotte_tab_open. There are no exclusions or prerequisites mentioned, so guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_toggleA
Toggle a checkbox or switch element. Returns full page representation after toggle.
| Name | Required | Description | Default |
|---|---|---|---|
| element_id | Yes | Target checkbox or switch element ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses the return format ('full page representation') and implies a state change by using 'toggle', but it does not detail behavior in edge cases (e.g., if already checked), potential side effects, or whether it waits for changes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that states the action, the target element type, and the return behavior. It is concise with no redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter browser automation tool, the description adequately covers its purpose and return value. It lacks explicit prerequisites like visibility or interactability, but those are likely implied by the platform context and the simple nature of the action.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the only parameter 'element_id' is described in the schema. The tool description adds no semantic detail beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool's action as toggling a checkbox or switch element and specifies the return value (full page representation). This distinguishes it from siblings like click, select, and submit by focusing on checkbox/switch toggle behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: when the target element is a checkbox or switch. However, it does not explicitly contrast it with the click tool or state exclusions, so it stops short of full guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_toolsA
Manage Charlotte tool visibility. Lists available tool groups and their status. Use 'enable' or 'disable' to control which tools are loaded. Disabled tools don't appear in the tool list — enable a group to access its tools. Groups: 'interaction' for form filling, clicking, and drag-and-drop. 'session' for cookie/auth management, tab switching, viewport, and network. 'dev_mode' for local development serving and audits. 'evaluate' for JavaScript execution. 'monitoring' for console and network request logs. 'dialog' for JavaScript dialog handling.
| Name | Required | Description | Default |
|---|---|---|---|
| group | No | Tool group to enable or disable | |
| action | No | "list" (default) — show all groups and status. "enable"/"disable" — toggle a group. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of disclosing behavior. It explains that disabled tools don't appear and that enabling groups is needed to use their tools. It also defines the scope of each group. It could mention persistence or side effects, but for a visibility toggle this is sufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and appropriately sized. The first sentence states the purpose, the second explains behavior, and the rest enumerates groups efficiently. Every sentence provides useful information with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with two parameters, no output schema, and no annotations, the description thoroughly covers purpose, behavior, and parameter semantics. It gives enough context for an agent to select the right group and action.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with enums, so the baseline is 3. The description adds significant value by describing what each group contains (e.g., 'interaction' for form filling/clicking) and clarifying the default 'list' action, which goes beyond the generic schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb+resource: 'Manage Charlotte tool visibility' and immediately explains the list/enable/disable actions. It is easy to distinguish from sibling tools, which perform specific browser actions like clicking or navigating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use the tool: to list groups/status and to enable or disable groups. It also clarifies the consequence (disabled tools don't appear) and that enabling is required to access tools. It does not explicitly mention alternatives or exclusions, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_typeA
Type text into an input element. Returns full page representation after typing.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | Text to enter | |
| slowly | No | Type one character at a time with a delay between keystrokes. Use for sites with autocomplete, search-as-you-type, or per-key validation (default: false) | |
| element_id | Yes | Target input element ID | |
| clear_first | No | Clear existing value before typing (default: true) | |
| press_enter | No | Press Enter after typing (default: false) | |
| character_delay | No | Milliseconds between keystrokes (implies slowly: true). Default when slowly is true: 50ms. Total typing time is capped at approximately 30s (including per-keystroke overhead); requests whose estimated duration exceeds that are rejected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of disclosing behavioral traits. It states the return type (full page representation) but does not disclose that typing may clear existing content by default, can trigger events, or requires the element to be visible/interactable. The default clearing behavior (clear_first: true) is only discoverable through the schema, not the description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with no wasted words. It front-loads the primary purpose and adds a useful behavioral note about the return representation, fitting the appropriate structure for a simple tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having 6 parameters and no output schema or annotations, the description is minimal but sufficient to understand the core action. The schema covers parameter semantics, and the description states the return type. However, it lacks context about default behaviors (e.g., clearing the field, pressing Enter) that would help an agent anticipate side effects, making it adequate but not comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description itself does not add meaning beyond the schema, but the schema parameters are well-documented with descriptions for each field. Since the description does not need to repeat schema details, a score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Type') and resource ('input element'), and explicitly notes it returns the full page representation after typing. This distinguishes it from sibling tools like charlotte_click, charlotte_select, and charlotte_submit, which perform different actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool should be used when text needs to be entered into an input element, but it does not explicitly discuss when to use this tool versus alternatives, nor does it mention exclusions or prerequisites. No guidance is given for distinguishing between typing and using charlotte_submit or charlotte_click for form interactions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
23 tool updates
v0.8.0- Changed
charlotte_back2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_click2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_click_at2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_diff2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_find4 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / output_fileAdded value: +{ + "description": "Write the full match results to this file path instead of returning them inline. Relative paths resolve against output_dir (see charlotte_configure). Returns only a confirmation with the file path and size. Use for broad selectors (e.g. 'div', '*') that match many elements.", + "type": "string" +} - changed
Input schema / properties / selector / descriptionPrevious value: -"CSS selector to query the DOM directly. Returns elements that may not be in the accessibility tree. Results include Charlotte element IDs for use with interaction tools."New value: +"CSS selector to query the DOM directly. Returns elements that may not be in the accessibility tree. Results include durable Charlotte element IDs (dom-…) that remain valid across subsequent renders and interactions, and work with fill_form; they are re-resolved against the live DOM by re-running the selector."
- Changed
charlotte_forward2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_navigate2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_observe2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_reload2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_screenshot3 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / full_pageAdded value: +{ + "description": "Capture the entire scrollable page (default: true). Set false to capture only the current viewport — much smaller output for long pages. Ignored when 'selector' is provided.", + "type": "boolean" +}
- Changed
charlotte_screenshot_delete2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_screenshot_get2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_screenshots1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
charlotte_scroll2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_select2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_submit2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_tab_close2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_tab_open2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_tab_switch2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_tabs1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
charlotte_toggle2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_tools2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_type7 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - changed
Input schema / properties / character_delay / descriptionPrevious value: -"Milliseconds between keystrokes (implies slowly: true). Default when slowly is true: 50ms"New value: +"Milliseconds between keystrokes (implies slowly: true). Default when slowly is true: 50ms. Total typing time is capped at approximately 30s (including per-keystroke overhead); requests whose estimated duration exceeds that are rejected." - removed
Input schema / properties / press_enter / $refRemoved value: -"#/properties/clear_first" - added
Input schema / properties / press_enter / typeAdded value: +"boolean" - removed
Input schema / properties / slowly / $refRemoved value: -"#/properties/clear_first" - added
Input schema / properties / slowly / typeAdded value: +"boolean"
23 tool updates
v0.6.3- First observed
charlotte_back - First observed
charlotte_click - First observed
charlotte_click_at - First observed
charlotte_diff - First observed
charlotte_find - First observed
charlotte_forward - First observed
charlotte_navigate - First observed
charlotte_observe - First observed
charlotte_reload - First observed
charlotte_screenshot - First observed
charlotte_screenshot_delete - First observed
charlotte_screenshot_get - First observed
charlotte_screenshots - First observed
charlotte_scroll - First observed
charlotte_select - First observed
charlotte_submit - First observed
charlotte_tab_close - First observed
charlotte_tab_open - First observed
charlotte_tab_switch - First observed
charlotte_tabs - First observed
charlotte_toggle - First observed
charlotte_tools - First observed
charlotte_type
TDQS
Scored across 23 tools
Each tool has a clearly distinct purpose: navigation, tab management, interaction, observation, and screenshot management are all separated. Even similar tools like charlotte_click and charlotte_click_at are explicitly differentiated by target type (element vs coordinates).
All tools share the 'charlotte_' prefix and use snake_case, but there is a mix of simple verbs (navigate, click, type) and compound verb_noun forms (tab_open, screenshot_get). This is mostly consistent but not perfectly uniform.
With 23 tools, the server is on the heavier side but still within a manageable range for a comprehensive browser automation tool. The count is justified by the breadth of features, though it approaches the threshold where it might feel overwhelming.
Core browser automation workflows are well covered: navigation, tab management, element interaction, observation, and screenshot handling. However, the description of charlotte_tools mentions groups for dialogs, drag-and-drop, and session management, but these tools are not present in the exposed set, leaving minor gaps.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Scrape, crawl and search the web for AI agents via MCP.
Docs: https://docs.keenable.ai/mcp-server Keenable is a free, remote MCP server that gives agents access to the web index. Search the web with ranked results and date/site filters, then fetch any indexed page as clean markdown. Works out of the box with no account or API key.
Reliable web access for AI agents: smart HTTP, rotating proxies, and full-browser rendering.
- RampifyOAuthdev.rampify
SEO MCP server: crawl your site, find AI-visibility gaps, and ship the fix from your coding agent.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceAn advanced MCP server for browser automation using Puppeteer, specifically optimized for token efficiency through minimal data returns and progressive enhancement. It enables agents to navigate pages, capture LLM-optimized screenshots, extract structured content, and perform batch interactions.3-
- AlicenseNot gradedqualityDmaintenanceMCP server for web scraping and browser automation, enabling AI agents to extract clean, token-efficient content from web pages.1MIT

Wickofficial
FlicenseNot gradedqualityAmaintenanceAn MCP server that provides browser-grade web access for AI agents, using Chrome's actual network stack to bypass anti-bot protections and return clean markdown.8-- AlicenseAqualityDmaintenanceAn MCP server providing AI agents with a stealth Chromium browser that uses hybrid accessibility-object-model and set-of-mark vision for token-lean snapshots and reliable action via ref ids.13601Apache 2.0
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/TickTockBent/charlotte'
If you have feedback or need assistance with the MCP directory API, please join our Discord server