Charlotte
Charlotte
Das Web, lesbar.
Dein KI-Agent verbrennt ~50.000 Zeichen Accessibility-Baum, nur um die Hacker-News-Startseite anzusehen. Charlotte schafft das in 364.
Charlotte ist ein MCP-Server, der KI-Agenten einen strukturierten, token-effizienten Zugriff auf das Web bietet. Statt bei jedem Aufruf den gesamten Accessibility-Baum auszugeben, liefert Charlotte nur das zurück, was der Agent braucht: eine kompakte Seitenzusammenfassung beim Eintreffen, gezielte Abfragen für bestimmte Elemente und volle Details nur auf ausdrückliche Anfrage. Auf inhaltsreichen Seiten ist diese Orientierung bis zu ~140x kleiner als ein vollständiger Accessibility-Baum-Schnappschuss von Playwright MCP; auf trivial kleinen Seiten sind beide ungefähr gleich groß.
Warum Charlotte?
Die meisten Browser-MCP-Server geben bei jedem Aufruf den gesamten Accessibility-Baum aus – ein flacher Textblock, der auf inhaltsreichen Seiten über eine Million Zeichen überschreiten kann. Agenten zahlen dafür, ob sie es brauchen oder nicht.
Charlotte zerlegt jede Seite in eine typisierte, strukturierte Darstellung – Landmarks, Überschriften, interaktive Elemente, Formulare, Inhaltszusammenfassungen – und lässt Agenten steuern, wie viel sie erhalten, mit drei Detailstufen. Wenn ein Agent zu einer neuen Seite navigiert, erhält er eine kompakte Orientierung (364 Zeichen für Hacker News) statt des vollständigen Element-Dumps (~50.000 Zeichen). Wenn er Details braucht, fragt er danach.
Benchmarks
Gemessen mit Charlotte v0.8.0 gegen Playwright MCP v0.0.79, nach Zeichen pro Tool-Aufruf auf echten Websites (npx tsx benchmarks/run-benchmarks.ts --suite comparison), 2026-08-08. Dieser Abschnitt ist eine Zusammenfassung – die kanonische Benchmarks-Seite (einschließlich Kosten pro Aufgabe und Release-Drift) ist charlotte.mintlify.site/benchmarks; Methodik, Instrumente und Rohdaten: benchmarks/.
Orientierungskosten (was ein Agent zahlt, um eine Seite beim Eintreffen zu „sehen“):
Ein Charlotte-navigate liefert standardmäßig eine nutzbare Orientierung – Landmarks, Überschriften und Anzahl interaktiver Elemente, gruppiert nach Seitenregion. Um das Äquivalent mit Playwright MCP zu erhalten, ruft ein Agent browser_snapshot auf, das den vollständigen Accessibility-Baum zurückgibt. (Playwrights browser_navigate allein liefert nur eine kurze Bestätigung, nicht den Seiteninhalt, daher ist es kein direkter Vergleich.)
Site | Charlotte | Playwright | Kleiner um |
example.com | 415 | 465 | 1.1x |
httpbin-Formular | 619 | 1,847 | 3.0x |
GitHub-Repository | 3,778 | 38,983 | 10x |
Wikipedia (KI-Artikel) | 22,134 | 1,137,928 | 51x |
Hacker News | 364 | 50,706 | 139x |
Der Vorteil skaliert mit der Seitenkomplexität: Auf inhaltsreichen Seiten ist die strukturierte Orientierung ~10–140x kleiner als der vollständige Schnappschuss, während auf einer trivial kleinen Seite wie example.com beide innerhalb von ~20% liegen (und auf einer so kleinen Seite kann die strukturierte Darstellung die größere der beiden sein – es gibt schlicht nichts zusammenzufassen). Charlottes Wert zeigt sich genau dort, wo Playwrights flacher Dump am meisten schmerzt. Wenn ein Agent mehr als die Orientierung benötigt, ruft er observe oder find für genau den Teil auf, den er möchte, statt den gesamten Baum im Voraus zu bezahlen.
Overhead der Tool-Definitionen (unsichtbare Kosten pro API-Aufruf):
Profil | Tools | Def. Tokens/Aufruf | Ersparnis vs. voll |
voll | 43 | 8.500 | — |
browse (Standard) | 23 | 4.372 | ~49% |
core | 7 | 2.186 | ~75% |
Tool-Definitionen werden bei jedem API-Round-Trip gesendet. Mit dem Standardprofil browse trägt Charlotte ~49% weniger Definitions-Overhead als beim Laden aller 43 Tools; das minimale core-Profil reduziert ihn um ~75%. Siehe den Profil-Benchmark-Bericht für vollständige Ergebnisse.
Der Workflow-Unterschied: Ein Playwright-Agent, der den vollständigen Schnappschuss liest, erhält ~50.000 Zeichen jedes Mal, wenn er Hacker News ansieht, egal ob er Schlagzeilen liest oder nach einem Login-Button sucht. Ein Charlotte-Agent erhält 364 Zeichen beim Eintreffen, ruft find({ type: "link", text: "login" }) auf, um genau das zu bekommen, was er braucht, und zahlt nie für den Rest.
Related MCP server: krwl3r
So funktioniert es
Charlotte unterhält eine persistente Headless-Chromium-Sitzung und fungiert als Übersetzungsschicht zwischen dem visuellen Web und dem textbasierten Denken des Agenten. Jede Seite wird in eine strukturierte Darstellung zerlegt:
┌─────────────┐ MCP Protocol ┌──────────────────┐
│ AI Agent │<────────────────────>│ Charlotte │
└─────────────┘ │ │
│ ┌────────────┐ │
│ │ Renderer │ │
│ │ Pipeline │ │
│ └─────┬──────┘ │
│ │ │
│ ┌─────▼──────┐ │
│ │ Headless │ │
│ │ Chromium │ │
│ └────────────┘ │
└──────────────────┘Agenten erhalten Landmarks, Überschriften, interaktive Elemente mit typisierten Metadaten, Begrenzungsrahmen, Formularstrukturen und Inhaltszusammenfassungen – alles abgeleitet aus dem, was der Browser bereits über jede Seite weiß.
Funktionen
Navigation — navigate, back, forward, reload
Beobachtung — observe (3 Detailstufen, strukturelle Baumansicht), find (räumliche + semantische Suche, CSS-Selektor-Modus, output_file für große Ergebnismengen), screenshot (mit persistentem Artefakt-Management), screenshots, screenshot_get, screenshot_delete, diff (struktureller Vergleich mit Schnappschüssen)
Interaktion (iframe-bewusst) — click, click_at (koordinatenbasiert), type (mit langsamer Eingabeunterstützung), select, toggle, submit, scroll, hover, drag, key (einzeln/sequenziell mit Element-Ziel), wait_for (asynchrones Bedingungs-Polling), upload (Dateieingabe), fill_form (Stapel-Formularausfüllung), dialog (JS-Dialoge akzeptieren/ablehnen)
Überwachung — console (alle Schweregrade, Filterung, Zeitstempel), requests (vollständiger HTTP-Verlauf, Filterung nach Methode/Status/Ressourcentyp)
Sitzungsverwaltung — tabs, tab_open, tab_switch, tab_close, viewport (generische Voreinstellungen oder benannte Geräte wie „iPhone 15“ mit DPR-, Touch- und User-Agent-Emulation), network (Drosselung, URL-Blockierung), set_cookies, get_cookies, clear_cookies, set_headers, configure
Entwicklungsmodus — dev_serve (statischer Server + Dateiüberwachung mit automatischem Neuladen), dev_inject (CSS/JS-Injektion), dev_audit (a11y, Leistung, SEO, Kontrast, defekte Links)
Dienstprogramme — evaluate (beliebige JS-Ausführung im Seitenkontext)
Tool-Profile
Charlotte bringt 43 Tools mit (42 registrierte + das Meta-Tool charlotte_tools), aber die meisten Workflows benötigen nur eine Teilmenge. Startprofile steuern, welche Tools in den Kontext des Agenten geladen werden, und reduzieren den Definitions-Overhead um bis zu ~75%.
charlotte --profile browse # 23 tools (default) — navigate, observe, interact, tabs
charlotte --profile core # 7 tools — navigate, observe, find, click, type, submit
charlotte --profile full # 43 tools — everything
charlotte --profile interact # 31 tools — full interaction + dialog + evaluate
charlotte --profile develop # 34 tools — interact + dev_serve, dev_inject, dev_audit
charlotte --profile audit # 14 tools — navigation + observation + dev_audit + viewportAgenten können während der Sitzung weitere Tools aktivieren, ohne neu zu starten:
charlotte_tools enable dev_mode → activates dev_serve, dev_audit, dev_inject
charlotte_tools disable dev_mode → deactivates them
charlotte_tools list → see what's loadedSelbst-Hosting (Charlotte Remote)
Führe Charlotte als Remote-MCP-Server aus und verbinde ihn mit claude.ai – ein Befehl:
docker run --cap-add SYS_ADMIN --shm-size 2g -p 3737:3737 ghcr.io/ticktockbent/charlotteEr gibt eine öffentliche Connector-URL und ein Operator-Token aus. In claude.ai: Einstellungen → Connectors → Benutzerdefinierten Connector hinzufügen – füge die URL ein, lasse die OAuth-Client-ID/Secret-Felder leer und gib das Token auf Charlottes Zustimmungsseite ein, wenn sie erscheint. Das war's; du surfst.
Die Demo-URL und das Token sind ephemer (beide rotieren beim Neustart). Für den echten Betrieb – stabile Domain, eigener Tunnel oder Reverse-Proxy, Docker Compose: Selbst-Hosting. Vertrauensmodell und Netzwerk-Schutz: Sicherheit. Container- und Sandbox-Interna: Docker.
Schnellstart
Voraussetzungen
Node.js >= 20
npm
Installation
Charlotte ist im MCP-Registry als io.github.TickTockBent/charlotte gelistet und auf npm als @ticktockbent/charlotte veröffentlicht:
npm install -g @ticktockbent/charlotteDocker-Images sind auf Docker Hub und GitHub Container Registry verfügbar:
# Alpine (default, smaller)
docker pull ticktockbent/charlotte:alpine
# Debian (if you need glibc compatibility)
docker pull ticktockbent/charlotte:debian
# Or from GHCR
docker pull ghcr.io/ticktockbent/charlotte:latestOder aus dem Quellcode installieren:
git clone https://github.com/ticktockbent/charlotte.git
cd charlotte
npm install
npm run buildAusführen
Charlotte kommuniziert über stdio mit dem MCP-Protokoll:
# If installed globally (default browse profile)
charlotte
# With a specific profile
charlotte --profile core
# If installed from source
npm startMCP-Client-Konfiguration
Claude Code
Erstelle .mcp.json im Projektstamm:
{
"mcpServers": {
"charlotte": {
"type": "stdio",
"command": "npx",
"args": ["@ticktockbent/charlotte"],
"env": {}
}
}
}Claude Desktop
Füge zu claude_desktop_config.json hinzu:
{
"mcpServers": {
"charlotte": {
"command": "npx",
"args": ["@ticktockbent/charlotte"]
}
}
}Cursor
Füge zu .cursor/mcp.json hinzu:
{
"mcpServers": {
"charlotte": {
"command": "npx",
"args": ["@ticktockbent/charlotte"]
}
}
}Windsurf
Füge zu ~/.codeium/windsurf/mcp_config.json hinzu:
{
"mcpServers": {
"charlotte": {
"command": "npx",
"args": ["@ticktockbent/charlotte"]
}
}
}VS Code (Copilot)
Füge zu .vscode/mcp.json hinzu:
{
"servers": {
"charlotte": {
"type": "stdio",
"command": "npx",
"args": ["@ticktockbent/charlotte"]
}
}
}Cline
Füge zu den Cline-MCP-Einstellungen hinzu (über die Cline-Seitenleiste > MCP-Server > Konfigurieren):
{
"mcpServers": {
"charlotte": {
"command": "npx",
"args": ["@ticktockbent/charlotte"]
}
}
}Amp
Füge zu ~/.amp/settings.json hinzu:
{
"mcpServers": {
"charlotte": {
"command": "npx",
"args": ["@ticktockbent/charlotte"]
}
}
}Siehe docs-internal/mcp-setup.md für die vollständige Einrichtungsanleitung, einschließlich Entwicklungsmodus, generischer MCP-Clients, Verifizierungsschritte und Fehlerbehebung.
Konfiguration
Charlotte löst Einstellungen aus vier Quellen auf, mit höchster Priorität zuerst: CLI-Argumente → Umgebungsvariablen → Konfigurationsdatei → eingebaute Standardwerte. Siehe docs/configuration.md für die vollständige Referenz.
Konfigurationsdatei
Übergib eine JSON-Konfigurationsdatei mit --config oder lege eine charlotte.config.json im Arbeitsverzeichnis ab, und Charlotte lädt sie automatisch:
charlotte --config charlotte.config.json{
"browser": { "headless": true, "noSandbox": false },
"tools": { "profile": "browse" },
"rendering": { "includeIframes": false, "iframeDepth": 3 },
"output": { "dir": "./charlotte-output" },
"limits": {
"maxInteractiveElements": 2000,
"maxFullContentChars": 200000,
"maxResponseBytes": 1000000,
"maxEvaluateBytes": 256000
}
}Jeder Abschnitt ist optional; ein leeres {} ist gültig. Die Datei wird mit zod validiert – unbekannte Schlüssel, falsche Typen oder ungültige Enum-Werte erzeugen eine klare Startfehlermeldung auf stderr, und Charlotte beendet sich mit einem Nicht-Null-Exit-Code. Vier Einstellungen haben auch Umgebungsvariablen: CHARLOTTE_NO_SANDBOX, CHARLOTTE_OUTPUT_DIR, CHARLOTTE_CDP_ENDPOINT und CHARLOTTE_INIT_SCRIPT. Skripte, die bei jedem neuen Dokument vor dem Seiten-JS ausgeführt werden sollen, kommen in browser.initScripts oder --init-script <Pfad> (wiederholbar); siehe Init-Skripte.
Die Chromium-Sandbox ist standardmäßig aktiviert
Verhaltensänderung in v0.7.0: Frühere Versionen haben
--no-sandboxin jeden Chromium-Start eingebaut. Seit v0.7.0 ist die Chromium-Sandbox standardmäßig aktiviert – die primäre Verteidigung zwischen einer nicht vertrauenswürdigen Seite und dem Konto, unter dem Charlotte läuft. Du musst explizit abwählen, wo die Kernel-Sandbox nicht verfügbar ist.
charlotte --no-sandbox # CLI flag
CHARLOTTE_NO_SANDBOX=1 charlotte # environment variable
# or "browser": { "noSandbox": true } in the config fileMigrationshinweis (Docker / Bare-Metal): Container können die Kernel-Sandbox normalerweise nicht einrichten, daher setzen die bereitgestellten Dockerfiles CHARLOTTE_NO_SANDBOX=1 für dich, und docker-compose.yml behält jetzt den Standard-seccomp-Filter von Docker bei (es läuft nicht mehr mit seccomp=unconfined). Wenn du Charlotte bare-metal als root ausführst, weigert sich Chromium, mit aktivierter Sandbox zu starten – führe es als Nicht-Root-Benutzer aus (empfohlen) oder übergib --no-sandbox. Bestehende Setups, die zuvor auf das implizite --no-sandbox angewiesen waren und in einer Umgebung laufen, in der die Sandbox nicht initialisiert werden kann, müssen jetzt CHARLOTTE_NO_SANDBOX=1 (oder das Flag/Konfigurationsäquivalent) setzen, um weiter zu funktionieren.
Das Ausführen von Charlotte Remote (HTTP-Modus) über das Netzwerk wirft zusätzliche Fragen zu Vertrauensgrenzen und Netzwerk-Schutz auf, die über die Sandbox hinausgehen – siehe Sicherheit.
Ausgabegrößen-Limits
Die limits.*-Schlüssel begrenzen, wie viel eine einzelne Tool-Antwort zurückgeben kann, damit eine pathologische Seite (100k Links, ein Endlos-Scroll-Feed, ein riesiger Dokumentkörper) den Kontextfenster des Agenten nicht überlaufen lässt. Wenn eine Seitenantwort maxResponseBytes überschreitet, wird sie zu einer kompakten Zusammenfassung degradiert und schlägt vor, das vollständige Ergebnis über output_file auf die Festplatte zu schreiben; charlotte_evaluate-Ergebnisse werden unabhängig durch maxEvaluateBytes begrenzt. Abgeschnittene Antworten tragen einen truncation-Marker. Siehe docs/configuration.md für die Schlüssel und Standardwerte.
Absturz-Wiederherstellung
Ein Chromium-Absturz blockiert den Server nicht mehr. Der nächste Tool-Aufruf startet den Browser automatisch neu, leert die Caches für tote Tabs und CDP-Sitzungen und öffnet einen neuen leeren Tab – so kann ein Agent nach einem Renderer-Absturz weiterarbeiten, ohne den MCP-Server neu zu starten.
Verwendungsbeispiele
Sobald die Verbindung steht, kann ein Agent Charlottes Tools verwenden:
Eine Website durchsuchen
navigate({ url: "https://example.com" })
// → 612 chars: landmarks, headings, interactive element counts
find({ type: "link", text: "More information" })
// → just the matching element with its ID
click({ element_id: "lnk-a3f1c2" })Ein Formular ausfüllen
navigate({ url: "https://httpbin.org/forms/post" })
find({ type: "text_input" })
type({ element_id: "inp-c7e29b", text: "hello@example.com" })
select({ element_id: "sel-e8a3f5", value: "option-2" })
submit({ form_id: "frm-b1d4e7" })Lokale Entwicklungs-Feedbackschleife
dev_serve({ path: "./my-site", watch: true })
observe({ detail: "full" })
dev_audit({ checks: ["a11y", "contrast"] })
dev_inject({ css: "body { font-size: 18px; }" })Seitenrepräsentation
Charlotte liefert strukturierte Repräsentationen mit drei Detailstufen, die es Agenten ermöglichen, zu steuern, wie viel Kontext sie verbrauchen:
Minimal (Standard für navigate)
Landmarks, Überschriften und Zählungen interaktiver Elemente, gruppiert nach Seitenregion. Entwickelt für die Orientierung – „Was ist auf dieser Seite?" – ohne jedes Element aufzulisten.
{
"url": "https://news.ycombinator.com",
"title": "Hacker News",
"viewport": { "width": 1280, "height": 720 },
"structure": {
"headings": [{ "level": 1, "text": "Hacker News", "id": "hdg-a1b2c3" }]
},
"interactive_summary": {
"total": 93,
"by_landmark": {
"(page root)": { "link": 91, "text_input": 1, "button": 1 }
}
}
}Summary (Standard für observe)
Vollständige Liste interaktiver Elemente mit typisierten Metadaten, Formularstrukturen und Inhaltszusammenfassungen.
{
"url": "https://example.com/dashboard",
"title": "Dashboard",
"viewport": { "width": 1280, "height": 720 },
"structure": {
"landmarks": [
{ "id": "rgn-b2c1d0", "role": "banner", "label": "Site header", "bounds": { "x": 0, "y": 0, "w": 1280, "h": 64 } },
{ "id": "rgn-d4e5f6", "role": "main", "label": "Content", "bounds": { "x": 240, "y": 64, "w": 1040, "h": 656 } }
],
"headings": [{ "level": 1, "text": "Dashboard", "id": "hdg-1a2b3c" }],
"content_summary": "main: 2 headings, 5 links, 1 form"
},
"interactive": [
{
"id": "btn-a3f1c2",
"type": "button",
"label": "Create Project",
"bounds": { "x": 960, "y": 80, "w": 160, "h": 40 },
"state": {}
}
],
"forms": []
}Full
Alles aus summary, plus den gesamten sichtbaren Textinhalt der Seite.
Detailstufen
Level | Tokens | Verwendungszweck |
| ~50-200 | Orientierung nach der Navigation. Welche Regionen gibt es? Wie viele interaktive Elemente? |
| ~500-5000 | Arbeiten mit der Seite. Vollständige Elementliste, Formularstrukturen, Inhaltszusammenfassungen. |
| variabel | Lesen des Seiteninhalts. Der gesamte sichtbare Text ist enthalten. |
Navigationstools verwenden standardmäßig minimal. Das observe-Tool verwendet standardmäßig summary. Beide akzeptieren einen optionalen detail-Parameter zum Überschreiben.
Element-IDs
Element-IDs sind über kleinere DOM-Mutationen hinweg stabil. Sie werden durch Hashing eines zusammengesetzten Schlüssels aus Elementtyp, ARIA-Rolle, zugänglichem Namen und DOM-Pfadsignatur erzeugt:
btn-a3f1c2 (button) inp-c7e29b (text input)
lnk-d4b910 (link) sel-e8a3f5 (select)
chk-f1a204 (checkbox) frm-b1d4e7 (form)
rgn-e0d2a8 (landmark) hdg-0f4063 (heading)
dom-b2c3d9 (DOM element, from CSS selector queries)v0.7.0-ID-Formatänderung: Element-ID-Hashes sind jetzt 6 Hexadezimalzeichen (z. B.
btn-a3f1c2), gegenüber 4 in früheren Versionen. Dies reduziert Hash-Kollisionen zwischen Elementen auf großen Seiten drastisch. Agenten, die 4-stellige IDs fest codiert oder per Muster abgeglichen haben, sollten Elemente nach dem Upgrade erneut perfindsuchen, anstatt zwischengespeicherte IDs wiederzuverwenden.
IDs überstehen unabhängige DOM-Änderungen und Neuanordnungen von Elementen innerhalb desselben Containers. Wenn ein Agent mit minimaler Detailstufe navigiert (ohne einzelne Element-IDs), verwendet er find, um Elemente anhand von Text, Typ oder räumlicher Nähe zu lokalisieren – die zurückgegebenen Elemente enthalten IDs, die für die Interaktion bereit sind.
Entwicklung
# Run in watch mode
npm run dev
# Run all tests
npm test
# Run only unit tests
npm run test:unit
# Run only integration tests
npm run test:integration
# Type check
npx tsc --noEmitProjektstruktur
src/
browser/ # Puppeteer lifecycle, tab management, CDP sessions
renderer/ # Accessibility tree extraction, layout, content, element IDs
state/ # Snapshot store, structural differ
tools/ # MCP tool definitions (navigation, observation, interaction, session, dev-mode)
dev/ # Static server, file watcher, auditor
types/ # TypeScript interfaces
utils/ # Logger, hash, wait utilities
tests/
unit/ # Fast tests with mocks
integration/ # Full Puppeteer tests against fixture HTML
fixtures/pages/ # Test HTML filesArchitektur
Die Renderer-Pipeline ist der Kern – sie ruft Extraktoren in Reihenfolge auf und setzt eine PageRepresentation zusammen:
Extraktion des Accessibility-Baums (CDP
Accessibility.getFullAXTree)Layout-Extraktion (CDP
DOM.getBoxModel)Extraktion von Landmarks, Überschriften, interaktiven Elementen und Inhalten
Generierung von Element-IDs (hashbasiert, stabil über Neu-Renderings hinweg)
Alle Tools laufen über renderActivePage(), das Snapshots, Reload-Ereignisse, Dialogerkennung und Antwortformatierung übernimmt.
Sandbox
Charlotte enthält eine Testwebsite in tests/sandbox/, die alle Tools ohne Zugriff auf das öffentliche Internet testet. Lokal bereitstellen mit:
dev_serve({ path: "tests/sandbox" })Fünf Seiten decken Navigation, Formulare, interaktive Elemente, Popups, verzögerten Inhalt, Scroll-Container und mehr ab. Siehe docs-internal/sandbox.md für die vollständige Seitenreferenz und eine toolweise Übungscheckliste.
Bekannte Probleme
Shadow DOM – Offenes Shadow DOM funktioniert transparent. Der Accessibility-Baum von Chromium durchdringt offene Shadow-Grenzen, sodass Webkomponenten (z. B. GitHub's <relative-time>, <tool-tip>) ihren Inhalt ohne spezielle Behandlung in Charlottes Repräsentation rendern. Geschlossene Shadow-Roots sind für den Accessibility-Baum undurchsichtig und werden nicht erfasst.
Roadmap
Sitzung & Konfiguration
Feature-Roadmap
Videoaufzeichnung – Interaktionen als Video aufzeichnen, um die vollständige Sequenz agentengesteuerter Navigation und Manipulation für Debugging, Dokumentation und Überprüfung zu erfassen.
Siehe docs-internal/playwright-mcp-gap-analysis.md für die vollständige Gap-Analyse gegenüber Playwright MCP, einschließlich Prioritäten mit niedrigerer Priorität (Vision-Tools, Testen/Verifizieren, Tracing, Transport, Sicherheit) und Bereichen, in denen Charlotte Vorteile hat.
Vollständige Spezifikation
Siehe docs-internal/CHARLOTTE_SPEC.md für die vollständige Spezifikation einschließlich aller Tool-Parameter, des Seitenrepräsentationsformats, der Elementidentitätsstrategie und der Architekturdetails.
Lizenz
Community
Öffnen Sie einen Fehlerbericht für reproduzierbare Fehler, Regressionen oder MCP-Client-spezifische Probleme.
Öffnen Sie eine Funktionsanfrage für Workflow-Verbesserungen oder neue Fähigkeiten.
Öffnen Sie eine Tool-Anfrage, wenn Sie ein neues Tool, eine Parameteroberfläche oder eine Profilplatzierung vorschlagen möchten.
Durchsuchen Sie offene Issues, um aktuelle Arbeiten und Diskussionen zu finden.
Prüfen Sie den geplanten Good-First-Issue-Filter, da Maintainer einsteigerfreundliche Aufgaben markieren.
Mitwirken
Siehe CONTRIBUTING.md für Richtlinien.
Teil einer wachsenden Suite literarisch benannter MCP-Server. Mehr unter github.com/TickTockBent.
Available Tools
23 toolscharlotte_backA
Navigate back in browser history. Returns page representation after navigation.
| Name | Required | Description | Default |
|---|---|---|---|
| detail | No | "minimal" (default), "summary" (includes content context), "full" (includes all text content) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the main behavior (navigating back) and the return value (page representation), which provides some context. However, it omits important edge-case details such as behavior when there is no history, whether it waits for page load, and the exact nature of the 'page representation,' leaving gaps in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences, front-loading the action and then adding the return type. There is no filler, redundancy, or unnecessary detail; every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with one well-documented parameter and no annotations or output schema. The description covers the core purpose and return, which is largely sufficient. It could be improved by mentioning how the 'detail' parameter affects the returned representation or failure behavior with an empty history, but overall it is complete for a simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides 100% coverage for the single optional 'detail' parameter, including its enum and description. The tool description adds no additional parameter information, but the baseline of 3 applies given the high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Navigate back') and resource ('browser history'), clearly distinguishing this from sibling tools like charlotte_forward and charlotte_navigate. It also states the return value ('page representation after navigation'), leaving no ambiguity about its function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by stating the action ('navigate back'), but it does not explicitly mention when to prefer this over alternatives (e.g., charlotte_forward, charlotte_navigate) or provide exclusions. There is no guidance on scenarios like empty history, but the action itself is a clear directive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_clickA
Click an interactive element on the page. Returns full page representation after the click.
| Name | Required | Description | Default |
|---|---|---|---|
| modifiers | No | Modifier keys to hold during click: ["ctrl"], ["shift"], ["alt"], ["meta"], or combinations like ["ctrl", "shift"] | |
| click_type | No | Click type: "left" (default), "right", "double" | |
| element_id | Yes | Target element ID from page representation |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full transparency burden. It does disclose that the tool returns a full page representation after the click, which is useful. However, it does not mention potential side effects like navigation, form submission, or waiting behavior, which are important for a click action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, concise and front-loaded with the action and return value. Every word earns its place, with no fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a basic click tool, covering the action and return value. However, it lacks any mention of the sibling tool charlotte_click_at, does not explain what 'full page representation' entails, and omits potential side-effect warnings. More context is needed for full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides descriptions for all three parameters (element_id, click_type, modifiers) with 100% coverage. The description itself adds no extra parameter context, so it relies on the schema, matching the baseline score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool clicks an interactive element on the page, which distinguishes it from scrolling, typing, and navigation. However, it does not explicitly differentiate from charlotte_click_at, which is a sibling tool likely for coordinate-based clicks, leaving the distinction implied through the word 'element'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'interactive element' implies element-based clicking, which suggests using this tool over charlotte_click_at, but it does not explicitly state when to use this tool versus alternatives or provide exclusions. The guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_click_atA
Click at specific page coordinates. Use when target elements are not in the accessibility tree (custom widgets, canvas, non-semantic interactive divs). Dispatches real CDP-level mouse events. Returns full page representation after the click.
| Name | Required | Description | Default |
|---|---|---|---|
| x | Yes | X coordinate in page pixels | |
| y | Yes | Y coordinate in page pixels | |
| modifiers | No | Modifier keys to hold during click: ["ctrl"], ["shift"], ["alt"], ["meta"], or combinations | |
| click_type | No | Click type: "left" (default), "right", "double" |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full responsibility for behavioral disclosure. It reveals that the click dispatches 'real CDP-level mouse events' and returns a 'full page representation' after the click. While it doesn't cover every edge case (e.g., out-of-viewport coordinates, waiting for navigation), it provides the core behavioral traits relevant to a coordinate click tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: action statement, usage context, and return behavior. No filler or redundancy. The structure is front-loaded with the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (4 params, no output schema), and the description covers the return value ('full page representation') and the type of events dispatched. It doesn't mention prerequisites like page load state, but such details are likely unnecessary for a click tool with clear semantics. Overall, it is sufficiently complete for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents all four parameters with clear descriptions, including enum values for modifiers and click_type (100% coverage). The description adds no additional parameter semantics beyond labeling x/y as 'page coordinates', so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Click at specific page coordinates' with a specific verb and resource. It also distinguishes itself from the sibling element-based click tool by noting it's for when elements are not in the accessibility tree (custom widgets, canvas, non-semantic interactive divs).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly provides when-to-use guidance: 'Use when target elements are not in the accessibility tree...' This implies not to use it when elements are accessible, effectively differentiating it from alternatives like charlotte_click. The mention of CDP-level events preempts expectations about event orchestration.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_diffA
Compare current page state to a previous snapshot. Returns structural diff showing added, removed, moved, and changed elements.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | No | "all" (default), "structure" (landmarks/headings), "interactive" (elements/forms), "content" (text/url/title) | |
| snapshot_id | No | Compare against a specific snapshot ID (default: previous snapshot) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the transparency burden. It discloses the output type (structural diff with added/removed/moved/changed elements) and implies a read-only comparison, but it doesn't explicitly state that it is non-destructive or mention prerequisites like the existence of a previous snapshot. Some behavioral context is added, but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is composed of two short sentences that front-load the primary action and the return value. There is no redundant phrasing or filler, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two optional, well-documented parameters and no output schema, the description provides the core purpose and a high-level summary of return categories. It does not explain snapshot prerequisites or error behavior, but the combination of description and schema is sufficient for the tool's apparent simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers both parameters 100%, with scope enum descriptions and snapshot_id details. The tool description itself adds no additional parameter semantics beyond what the schema already provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Compare' with a clear resource ('current page state to a previous snapshot') and states the return type ('structural diff showing added, removed, moved, and changed elements'). This distinguishes it from sibling tools that operate on the live page (e.g., click, type, observe) and makes the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear context: use this tool when you need to see differences between the current page state and a previous snapshot. It doesn't explicitly list exclusions or alternatives, but the diff focus is obvious among the siblings, which are mostly navigation and interaction tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_findA
Search for elements matching criteria. Filters interactive elements by text, role, type, or spatial proximity. Use the selector parameter to find DOM elements by CSS selector — this reaches elements not in the accessibility tree (custom widgets, non-semantic divs). Selector results return Charlotte element IDs usable with click, hover, drag, etc.
| Name | Required | Description | Default |
|---|---|---|---|
| near | No | Element ID — find elements spatially near this one (within ~200px) | |
| role | No | ARIA role filter | |
| text | No | Text content to search for (case-insensitive substring match) | |
| type | No | Interactive element type filter (button, link, text_input, select, checkbox, etc.) | |
| within | No | Element ID — find elements geometrically contained within this one's bounds | |
| selector | No | CSS selector to query the DOM directly. Returns elements that may not be in the accessibility tree. Results include durable Charlotte element IDs (dom-…) that remain valid across subsequent renders and interactions, and work with fill_form; they are re-resolved against the live DOM by re-running the selector. | |
| output_file | No | Write the full match results to this file path instead of returning them inline. Relative paths resolve against output_dir (see charlotte_configure). Returns only a confirmation with the file path and size. Use for broad selectors (e.g. 'div', '*') that match many elements. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden and discloses important behaviors: selector reaches non-accessibility-tree elements, results return durable Charlotte element IDs re-resolved against the live DOM, and output_file writes to a file with a confirmation. It omits details on behavior when no filters are provided or whether hidden elements are included, leaving minor gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: lead purpose statement, filter list, selector explanation, then output_file behavior. Every clause adds functional detail without redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main search use case, the unique selector capability, and output_file output, but does not specify the return format for inline results or behavior when no criteria are supplied. This is a minor gap given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds meaning beyond the schema for selector (durable IDs, re-resolution, works with fill_form) and output_file (confirmation with path and size, use for broad selectors), justifying a score above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Search for elements matching criteria' with specific filters (text, role, type, spatial proximity) and highlights the selector parameter for DOM access. This distinguishes it from sibling action tools like click, type, and navigate by focusing on search/find.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly recommends using the selector parameter to find elements not in the accessibility tree, providing clear contextual guidance. However, it does not explicitly state when not to use the tool or compare it to alternatives, so it stops short of full exclusionary usage advice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_forwardA
Navigate forward in browser history. Returns page representation after navigation.
| Name | Required | Description | Default |
|---|---|---|---|
| detail | No | "minimal" (default), "summary" (includes content context), "full" (includes all text content) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the return behavior ('Returns page representation after navigation'), but does not specify edge cases like what happens when there is no forward history, or how 'page representation' is structured. This is adequate but minimal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with no wasted words. It front-loads the primary action and then describes the return value, making it easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (one optional parameter, no output schema), and the description covers the essential purpose and return behavior. It could mention edge cases, but given the low complexity, it is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single optional 'detail' parameter, with enum values and descriptions. The tool description does not add parameter information, but the schema fully documents it, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Navigate forward in browser history', using a specific verb and resource. It distinguishes itself from sibling tools like charlotte_back (backward navigation) and charlotte_navigate (URL navigation).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use this tool (when the user wants to go forward in browser history). It doesn't explicitly mention alternatives or exclusions, but the context is clear enough given the sibling tool names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_observeA
Get current page state without performing any action. Use detail levels to control verbosity: "minimal" for landmarks, headings, and interactive element counts by landmark (use charlotte_find to get specific elements with actionable IDs, or observe({ detail: "summary" }) to see all elements), "summary" (default) for content summaries and full element list, "full" for all text content. Use view: "tree" for a compact structural outline (cheapest orientation tool), or view: "tree-labeled" to include labels on interactive elements (still much cheaper than minimal JSON, and shows which button/link/input is which).
| Name | Required | Description | Default |
|---|---|---|---|
| view | No | "default" (structured JSON), "tree" (compact structural outline — element types only, cheapest), or "tree-labeled" (structural outline with interactive element labels — shows which button/link/input is which, still ~70% cheaper than minimal JSON) | |
| detail | No | "summary" (default), "full" (includes all text content), "minimal" (landmarks + interactive only) | |
| selector | No | CSS selector to scope observation to a subtree | |
| output_file | No | Write observation data to this file path instead of returning inline. Relative paths resolve against output_dir (see charlotte_configure). Returns only a confirmation with the file path and size. | |
| include_styles | No | Include computed styles for visible elements (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It states 'without performing any action' (non-mutating), explains cost trade-offs ('cheapest', 'still much cheaper than minimal JSON'), and reveals output behavior for output_file ('Returns only a confirmation with the file path and size'). This provides meaningful context beyond the schema, though it omits potential error conditions or page-load requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured: it leads with the main purpose, then explains optional parameters in a logical flow. Each sentence adds operational value without filler. Slightly longer than ideal, but the complexity of 5 parameters and absence of annotations justify the length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (5 params, no output schema, no annotations), the description covers the essential aspects: purpose, parameter behaviors, alternatives, and cost considerations. It lacks explicit return format details for the default view, but for a read-only observation tool, the provided information is sufficient for an agent to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds significant value by explaining detail level semantics (minimal/summary/full), view trade-offs (tree vs tree-labeled with cost estimates), and output_file resolution ('relative paths resolve against output_dir'). This goes beyond the bare enum names and provides actionable guidance for parameter selection.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Get current page state without performing any action.' The verb 'get' and resource 'page state' precisely convey the read-only nature, distinguishing it from action-oriented siblings like charlotte_click and charlotte_type. Mentioning charlotte_find and observe variants further differentiates it from element-finding tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly directs when to use alternatives: 'use charlotte_find to get specific elements with actionable IDs' and 'or observe({ detail: "summary" }) to see all elements.' Recommends view: 'tree' as the 'cheapest orientation tool,' giving clear context on when to choose this tool over others.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_reloadA
Reload the current page. Returns page representation after reload.
| Name | Required | Description | Default |
|---|---|---|---|
| hard | No | Bypass cache (default: false) | |
| detail | No | "minimal" (default), "summary" (includes content context), "full" (includes all text content) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It does state that the tool returns a page representation after reload, which is valuable. However, it does not mention potential side effects such as losing unsaved form state, and the 'hard' parameter's cache-bypass behavior is only in the schema, not the description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One concise sentence that immediately states the action and outcome. Every word earns its place with no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple reload action with two fully documented optional parameters and no output schema, the description provides adequate context: it says what the tool does and what it returns. Slightly more detail on how 'detail' affects the returned representation could improve it, but the schema already covers that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: both 'hard' and 'detail' already have meaningful descriptions in the input schema. The tool description adds no additional parameter semantics, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Reload') with a clear resource ('current page') and explicitly states the return value ('Returns page representation after reload'). It is distinct from sibling navigation tools like charlotte_navigate, charlotte_back, and charlotte_forward.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage context is implied: reload the current page when a refresh is needed. However, there are no explicit guidelines on when to choose this over alternatives like navigate or back/forward, nor any exclusionary conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_screenshotA
Capture a visual screenshot. Fallback for when structured representation isn't sufficient (complex visualizations, canvas elements, images). Use save: true to persist as a file artifact that can be referenced later.
| Name | Required | Description | Default |
|---|---|---|---|
| save | No | Save as a persistent file artifact (default: false). When true, the screenshot is written to disk and artifact metadata is returned alongside the image. | |
| format | No | "png" (default), "jpeg", "webp" | |
| quality | No | 1-100 for jpeg/webp quality | |
| selector | No | CSS selector to capture specific element (default: full page) | |
| full_page | No | Capture the entire scrollable page (default: true). Set false to capture only the current viewport — much smaller output for long pages. Ignored when 'selector' is provided. | |
| output_file | No | Write screenshot to this file path instead of returning base64 inline. Relative paths resolve against output_dir (see charlotte_configure). Returns only a confirmation with the file path and size. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full transparency burden. It adds useful context about the persistent behavior of save:true and the tool's fallback role. However, it does not disclose the default return format (e.g., inline base64), whether the operation has side effects beyond saving, or that it is a read-only action. These are notable omissions for an unannotated tool, but the provided behavior hints (persistence, fallback) prevent a lower score.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long, front-loaded with the primary verb behavior ('Capture a visual screenshot'), then a use-case clause, then a targeted parameter tip. Every sentence earns its place; no filler or redundant restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 6 parameters, no output schema, and no annotations, the description does more than a minimal effort by establishing purpose and usage guidance. However, it leaves important contextual gaps: the default return format (base64) is not mentioned, and there is no hint about how to later retrieve saved artifacts (e.g., charlotte_screenshot_get). This is adequate but not fully complete for an agent making an uninformed invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all six parameters with meaningful descriptions. The description adds a bit of extra context for 'save' (persisting as a file artifact) but largely repeats what the schema explains. It neither contradicts nor significantly enhances the schema's parameter documentation, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource phrase ('Capture a visual screenshot'), clearly stating the tool's core function. It further distinguishes the tool from siblings by framing it as a 'fallback for when structured representation isn't sufficient,' referencing complex visualizations, canvas elements, and images—which implicitly contrasts with observation-oriented siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use context: 'Fallback for when structured representation isn't sufficient' and gives concrete example cases (complex visualizations, canvas elements, images). It also offers conditional parameter guidance ('Use save: true to persist as a file artifact'). It stops short of naming a specific alternative sibling tool or outright saying 'do not use for simple text/structured content,' so it earns a 4 rather than 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_screenshot_deleteA
Delete a saved screenshot artifact by its ID. Removes the file from disk.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Screenshot artifact ID to delete |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of disclosing side effects. It explicitly states 'Removes the file from disk,' which conveys the destructive, persistent nature of the operation. It does not detail error handling or permissions, but for a simple delete operation, this is substantial disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences, front-loads the core action, and contains no filler or repetition. Every word contributes meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple deletion tool with one parameter and no output schema, the description adequately covers the action and its effect. It does not mention how to obtain the ID or that the deletion is permanent, but the essentials are present, and the simplicity of the tool reduces the need for further context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already has 100% coverage, documenting the single 'id' parameter as 'Screenshot artifact ID to delete.' The description adds little beyond restating 'by its ID,' so the schema carries the semantic weight. This matches the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Delete') and clearly identifies the resource ('saved screenshot artifact') and the required identifier ('by its ID'). It is unambiguous and distinguishes this tool from sibling tools like charlotte_screenshot_get, which retrieves rather than deletes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool—when a saved screenshot artifact needs to be deleted—but gives no explicit guidance about when not to use it or alternatives. It does not mention that this is the only tool for deletion or that retrieval tools should be used for other purposes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_screenshot_getA
Retrieve a previously saved screenshot artifact by its ID. Returns the image data and metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Screenshot artifact ID (e.g. ss-20260224103000-a1b2c3) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. The verb 'Retrieve' clearly implies a non-destructive read-only operation, and 'Returns the image data and metadata' sets expectations for the output. No side effects or special requirements are mentioned, but for a simple retrieval tool this is sufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that states the action, the target resource, and the return value without any redundant wording or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has only one parameter and no output schema, so the description needs to explain the outcome, which it does ('Returns the image data and metadata'). For a simple getter, this is fully complete and requires no additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers the only parameter 'id' with a clear description and format example. The tool description adds only 'by its ID', which reinforces the schema without adding new semantic meaning. Given 100% schema coverage, the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Retrieve') and resource ('saved screenshot artifact by its ID'), making the operation clear. It also distinguishes from sibling tools like charlotte_screenshot (which likely captures a new screenshot) and charlotte_screenshot_delete (which deletes) by focusing on retrieval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context that this tool is for previously saved screenshots, implying it should be used when an artifact ID is already available. It does not explicitly name alternatives or exclusions, but the sibling tool names and the phrase 'previously saved' give practical usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_screenshotsA
List all saved screenshot artifacts. Returns metadata for each saved screenshot including ID, filename, page URL, and timestamp.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses that the tool returns metadata (not image contents), which is useful. However, it doesn't mention pagination, limiting, ordering, or whether the list spans multiple sessions. The read-only nature is implied by 'List' but not explicitly stated. There is no contradiction with annotations since none exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description consists of two short sentences that are direct and information-dense. Every word contributes: 'List', 'all saved screenshot artifacts', and the specific metadata fields. No fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (no parameters, no output schema), the description is largely complete. It clearly states the tool lists all saved artifacts and specifies the metadata fields returned. It does not explain edge cases like empty lists or error scenarios, but for a list operation with no inputs, the provided information is sufficient for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema coverage is 100% (empty properties). The description adds no parameter details, but none are needed. By the rubric, zero params gives a baseline of 4, and the description does provide additional context about what the returned metadata includes (ID, filename, page URL, timestamp) which is helpful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('List all saved screenshot artifacts') and the resource (saved screenshots). It distinguishes from sibling tools like charlotte_screenshot (capture) and charlotte_screenshot_get (fetch one) by focusing on listing all saved artifacts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly indicates when to use the tool: when you need to list all saved screenshot artifacts. It doesn't explicitly mention exclusions or alternatives, but the context is unambiguous given the sibling tool names. A minor gap is lack of guidance on when to use this vs. charlotte_screenshot_get or charlotte_screenshot_delete, but the primary use case is evident.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_scrollA
Scroll the page or a specific container. Returns full page representation after scrolling.
| Name | Required | Description | Default |
|---|---|---|---|
| amount | No | Scroll amount: "page" (default), "half", or pixel value (e.g. "200") | |
| direction | Yes | Scroll direction | |
| element_id | No | Scroll within a specific container element |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It does disclose that the tool returns a full page representation after scrolling, which is useful behavioral context. However, it does not mention any potential side effects, prerequisites, or the nature of the scroll (e.g., instant, smooth). The safety/read-only nature is inferred but not stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences: one for the action and one for the return value. There is no redundant or filler content, and the key information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is relatively simple with 3 parameters, all well-documented in the schema. The description covers the action and explicitly mentions the return value, which is important since there is no output schema. However, 'full page representation' is a bit vague, and the lack of annotations leaves safety assumptions implicit. Overall, it is reasonably complete for a scroll action.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with each parameter (amount, direction, element_id) already described. The description adds little beyond the schema, only hinting at element_id via 'specific container'. Since the schema does the heavy lifting, a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Scroll') and a resource ('page or a specific container'), which clearly differentiates it from sibling tools like charlotte_navigate or charlotte_observe. It also states the outcome ('Returns full page representation after scrolling'), making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance is given about when to use this tool versus alternatives. The usage is implied by the action itself (scrolling), but there is no mention of exclusions or alternative tools, such as using navigation for moving between pages. This is a minimum viable level of guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_selectA
Select an option in a select/dropdown element. Returns full page representation after selection.
| Name | Required | Description | Default |
|---|---|---|---|
| value | Yes | Value or text of the option to select | |
| element_id | Yes | Target select element ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that the tool returns a full page representation after selection, which is useful, but it does not mention side effects, prerequisites (e.g., element visibility), or event triggering. This is a basic but not comprehensive behavioral description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences that immediately state the core function and the return behavior. It is front-loaded with the action, contains no fluff, and is appropriately sized for the tool's simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema, the description covers the essential points: what it does and what it returns. It could be marginally improved by noting that it is specifically for dropdowns, but the description and schema together provide sufficient context for an agent to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already covers 100% of parameters with clear descriptions ('Value or text of the option to select' and 'Target select element ID'). The tool description adds no additional semantics beyond restating the purpose, so the schema is the primary source of parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Select an option') and the target resource ('select/dropdown element'), making it distinct from sibling tools like charlotte_click or charlotte_type. It also notes the return behavior, further clarifying its function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for interacting with select/dropdown elements, which provides context on when to use it. However, it does not explicitly contrast it with alternatives like charlotte_click or charlotte_toggle, nor does it mention scenarios where it should not be used.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_submitB
Submit a form. Can submit by form ID or by clicking its submit button. Returns full page representation after submission.
| Name | Required | Description | Default |
|---|---|---|---|
| form_id | Yes | Form ID from page representation |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It does state the return value ('Returns full page representation after submission'), which is useful. However, it does not disclose that submitting a form is a mutating action with potential side effects (e.g., data changes, navigation, or irreversible submissions). The description mentions a 'clicking' method without clarifying whether it simulates a user click or requires the button to be visible, which is a behavioral gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise at two sentences, with the primary action front-loaded. The second sentence adds return information without excessive detail. The 'Can submit by form ID or by clicking its submit button' clause is somewhat ambiguous but does not significantly bloat the description. Overall, it is appropriately sized for a simple tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one parameter and no output schema, the description covers the core action and return value, which is fairly complete. However, it omits prerequisites (e.g., needing to be on a page with a form, ensuring the form_id is valid) and does not clarify the 'clicking' method. The mention of an alternative submission method without explaining how to invoke it via the schema reduces completeness. Given the tool's simplicity, a score of 3 is appropriate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers the single parameter 'form_id' with a description ('Form ID from page representation'), giving a baseline of 3. However, the description introduces an alternative submission method ('or by clicking its submit button') that is not represented in the schema, making the parameter semantics confusing. It doesn't add meaningful detail about how form_id is used or obtained, and the alternative method could mislead the agent into expecting an additional parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Submit a form.' This specifies the action (submit) and the resource (form), distinguishing it from sibling tools like charlotte_click or charlotte_type. However, the added 'Can submit by form ID or by clicking its submit button' introduces ambiguity about how submission is performed, detracting from full clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no explicit guidance on when to use this tool versus alternatives like charlotte_click. It mentions two submission methods (by form ID or clicking the submit button) but does not explain when one should be preferred, nor does it contrast with sibling tools that could also perform similar actions. This leaves the agent without clear selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_tab_closeA
Close a browser tab by its ID. If the closed tab was active, switches to the first remaining tab.
| Name | Required | Description | Default |
|---|---|---|---|
| tab_id | Yes | ID of the tab to close |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the behavioral disclosure burden. It adds useful context beyond the action itself: if the closed tab was active, it switches to the first remaining tab. This discloses a side effect that an agent would not otherwise know.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no redundant words. It front-loads the core action and then adds one key behavioral detail, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with no output schema, the description is largely complete. It could benefit from mentioning where the tab_id comes from (e.g., from charlotte_tabs), but the sibling context and clear action make this a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% because the only parameter is tab_id with a clear description. The description's 'by its ID' merely restates the schema, adding no additional semantic detail about the parameter's format or source.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the specific verb 'Close' and resource 'browser tab', clearly distinguishing this tool from siblings like charlotte_tab_open and charlotte_tab_switch. The addition 'by its ID' specifies the exact input needed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the usage context obvious: use this when you want to close a browser tab. It doesn't explicitly mention alternatives, but sibling tool names (tab_open, tab_switch) provide enough differentiation to infer when this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_tab_openA
Open a new browser tab. Optionally navigate to a URL. The new tab becomes the active tab.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL to navigate to (default: blank page) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It effectively discloses the key behavioral traits: a new tab is created, optional navigation occurs, and the new tab becomes active. For a simple tool, this is sufficient, though it could mention that the previous tab remains open.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences, with key information front-loaded. Every word earns its place, avoiding unnecessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with no output schema, the description fully covers the purpose and behavior. It is complete enough for an agent to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with the url parameter already described as 'URL to navigate to (default: blank page)'. The description does not add significant new semantics beyond repeating the optionality, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Open') and resource ('browser tab'), clearly stating the action. It also distinguishes itself from sibling tools like charlotte_tab_switch and charlotte_tab_close by focusing on opening a new tab.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: opens a new tab and optionally navigates to a URL, with the new tab becoming active. However, it does not explicitly mention when to use this tool instead of charlotte_navigate (which likely navigates the current tab), leaving a slight gap in alternative guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_tabsA
List all open browser tabs with their URLs, titles, and active status.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It clearly indicates a read-only listing operation and the content of the results. It does not mention pagination or ordering, but for a zero-parameter tool with straightforward behavior, this is adequately transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the action and resource. Every word adds value, with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, no-output-schema tool, the description fully specifies what the tool does and what information is returned. There are no missing details that would prevent an agent from using it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema coverage is 100% (vacuously). Baseline for zero parameters is 4. The description adds no parameter-specific detail because none exists, which is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description uses the specific verb 'List' and clearly identifies the resource as 'all open browser tabs', while also specifying the fields returned (URLs, titles, active status). This distinguishes it from sibling tools such as charlotte_tab_open, charlotte_tab_switch, and charlotte_tab_close.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for obtaining an overview of tabs, but it does not explicitly contrast with alternatives or state when to use this tool versus opening, switching, or closing tabs. Usage context is present but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_tab_switchA
Switch to a different browser tab by its tab ID. Returns the page representation of the activated tab.
| Name | Required | Description | Default |
|---|---|---|---|
| tab_id | Yes | ID of the tab to switch to |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of disclosure. It does reveal an important behavioral aspect: the tool returns the page representation of the activated tab. However, it does not mention what happens if the tab ID is invalid, whether focus changes, or any side effects beyond the switch. This is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences, no filler, and front-loads the action. Every word contributes to understanding the tool's purpose and return value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one parameter and no output schema, the description provides the essential information: what it does and what it returns. Some details like error cases are not covered, but the tool's simplicity makes the description sufficiently complete for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides 100% coverage for the single parameter tab_id with the description 'ID of the tab to switch to'. The tool description adds minimal extra meaning beyond restating 'by its tab ID', so it does not improve on the schema. Baseline 3 is appropriate given high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Switch') and resource ('different browser tab') with the input identifier ('by its tab ID'). It clearly distinguishes itself from sibling tools like charlotte_tab_open and charlotte_tab_close by focusing on switching to an existing tab, and it also specifies the return value.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the basic usage clear (switch to a tab by ID) but does not explicitly state when to prefer this over alternatives like charlotte_tabs or charlotte_tab_open. There are no exclusions or prerequisites mentioned, so guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_toggleA
Toggle a checkbox or switch element. Returns full page representation after toggle.
| Name | Required | Description | Default |
|---|---|---|---|
| element_id | Yes | Target checkbox or switch element ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses the return format ('full page representation') and implies a state change by using 'toggle', but it does not detail behavior in edge cases (e.g., if already checked), potential side effects, or whether it waits for changes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that states the action, the target element type, and the return behavior. It is concise with no redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter browser automation tool, the description adequately covers its purpose and return value. It lacks explicit prerequisites like visibility or interactability, but those are likely implied by the platform context and the simple nature of the action.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the only parameter 'element_id' is described in the schema. The tool description adds no semantic detail beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool's action as toggling a checkbox or switch element and specifies the return value (full page representation). This distinguishes it from siblings like click, select, and submit by focusing on checkbox/switch toggle behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: when the target element is a checkbox or switch. However, it does not explicitly contrast it with the click tool or state exclusions, so it stops short of full guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_toolsA
Manage Charlotte tool visibility. Lists available tool groups and their status. Use 'enable' or 'disable' to control which tools are loaded. Disabled tools don't appear in the tool list — enable a group to access its tools. Groups: 'interaction' for form filling, clicking, and drag-and-drop. 'session' for cookie/auth management, tab switching, viewport, and network. 'dev_mode' for local development serving and audits. 'evaluate' for JavaScript execution. 'monitoring' for console and network request logs. 'dialog' for JavaScript dialog handling.
| Name | Required | Description | Default |
|---|---|---|---|
| group | No | Tool group to enable or disable | |
| action | No | "list" (default) — show all groups and status. "enable"/"disable" — toggle a group. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of disclosing behavior. It explains that disabled tools don't appear and that enabling groups is needed to use their tools. It also defines the scope of each group. It could mention persistence or side effects, but for a visibility toggle this is sufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and appropriately sized. The first sentence states the purpose, the second explains behavior, and the rest enumerates groups efficiently. Every sentence provides useful information with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with two parameters, no output schema, and no annotations, the description thoroughly covers purpose, behavior, and parameter semantics. It gives enough context for an agent to select the right group and action.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with enums, so the baseline is 3. The description adds significant value by describing what each group contains (e.g., 'interaction' for form filling/clicking) and clarifying the default 'list' action, which goes beyond the generic schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb+resource: 'Manage Charlotte tool visibility' and immediately explains the list/enable/disable actions. It is easy to distinguish from sibling tools, which perform specific browser actions like clicking or navigating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use the tool: to list groups/status and to enable or disable groups. It also clarifies the consequence (disabled tools don't appear) and that enabling is required to access tools. It does not explicitly mention alternatives or exclusions, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
charlotte_typeA
Type text into an input element. Returns full page representation after typing.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | Text to enter | |
| slowly | No | Type one character at a time with a delay between keystrokes. Use for sites with autocomplete, search-as-you-type, or per-key validation (default: false) | |
| element_id | Yes | Target input element ID | |
| clear_first | No | Clear existing value before typing (default: true) | |
| press_enter | No | Press Enter after typing (default: false) | |
| character_delay | No | Milliseconds between keystrokes (implies slowly: true). Default when slowly is true: 50ms. Total typing time is capped at approximately 30s (including per-keystroke overhead); requests whose estimated duration exceeds that are rejected. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of disclosing behavioral traits. It states the return type (full page representation) but does not disclose that typing may clear existing content by default, can trigger events, or requires the element to be visible/interactable. The default clearing behavior (clear_first: true) is only discoverable through the schema, not the description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with no wasted words. It front-loads the primary purpose and adds a useful behavioral note about the return representation, fitting the appropriate structure for a simple tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having 6 parameters and no output schema or annotations, the description is minimal but sufficient to understand the core action. The schema covers parameter semantics, and the description states the return type. However, it lacks context about default behaviors (e.g., clearing the field, pressing Enter) that would help an agent anticipate side effects, making it adequate but not comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description itself does not add meaning beyond the schema, but the schema parameters are well-documented with descriptions for each field. Since the description does not need to repeat schema details, a score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Type') and resource ('input element'), and explicitly notes it returns the full page representation after typing. This distinguishes it from sibling tools like charlotte_click, charlotte_select, and charlotte_submit, which perform different actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool should be used when text needs to be entered into an input element, but it does not explicitly discuss when to use this tool versus alternatives, nor does it mention exclusions or prerequisites. No guidance is given for distinguishing between typing and using charlotte_submit or charlotte_click for form interactions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
23 tool updates
v0.8.0- Changed
charlotte_back2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_click2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_click_at2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_diff2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_find4 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / output_fileAdded value: +{ + "description": "Write the full match results to this file path instead of returning them inline. Relative paths resolve against output_dir (see charlotte_configure). Returns only a confirmation with the file path and size. Use for broad selectors (e.g. 'div', '*') that match many elements.", + "type": "string" +} - changed
Input schema / properties / selector / descriptionPrevious value: -"CSS selector to query the DOM directly. Returns elements that may not be in the accessibility tree. Results include Charlotte element IDs for use with interaction tools."New value: +"CSS selector to query the DOM directly. Returns elements that may not be in the accessibility tree. Results include durable Charlotte element IDs (dom-…) that remain valid across subsequent renders and interactions, and work with fill_form; they are re-resolved against the live DOM by re-running the selector."
- Changed
charlotte_forward2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_navigate2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_observe2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_reload2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_screenshot3 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - added
Input schema / properties / full_pageAdded value: +{ + "description": "Capture the entire scrollable page (default: true). Set false to capture only the current viewport — much smaller output for long pages. Ignored when 'selector' is provided.", + "type": "boolean" +}
- Changed
charlotte_screenshot_delete2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_screenshot_get2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_screenshots1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
charlotte_scroll2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_select2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_submit2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_tab_close2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_tab_open2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_tab_switch2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_tabs1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
charlotte_toggle2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_tools2 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false
- Changed
charlotte_type7 fields changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema" - removed
Input schema / additionalPropertiesRemoved value: -false - changed
Input schema / properties / character_delay / descriptionPrevious value: -"Milliseconds between keystrokes (implies slowly: true). Default when slowly is true: 50ms"New value: +"Milliseconds between keystrokes (implies slowly: true). Default when slowly is true: 50ms. Total typing time is capped at approximately 30s (including per-keystroke overhead); requests whose estimated duration exceeds that are rejected." - removed
Input schema / properties / press_enter / $refRemoved value: -"#/properties/clear_first" - added
Input schema / properties / press_enter / typeAdded value: +"boolean" - removed
Input schema / properties / slowly / $refRemoved value: -"#/properties/clear_first" - added
Input schema / properties / slowly / typeAdded value: +"boolean"
23 tool updates
v0.6.3- First observed
charlotte_back - First observed
charlotte_click - First observed
charlotte_click_at - First observed
charlotte_diff - First observed
charlotte_find - First observed
charlotte_forward - First observed
charlotte_navigate - First observed
charlotte_observe - First observed
charlotte_reload - First observed
charlotte_screenshot - First observed
charlotte_screenshot_delete - First observed
charlotte_screenshot_get - First observed
charlotte_screenshots - First observed
charlotte_scroll - First observed
charlotte_select - First observed
charlotte_submit - First observed
charlotte_tab_close - First observed
charlotte_tab_open - First observed
charlotte_tab_switch - First observed
charlotte_tabs - First observed
charlotte_toggle - First observed
charlotte_tools - First observed
charlotte_type
TDQS
Scored across 23 tools
Each tool has a clearly distinct purpose: navigation, tab management, interaction, observation, and screenshot management are all separated. Even similar tools like charlotte_click and charlotte_click_at are explicitly differentiated by target type (element vs coordinates).
All tools share the 'charlotte_' prefix and use snake_case, but there is a mix of simple verbs (navigate, click, type) and compound verb_noun forms (tab_open, screenshot_get). This is mostly consistent but not perfectly uniform.
With 23 tools, the server is on the heavier side but still within a manageable range for a comprehensive browser automation tool. The count is justified by the breadth of features, though it approaches the threshold where it might feel overwhelming.
Core browser automation workflows are well covered: navigation, tab management, element interaction, observation, and screenshot handling. However, the description of charlotte_tools mentions groups for dialogs, drag-and-drop, and session management, but these tools are not present in the exposed set, leaving minor gaps.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Scrape, crawl and search the web for AI agents via MCP.
Docs: https://docs.keenable.ai/mcp-server Keenable is a free, remote MCP server that gives agents access to the web index. Search the web with ranked results and date/site filters, then fetch any indexed page as clean markdown. Works out of the box with no account or API key.
Reliable web access for AI agents: smart HTTP, rotating proxies, and full-browser rendering.
- RampifyOAuthdev.rampify
SEO MCP server: crawl your site, find AI-visibility gaps, and ship the fix from your coding agent.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceAn advanced MCP server for browser automation using Puppeteer, specifically optimized for token efficiency through minimal data returns and progressive enhancement. It enables agents to navigate pages, capture LLM-optimized screenshots, extract structured content, and perform batch interactions.3-
- AlicenseNot gradedqualityDmaintenanceMCP server for web scraping and browser automation, enabling AI agents to extract clean, token-efficient content from web pages.1MIT

Wickofficial
FlicenseNot gradedqualityAmaintenanceAn MCP server that provides browser-grade web access for AI agents, using Chrome's actual network stack to bypass anti-bot protections and return clean markdown.8-- AlicenseAqualityDmaintenanceAn MCP server providing AI agents with a stealth Chromium browser that uses hybrid accessibility-object-model and set-of-mark vision for token-lean snapshots and reliable action via ref ids.13601Apache 2.0
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/TickTockBent/charlotte'
If you have feedback or need assistance with the MCP directory API, please join our Discord server