screen-context
Summary: This MCP server provides read-only tools to search and summarize captured screen/OCR history as untrusted observed data.
Full-text search previously captured window OCR text by query, with limit and optional time filter.
Summarize recent screen activity into consecutive-frame blocks with app, window title, time range, and text.
Get activity blocks before and after a Unix timestamp.
Get bounded screen observations for a specific local calendar day to draft reports or diaries, with timezone, limit, cursor, and purpose options.
All listed tools are read-only, non-destructive, and return untrusted observed data; they do not write or modify history.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@screen-contextfind the error message I was looking at in the browser a moment ago"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
ScreenContext records the foreground window, reads its text with the operating system's built-in OCR, keeps an encrypted local history, and exposes that history to AI assistants through the Model Context Protocol (MCP). Any MCP client can use it: Claude Code, Claude Desktop, Cursor, VS Code, Windsurf, Codex CLI, Gemini CLI and others.
Ask your assistant things like "find the error message I was looking at in the browser a moment ago" or "what was the spec page I read this morning?".
Platform | Status |
macOS 14+ | Supported (menu bar app + CLI). Built from source; there is no notarized download. |
Windows 10/11 | Beta (control window + CLI). Checked on one Windows 11 machine; see Windows (beta). |
Current version: 0.2.0. Changes: CHANGELOG.md. License: MIT.
What 0.2 adds
0.2 is about trust: you decide which assistants may read your history, what is never recorded, and how long anything is kept.
Per-client access: each MCP client gets its own token and must be approved once in the app. You can list clients, see what each one read, and revoke any of them.
Audit log: every query, with the frames it returned, is recorded in the encrypted database. Only you can read it (CLI), never an assistant.
Sensitive input is not stored: card numbers and My Number drop the whole frame; a name next to an address, phone number, date of birth or e-mail address is redacted. Contacts and checkout pages are excluded by default.
Data control: purge by time, app or keyword; retention limits; usage report; full wipe; passphrase-protected backup and restore; export of a time range.
Upgrading from 0.1 takes a few steps: see Upgrading from 0.1.
Related MCP server: ContextPulse
How it works
Three processes, each with a narrow job:
Capture (menu bar app on macOS, control window on Windows) captures only the foreground window at native resolution. Similar frames are skipped with a perceptual hash. Frames go to an encrypted spool.
Indexer runs OCR (Apple Vision on macOS,
Windows.Media.Ocron Windows), applies your exclusion and sensitive-input rules, and stores text in an SQLCipher database with a trigram FTS5 index.MCP server (
screen-context serve) is started by your MCP client. It never imports capture code, reads screen data only (it writes nothing but audit rows and proposals), requires a per-client token, and labels every result as untrusted observed data.
More detail: docs/ARCHITECTURE.md. Planned work: docs/ROADMAP.md.
Requirements
Python 3.11 or later and uv
macOS 14 or later, or Windows 10/11 (beta)
To build the macOS
.appwith py2app, use a python.org or Homebrew Python. Some standalone Python builds fail in py2app becausezlibis built in.
Quick start (macOS)
1. Install and create the encrypted store
git clone https://github.com/ikeikeikeda66/screen-context-agent.git
cd screen-context-agent
uv sync --locked --extra macos --extra encrypted --extra dev
.venv/bin/screen-context initinit creates ~/Library/Application Support/ScreenContext and stores a random encryption key in the macOS Keychain. If you lose the key, the history cannot be decrypted; screen-context backup keeps a passphrase-protected copy.
2. Start the indexer
.venv/bin/screen-context index --watchTo start it at login instead, see packaging/macos/README.md.
3. Build, sign and start the menu bar app
macOS grants Screen Recording permission to a signed app, so capture runs from the app, not from the terminal. Sign every build with the same identity and the permission survives rebuilds:
sh packaging/macos/signing_identity.sh # once per Mac: a self-signed identity in your login keychain
SCREEN_CONTEXT_SIGN_IDENTITY="ScreenContext Local Signing" sh packaging/build_mac.sh
open dist/ScreenContext.appAllow ScreenContext in System Settings > Privacy & Security > Screen Recording, then open the app again. After a rebuild, macOS may ask once whether codesign may use the signing key; choose Always Allow.
A Developer ID identity works the same way.
SCREEN_CONTEXT_SIGN_IDENTITY=-(ad hoc) also builds, but the permission must be granted again after every rebuild.The app is not notarized. Build it on the Mac that runs it: a copy downloaded or moved from another Mac is blocked by Gatekeeper.
To try capture once with the terminal's own permission:
.venv/bin/screen-context capture --once.
4. Connect your assistant: see Connect an MCP client.
Menu bar

The menu bar shows SC Rec while recording and SC Paused while paused.
Pause Capture / Resume Capture: use this while watching video or showing private content. The state is saved.
Capture Interval: 5 s, 15 s (default), 30 s, 1 min, 2 min or 5 min. The change applies without a restart. Unchanged screens are not saved again.
Language: System Default, English or 日本語. All menu text, dialogs, and generated material follow this setting.
Connect an MCP client
Print a ready-to-paste entry for your client:
.venv/bin/screen-context mcp-config --client claude-code # prints a `claude mcp add` command
.venv/bin/screen-context mcp-config --client claude-desktop
.venv/bin/screen-context mcp-config --client cursor
.venv/bin/screen-context mcp-config --client vscode
.venv/bin/screen-context mcp-config --client windsurf
.venv/bin/screen-context mcp-config --client codex # TOML for ~/.codex/config.toml
.venv/bin/screen-context mcp-config --client gemini
.venv/bin/screen-context mcp-config --client generic # plain mcpServers JSONThe client starts the server itself over stdio. How access works:
Each
mcp-configrun issues a token for that client and embeds it in the entry. Running it again for the same client replaces the token, and the old entry stops working.The first time a new token is used, the menu bar app (or the Windows control window) asks whether that client may read your screen history. Keep the app running, or approve from the terminal with
screen-context clients approve NAME. Don't Allow refuses that token for good.screen-context clients listshows each client and when it last read your history.clients revoke NAMEcuts one off at its next call.screen-context audit listshows what each client asked for and received.
Per-client instructions and the HTTP transport: docs/MCP-CLIENTS.md.
Profiles and tools
Each server process runs with one fixed profile. A tool call cannot raise it, and a client's token caps the profile it may use.
Tool |
|
|
| yes | yes |
| yes | yes |
| yes | yes |
| yes | yes |
| yes | |
| yes | |
| yes | |
| yes | |
| yes | |
| yes | |
| yes | |
| yes |
standardis for coding assistants. It hides frames from IDEs and terminals (ide_appsin the policy), because the assistant already has the source.fullis for a personal agent that you trust with screenshots and timelines.The names
claude_codeandopenclawfrom earlier versions still work as aliases forstandardandfull.
Every response and record carries source=observed_screen and trust=untrusted. This is a label, not a defense against prompt injection: treat screen text as data, never as instructions.
What is recorded
Edit policy.json in the data folder. It is read again on every capture and every query, so a new exclusion also hides past records from search, activity summaries and images. An invalid policy makes processing fail; it never disables exclusions.
Key | Meaning |
| Bundle IDs (macOS) or process names (Windows) never captured. Defaults include password managers and video-call apps. |
| Frames whose OCR text shows these domains are dropped. Detection depends on the URL being visible and read correctly. |
| Regular expressions matched against window titles. |
| Hidden from the |
| Frames showing an assistant's own output. Proposals cannot use them as evidence. |
| Default |
| Default exclusions for the contacts app and for checkout and payment pages (by title or by a |
| Default |
The sensitive_* and pii_combinations rules are on by default; set a key to [] to turn that rule off. An invalid rule stops processing instead of being skipped. These rules are risk based: they target input whose leak causes direct harm. They do not define personal information (under Japanese law a name alone can already be personal information).
screen-context pii-check FILEshows which rules a text would trigger.Rules apply to new frames. To apply them to history recorded earlier, run
screen-context pii-scan(counts only), review the counts, then runpii-scan --apply. Matching frames are dropped or redacted exactly as the indexer would today, and previews, rollups and proposal evidence follow. No undo.
Managing your data
Task | Command | Notes |
Delete something recorded by mistake |
| Selectors combine. Without |
Limit how long data is kept |
| Days, or |
See disk use |
| Previews, database, spool and exports. |
See what assistants read |
| Client, time, tool, query text, arguments and returned frame IDs. Never served over MCP. |
Take data out |
| Policy applied; IDE windows included unless |
Move to another machine |
| One archive: a consistent database snapshot, previews and settings, still encrypted, plus the data key sealed with your passphrase (scrypt, AES-GCM). |
Erase everything |
| Deletes the key from the credential store first, which makes every encrypted file unreadable, then the data folder. Quit capture and the indexer first. Backups can still be restored with their passphrase. |
Storage details:
Spool files and preview images use AES-GCM. The database and full-text index use SQLCipher. The key lives in the OS credential store (Keychain or Windows Credential Manager), or in
SCREEN_CONTEXT_KEYfor headless use.The spool stops accepting frames at 100 files or 512 MB. Unprocessed frames older than 24 hours are deleted by
maintain.SCREEN_CONTEXT_PLAINTEXT=1is for development tests only. ScreenContext never falls back to plaintext on its own.
Threat model
ScreenContext protects against:
Someone who has the files but not the key: a stolen disk, a copied data folder, a synced backup. The database, spool and previews are encrypted, and
backuparchives need their passphrase.An MCP client reading more than you allowed: each client has its own token, is approved once by you, is limited to its profile, and can be revoked. The audit log shows what each client asked for and received.
Recording what should never be kept: exclusions, the card-number and My Number detectors, the personal-data combination rule, and
purge.
It does not protect against:
Other programs running as your user. They can start
screen-context servewith a token copied from a client's configuration file, or read the key from the credential store, and so read your history without Screen Recording permission. A self-signed build also turns off library validation (the embedded Python cannot load without it), so such a program could plant code in the app and use its Screen Recording permission. Keeping the key inside a signed app is planned only if a Developer ID is adopted (roadmap).A local administrator. Managed settings prevent mistakes and policy violations; they are not DRM.
Instructions shown on screen (prompt injection). Results are labeled untrusted; clients must treat them as data.
Copies outside the store. Exports and results already returned to a client are not reached by later exclusions, purges or retention.
purgetells you when such copies exist.What OCR or the rules miss. Detectors are pattern based: a misread card number or an unlabeled name is stored.
Known limitations
macOS distribution: the app is not notarized, so it must be built on the Mac that runs it (see Quick start).
Password fields: capture is not yet paused while macOS Secure Input is on (#16). Password managers are excluded by default, and password fields show masked characters.
Personal-data redaction depends on OCR. OCR sometimes returns a second, truncated reading of the same line; such a fragment can escape redaction (for example a bare domain from a redacted e-mail address) (#55).
Windows is beta: see below.
Upgrading from 0.1
The database schema, MCP entries and approval flow changed. Existing history is kept.
Quit the app and the indexer, and copy the data folder somewhere safe.
Update:
git pull, thenuv sync --locked --extra macos --extra encrypted --extra dev(--extra windowson Windows).Run
screen-context init. It migrates the database to schema v3 and moves the oldaudit.jsonlinto the encrypted audit log. Until then,healthreportsneeds_init.On macOS, rebuild the app (step 3 of the Quick start).
Run
mcp-configagain for every client and replace its ScreenContext entry. Entries from 0.1 no longer start the server. Approve each client at its first call.Optional: run
screen-context pii-scan, thenpii-scan --apply, to apply the new sensitive-input rules to history recorded before 0.2.
Configuration
Environment variable | Purpose |
| Data folder. Default: |
| 64 hex characters. Replaces the OS credential store (headless use). |
|
|
| Comma-separated OCR languages, for example |
| The client's token, set by |
The language can also be set with screen-context language en|ja|system.
CLI reference
# Setup and workers
screen-context init create the data folder, key and database (also migrates)
screen-context index [--watch] OCR and store spooled frames
screen-context capture [--once] capture from the terminal (development)
screen-context pause | resume stop or restart new captures
screen-context status | health queue and worker state
screen-context maintain retention and daily rollups
screen-context language [system|en|ja]
# MCP clients
screen-context serve [--profile standard|full] [--transport stdio|http] [--port 8765]
screen-context mcp-config [--client NAME] [--profile standard|full] [--name TOKEN_NAME]
screen-context clients list | approve NAME | revoke NAME
screen-context audit list|export [--client NAME] [--since YYYY-MM-DD] [--limit N]
# Your data
screen-context purge [--from T] [--to T] [--last 15m] [--app ID] [--keyword TEXT] [--block ID] [--excluded] [--yes]
screen-context retention [--preview D] [--text D] [--audit D] days to keep, or none
screen-context usage disk space by kind of data
screen-context pii-check FILE which sensitive-input rules a text would trigger
screen-context pii-scan [--apply] apply the rules to frames recorded earlier
screen-context export --from T [--to T] [--format jsonl|md|csv|viking] [--out DIR] [--exclude-ide]
screen-context backup FILE | restore FILE [--replace] passphrase-protected archive
screen-context wipe delete the key, then all data (no undo)
# Optional
screen-context diary-material DATE [--budget 6000] [--lang en|ja]
screen-context proposal prepare|simulate|finish
screen-context push DATE send one day to a local OpenViking serverexport DATE (one day for OpenViking) still works in 0.2 but is deprecated; use --format viking.
Optional: diary material and periodic proposals
diary-material DATEprints a compact, bounded Markdown summary of one day, for use as input to a diary or daily report prompt.proposal prepareis designed as a pre-run script for a scheduler (cron or an agent framework). It prints material only when there are new observations. Otherwise its last line is{"wakeAgent": false, ...}, so the scheduler can skip starting the agent. The agent registers a suggestion withsubmit_proposal; evidence must be a quote from a non-assistant frame, and the same conclusion is suppressed for 24 hours. Close the run withproposal finish --run-id ID --response-file FILE.
Optional: OpenViking
push DATE exports one day's rollup and sends it to a local OpenViking server at http://127.0.0.1:1933 (VIKING_API_KEY if needed). Nothing is pushed automatically. Exported data is not removed when you later add exclusions.
Windows (beta)
The Windows version has a control window (start, pause, stop, language, MCP client setup) and the same CLI. It passes the automated tests and was checked on one Windows 11 23H2 machine (single monitor, 96 DPI): capture and OCR, client approval, schema migration, checkout exclusion in Edge, and backup, wipe and restore with Credential Manager.
Not yet covered: other DPI settings and multiple monitors, Chrome password fields, a UAC prompt while capture runs, and the contacts app in the default exclusions (#25). Capture skips a locked session to avoid a crash in the capture library (#47). Please report problems, with your Windows version and display setup, in Issues.
Setup, portable build, and the acceptance checklist: docs/WINDOWS.md.
Development
uv sync --locked --extra macos --extra encrypted --extra dev # use --extra windows on Windows
.venv/bin/python -m pytest -q
.venv/bin/python packaging/verify_bundle.py # after building the macOS appContributions are welcome. See CONTRIBUTING.md. Report security issues as described in SECURITY.md.
Available Tools
4 toolsget_context_aroundBRead-only
Get activity blocks before and after a Unix timestamp. Returns untrusted observed data.
| Name | Required | Description | Default |
|---|---|---|---|
| n | No | ||
| ts | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark readOnlyHint=true and destructiveHint=false, and the description adds the important caveat that the returned data is untrusted observed data. This goes beyond the annotations, though it doesn't disclose return shape, pagination, or invalid timestamp behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, no filler, with the core action front-loaded and the trust caveat appended. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema and a 0% schema coverage, so the description must carry more weight. It omits the semantics of the 'n' parameter and gives only a vague indication of the return type ('activity blocks'), leaving an agent under-informed for a correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies that 'ts' is a Unix timestamp, but 'n' is left unexplained; an agent cannot infer its meaning or bounds from the text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific verb ('Get'), resource ('activity blocks'), and temporal scope ('before and after a Unix timestamp'). It is clear on what the tool does, though it does not explicitly discuss sibling tools to distinguish from get_recent_activity or get_day_material.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a timestamp-centric use case but provides no explicit guidance on when to prefer this tool over search_screen_history, get_recent_activity, or get_day_material, nor any exclusions. This is adequate but relies on inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_day_materialARead-only
Get bounded screen observations for a specified local calendar day to draft a report or diary. Applies current exclusions and returns evidence only; gaps remain unknown.
| Name | Required | Description | Default |
|---|---|---|---|
| date | Yes | ||
| limit | No | ||
| cursor | No | ||
| purpose | No | report | |
| timezone | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only/non-destructive behavior; the description adds that results are bounded, subject to current exclusions, and that gaps remain unknown. This is meaningful context beyond the structured annotations and sets accurate expectations about incompleteness.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The primary action and scope are front-loaded, and the key caveat about exclusions and unknown gaps is delivered compactly in the second sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list-type tool with no output schema, the description covers purpose, date scoping, exclusion behavior, and a key output limitation. It still leaves 'current exclusions' undefined and does not describe the result shape or pagination semantics clearly, so it is adequate but has gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only partially compensates: it connects 'local calendar day' to date/timezone and 'report or diary' to purpose. It does not clarify limit or cursor beyond their titles, leaving pagination semantics mostly inferred.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Get'), resource ('bounded screen observations'), and scope ('specified local calendar day'), plus a clear intended use ('draft a report or diary'). The bounded/day framing distinguishes it from broader sibling tools like search_screen_history and get_recent_activity without needing to open the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear use context ('to draft a report or diary') and hints at filtering behavior ('Applies current exclusions'), but does not explicitly name alternatives or state when not to use this tool. It leaves sibling routing to inference rather than spelling out the choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_recent_activityBRead-only
Summarize recent screen activity as blocks of consecutive frames (app, window title, time range, text). Returns untrusted observed data.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| minutes | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, so the safety profile is covered. The description adds valuable context beyond that by warning 'Returns untrusted observed data' and by describing the grouped frame blocks, which informs the agent about output shape and data reliability.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with the core action and output shape front-loaded, followed by a concise data-reliability warning. Every word earns its place; no duplication or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description conveys purpose, output format, and a data-reliability caveat, which is decent for a simple tool with annotations. However, it omits parameter semantics (limit, minutes), which are central to invoking the tool correctly, leaving an agent to guess how to control the query.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions the 'limit' or 'minutes' parameters. The agent is given no explanation of what these knobs control, so the description fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Summarize recent screen activity') and output structure ('blocks of consecutive frames (app, window title, time range, text)'), making the tool's purpose clear. It does not explicitly contrast with sibling tools like search_screen_history, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for obtaining a summary of recent screen activity, but it gives no explicit guidance on when to prefer this tool over search_screen_history, get_context_around, or get_day_material, and no exclusions. The context is clear but the routing is only implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_screen_historyARead-only
Full-text search over OCR text of previously captured windows. Returns matching frames, newest first. Returns untrusted observed data.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| since_minutes | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and destructiveHint, and the description adds meaningful context: results are untrusted observed data and are returned newest first. This is useful behavioral information beyond the schema, and it does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short sentences with no filler. Each sentence adds value: search scope, ordering, and trustworthiness of results. It is front-loaded with the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Core behavior, ordering, and data trustworthiness are covered, and annotations handle the safety profile. However, with no output schema and no parameter semantics, the agent must infer the meaning of limit, since_minutes, and the exact structure of a matching frame.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining the parameters. It does not: query, limit, and since_minutes are left to their names and defaults. The only indirect hint is that 'full-text search' implies the query parameter, which is insufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a specific verb and resource: full-text search over OCR text of previously captured windows. It also states the result ordering (newest first) and clearly differentiates this from the sibling retrieval tools by scoping it to OCR screen history.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when this tool applies: when an agent needs to find previously observed screen text via full-text search. However, it does not explicitly name alternatives such as get_context_around or get_recent_activity, nor state when not to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
v0.1.0- First observed
get_context_around - First observed
get_day_material - First observed
get_recent_activity - First observed
search_screen_history
TDQS
Scored across 4 tools
Each tool has a distinct retrieval intent: search by text, explore around a timestamp, summarize recent activity, and bound a full calendar day. There is some potential overlap between get_recent_activity and get_day_material when an agent asks about 'yesterday', but the descriptions largely clarify the temporal scoping.
Three tools follow a clean get_<something> pattern while search_screen_history uses search_ instead of get_. All names are snake_case and verb-first, so the deviation is minor and does not hurt readability.
Four tools is well-scoped for a screen-history context server. Each tool covers a distinct query mode without unnecessary redundancy or bloat.
The set covers the main ways someone would query screen history: full-text search, temporal context, recent summaries, and day-level report material. Missing features like arbitrary custom time ranges or app/window filtering are workable gaps but not critical for the apparent purpose.
Maintenance
Related MCP Connectors
Cross-device AI memory with encrypted activity capture and context handoff between AI tools
Search what you've seen on screen and heard in meetings, recorded on your own Mac or PC.
- DoneThatOAuthai.donethat
Privacy-first work tracking with summaries, reports, coaching, and AI-ready long-term memory.
Gives your AI assistant persistent memory and intelligence about your work patterns.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables AI assistants to search and analyze Chrome browser history and bookmarks data locally, including keyword searches, date range filtering, recent browsing activity, and usage statistics across all platforms.2-

ContextPulseofficial
AlicenseAqualityBmaintenanceLets AI assistants understand what you're working on — current screen content, recent dictation, clipboard, and saved notes — running entirely on your own machine with nothing sent to the cloud.36AGPL 3.0- AlicenseAqualityDmaintenanceEnables AI assistants to query your local browsing history using keyword search, RAG Q&A, and activity stats.6MIT
- FlicenseNot gradedqualityBmaintenanceEnables AI assistants to search and retrieve memories from your Mac, including screen captures, meeting transcripts, and browsing history, all locally and privately.1-