constant-watch
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@constant-watchwhat did I see about Atlas earlier today?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Constant Watch
Open-source screen memory for macOS and Windows. Accessibility + OCR → local Qwen3.5 0.8B → continuous daily Markdown → read-only MCP.
The Mac app uses SwiftUI and a menu-bar companion. The Windows app uses a native WebView2 window with Microsoft UI Automation and Windows OCR. Both share the local Python journal and model backend.
Download installers · Website · MIT license · Contributing
Preview requirements: Apple Silicon / macOS 14+, or Windows 11 Intel/AMD x64. Ollama and the model are separate downloads. The Mac package is signed but not notarized; the Windows installer is unsigned. The transparent desktop orb intro is currently Mac-only.
The default view is Daily flow: one chronological journal across apps, with adjacent captures from the same app and window grouped into sessions. A white interface, guided onboarding, and a local landing page introduce the product without sign-up or authentication.
First run
The welcome opens with a transparent, full-desktop orb-to-logo reveal, followed by a short animated greeting and an optional first name saved on this Mac. Music on/off controls a bundled ambient soundtrack; it stops when you leave onboarding. Replay welcome reopens the greeting. macOS Reduce Motion disables staged reveals.
Choose a use case: recover context, review the day, or connect an assistant.
Enable Accessibility and Screen Recording separately. Each permission has a live status and a link to its macOS settings page.
Check the local Qwen model, or download it through the app when Ollama is running.
Open the practice note. Leave it visible for about 15 seconds, then return to see the actual captured result. Search Atlas to find it again.
New installations start paused. Explore first opens existing journals with capture paused; Open my daily flow completes setup and starts capture. The guide can be reopened from the sidebar or menu bar. Existing saved pause preferences are preserved on upgrade.
Practical uses: recover a detail seen in another app, return to an interrupted project, review the day's research and decisions, export a daily work journal, and supply a connected assistant with searchable context. This is text recall, not an activity-duration tracker or an automatic agent acting on your behalf.
The foreground app is sampled every 10 seconds by default. This is periodic text capture, not a video recorder: brief changes between samples can be missed. Only the foreground app's window is captured, not every display or background app.
Related MCP server: neurakeep
Install the packaged app
Open dist/Constant-Watch-0.1.0-macOS-arm64.dmg, drag Constant Watch into Applications, eject the disk image, and launch the installed app. The package includes its Python interpreter and dependencies; no developer tools or source checkout are required. Apple Silicon and macOS 14+ are required. Ollama and the Qwen model are separate first-run downloads. The DMG does not enable login startup automatically.
The package is Developer ID signed when that identity is available. Check its release report for notarization status; signing alone is not notarization. Upgrading from a build signed with a different certificate may require granting the new app's macOS permissions again.
Build and verify a package:
bash scripts/package-app.sh
.venv/bin/python scripts/smoke-package.py "$HOME/Library/Caches/Constant Watch/build/Constant Watch.app"Set CODE_SIGN_IDENTITY to select a certificate explicitly. Set CW_NOTARY_PROFILE to an existing notarytool Keychain profile to submit, staple, and validate notarization during packaging. Outputs include a compressed DMG and SHA-256 checksum. The smoke test uses a temporary journal, starts paused, checks packaged web assets, and exercises the actual bundled stdio MCP server.
Run from source
Requirements: Apple Silicon, macOS 14+, Xcode Command Line Tools (xcode-select --install), uv, and running Ollama.
bash scripts/setup.sh
bash scripts/install-app.shOpen ~/Applications/Constant Watch.app. The installer starts it and configures login startup. If access is missing, use Request access, then enable Constant Watch in System Settings → Privacy & Security → Accessibility and Screen Recording. The native permission card links directly to both settings pages. The native app itself requests and uses both permissions; Python and the command-line helper do not need access for the installed app. Quit/reopen if macOS requests it. If an older development build shows enabled but is not recognized, remove its stale entry and add the installed app again using Show installed app. The permission check does not capture text.
.venv/bin/constant-watch doctor
.venv/bin/constant-watch permissionsThe native app owns the local Python service. Closing the window leaves the menu-bar app and capture running; Quit Constant Watch stops both. Use the menu-bar icon to reopen the journal, pause/resume, show notes in Finder, or copy MCP configuration.
To update the installed app after source changes, quit it, then rerun bash scripts/install-app.sh. To refresh only its backend runtime and login startup:
.venv/bin/python scripts/service.py installThis installs a separate runtime under ~/Library/Application Support/Constant Watch/runtime so launch does not depend on access to the protected Documents directory. Stop a manually running server before installing. A per-data-directory lock prevents two capture loops running together. The native app starts the service as its child and performs all macOS permission checks, Accessibility reads, and OCR capture in its own process. The backend sends jobs over a local bridge authenticated with a fresh per-launch token; it never launches a capture helper in this mode.
To remove login startup (retains journals, app, and runtime), quit Constant Watch first, then:
.venv/bin/python scripts/service.py uninstallLogs: ~/Library/Logs/Constant Watch/. Pause capture persists across restarts; existing queued summaries can finish while capture is paused. Ollama must also be running for summaries; source capture continues and queued summaries retry when it returns.
The landing page is at http://127.0.0.1:8765 while the app is running. Its interactive example journal uses clearly labeled fictional content and never loads your private history. Native app links use the constantwatch:// URL scheme. The optional browser journal is at http://127.0.0.1:8765/journal. For backend-only development, run .venv/bin/constant-watch serve instead of the native app.
Your data
~/Library/Application Support/Constant Watch/
settings.json
memory.sqlite3
days/
2026-09-28.md
apps/
com.apple.Safari-<stable hash>/2026-09-28.md
com.microsoft.VSCode-<stable hash>/2026-09-28.mdApp folders use bundle IDs and a stable hash to avoid unsafe paths and naming collisions. Every observation includes its timestamp, window title, generated summary, separate accessibility/OCR source text, and capture warnings. SQLite is the source of truth; Markdown is rewritten atomically and can be regenerated with constant-watch rebuild. Do not edit generated Markdown expecting those changes to persist.
days/YYYY-MM-DD.md is the continuously updated cross-app journal. It includes an overview, chronological sessions, summaries, and expandable source evidence. Returning to an unchanged app is preserved as a transition. Consecutive unchanged samples extend the last-seen timestamp; a different app/window or a gap over five minutes starts another session. Time ranges describe observations, not continuous active time. Local session summaries consolidate multiple captured moments and refresh at most once a minute while the session changes; pending updates are indicated. App-specific journals still retain every original observation.
Search covers captured text, titles, and summaries. Select an app and date to download its daily Markdown. Settings control interval (3–300 seconds), exclusions by bundle ID, and retention (1–365 days, default 30). Reducing retention deletes old database records and Markdown exports. Retention is logical deletion, not forensic secure erasure; backups may retain copies.
For command-line operation, override the storage root with CONSTANT_WATCH_DATA and the helper executable with CONSTANT_WATCH_HELPER. The installed native app uses its own in-process capture core. All clients reading the same journal must use the same data root.
MCP
Use Copy MCP configuration in the native app's settings or menu bar, or mcp-config.example.json in this checkout, for a stdio MCP client. Replace the example command with your installed executable path. Generic configuration:
{
"mcpServers": {
"constant-watch": {
"command": "/absolute/path/to/constant-watch/.venv/bin/constant-watch",
"args": ["mcp"]
}
}
}For the installed service, the command can instead be /Users/YOUR_USER/Library/Application Support/Constant Watch/runtime/venv/bin/constant-watch. The MCP process reads the shared journal independently of the capture daemon. It does not start capture or expose capture controls.
Tools:
Tool | Purpose |
| Questions answered with exact captured excerpts, observation citations, app/time/document/source metadata; optional app/date/topic filters |
| Suggested groups across apps based on explicit project names and document titles |
| Chronological, paginated captures for an exact topic key |
| Verify a citation against original Accessibility/OCR text |
| One complete chronological daily Markdown journal across apps |
| Paginated grouped sessions with summaries and source evidence |
| Applications, bundle IDs, counts, available dates |
| Full-text AND search; optional app and date filters |
| Latest observations; pagination with |
| Complete daily Markdown for one application |
Resources: watch://day/{day}, watch://apps, watch://app/{app_id}/{day}, and watch://observation/{observation_id}. Dates use local YYYY-MM-DD. Search/recent tools return at most 50 observations and day_sessions at most 50 sessions per call. Complete journals can be large; prefer pagination for high-volume days.
Ask your day and follow a topic
The native workspace opens with Ask your day. Ask about a project or detail, use Today only or an app filter, then open View captured text to verify the source. Topics follows matching names across applications in chronological order. The daily journal and per-app Markdown remain available; daily Markdown also lists suggested threads spanning multiple apps.
Recall currently retrieves and quotes passages rather than generating a new factual answer. It handles question filler, a small set of common aliases, dates, and word prefixes; it is not semantic search and may miss paraphrases. Every meaningful query term must match. Qwen3.5 continues to generate background observation/session summaries using cleaner text and a prompt focused on concrete facts. Older summaries keep their attribution.
Raw Accessibility/OCR text is preserved after existing redaction. A separate searchable representation removes exact duplicate lines and common interface labels without collapsing signs or changed numbers. Existing captures are migrated automatically. Current-document URLs come only from Accessibility document/web-area metadata; unavailable historical links are never guessed. Secret query parameters and URL credentials are removed. Open source opens HTTP(S) links; Reveal document shows a captured local document in Finder. constantwatch://observation/123 opens a cited capture in the native app.
Topic labels are deterministic suggestions based on visible names/titles, not confirmed project membership. A screenshot or visible message does not prove an action was completed. Retention removes the associated evidence and search entries; opening an expired citation reports it as unavailable.
Captured content is marked as untrusted reference data, not instructions. Any assistant connected to this MCP can read your journal; its own data handling policies apply. MCP registration is deliberately provided as a file so you can choose which assistant receives access.
Architecture
Shared Swift capture core: AppKit identifies the foreground app; AXUIElement reads its focused window, with depth, node, and time bounds. ScreenCaptureKit captures that app's foreground window in memory, and Apple Vision recognizes text. A focus change during capture discards the sample. No screenshots are saved. The installed app calls this core directly; a separate helper exists only for command-line operation.
Native SwiftUI app: NavigationSplitView app sidebar, real application icons, searchable journal, source disclosure, native date picker and save panel, settings sheet, and MenuBarExtra controls. It starts/stops the Python backend as a child process. No WebView or Electron runtime.
Python daemon: redacts common token/password patterns, merges duplicate AX/OCR lines, detects unchanged text per app and day, stores observations, and exports Markdown. Capture and summary workers run separately so model latency doesn't block sampling. Native bridge requests have bounded queues and deadlines; cancelled or expired results are discarded. Standalone helper processes have a timeout and are terminated on shutdown.
Ollama: Qwen3.5 0.8B is a 0.8B-parameter model (approximately 1 GB download). Although the model supports images, this app supplies only Accessibility and OCR text. Thinking is disabled for short background summaries. Output is limited to 160 generated tokens per observation, with a 4K context and bounded input. Raw source text is retained even when no model is available. Existing summaries retain their original model attribution when you change the active model; new summaries use the selected model.
FastAPI dashboard: loopback-only, same-origin controls, host validation, no remote fonts or scripts. Captured strings are inserted as text, never executable HTML.
Official MCP Python SDK: pinned to the supported v1 line (
mcp<2), serving stdio tools/resources.uv.lockrecords the development environment.
Capture limits and privacy
Password managers are excluded by default. Exclusions are checked before accessibility traversal or image capture. Locked or inactive desktop sessions and windows with detected secure text fields are skipped. Secret-pattern redaction happens before persistence and inference.
These are best-effort protections: some apps expose little accessibility text, may not label sensitive fields correctly, and OCR may still read sensitive information. Incognito/private windows are not automatically detected. Exclude apps whose contents you do not want stored. The app does not send screen content to cloud services; only the initial model/dependency downloads require the network. Local files are permission-restricted, but not separately encrypted. The loopback API is available to other processes on this user account and is not intended for shared-machine or remote deployment.
A 0.8B model can hallucinate or produce weak summaries, especially on dense screens or malicious instructions in captured text. The UI exposes source text for checking. Large windows and accessibility trees are bounded/truncated to limit load. Capture interval is measured after capture work, so slow OCR can lengthen the effective interval. The dashboard avoids re-ingesting its own journal when its title is visible in the foreground window.
Development and verification
uv sync --python 3.12 --extra dev
bash scripts/build-app.sh
.venv/bin/pytest -qTests exercise real SQLite/Markdown persistence, full-text index updates, retention, redaction, chronological session grouping, app-return transitions, consolidated session summaries, pause races, onboarding persistence, separate permission requests, native bridge routing and authentication, cancellation/timeout cleanup, local API access controls, single-daemon locking, and an actual stdio MCP client/server exchange. Screen capture needs a logged-in macOS desktop and its permissions, so tests mock the OS boundary. Run the native app for live validation; command-line doctor checks the launching terminal's permission context, which may differ from the native app.
Builds use the sole available Apple Development certificate when one exists, keeping the app identity stable across rebuilds. Set CODE_SIGN_IDENTITY to select a certificate explicitly; set it to - for ad-hoc signing. With no unique development certificate, builds fall back to ad-hoc signing. The release packaging script bundles the runtime. Developer ID signing and Apple notarization require your own credentials; the development installer does not perform notarization.
Changing signing identity or rebuilding an ad-hoc-signed binary can cause macOS to require a new permission grant. The native onboarding reports the actual permission status and cannot grant it for you. Build staging happens outside Documents to avoid cloud-file-provider metadata interfering with signing.
Implementation references: Apple ScreenCaptureKit, Ollama chat API, Qwen model specification, MCP Python SDK v1.
Day and week reviews
The native app opens to Review: choose a day or week, inspect source-linked highlights, and prepare an editable Markdown status draft. Counts describe captured context, not work duration. Draft outcomes/priorities need your confirmation; drafts are only copied on request and edits are temporary. MCP read_review(start, end) and GET /api/review?start=YYYY-MM-DD&end=YYYY-MM-DD support inclusive ranges up to 31 days. Counts represent sampled context, not continuous activity duration.
Replay and public website
Use Replay orb intro in the sidebar, application menu, or menu bar to restart the transparent desktop introduction. The shortcut is ⌘⇧R; constantwatch://replay opens the same flow from the website. Replaying preserves capture settings, permissions, journal data, and the preferred name.
The public website source lives in site/, with the hosted version at https://constant-watch.yuggupta.chatgpt.site. It contains fictional demonstration text and download links, never the local journal API or captured history. Production hosting configuration is managed separately.
bash scripts/package-app.sh builds the app and the branded Retina drag-install DMG. bash scripts/package-dmg.sh repackages an already signed app from the build cache. The icon and native wordmark share native/BrandMark.swift, matching the seven-sphere reveal. Installer Finder settings deliberately use supported 128-point icons and 16-point labels.
Windows desktop preview
The Windows edition uses a native WebView2 desktop window, Microsoft UI Automation and Windows OCR, with the same Python journal, Qwen model, and read-only MCP server. Windows 11 x64 is the initial supported target. The Mac SwiftUI app remains available separately.
Run Constant-Watch-0.1.0-Windows-x64-Setup.exe. It installs for the current user without administrator access and adds a Start menu shortcut. Python and service dependencies are bundled. Windows 11 normally includes the required Microsoft Edge WebView2 Runtime. Install and open Ollama through the setup guide, then choose Download local model. Choose Start my journal explicitly to enable capture. Closing the desktop window stops its capture service; minimize it to keep watching.
Windows data lives under %LOCALAPPDATA%\Constant Watch. Exclusions use lowercase executable IDs such as windows:chrome.exe and windows:bitwarden.exe; the app list shows the IDs that were actually captured. Locked desktops, excluded applications, and windows with visible password controls are skipped. UI Automation privacy inspection has a node/time limit; oversized or changing trees are skipped instead of bypassing inspection. OCR operates on the visible foreground-window rectangle in memory and may include overlapping visible windows. Elevated/protected apps may be unreadable. Windows language settings must have an OCR-supported language installed; otherwise accessibility capture continues with an OCR warning.
Copy the MCP configuration from Capture settings or %LOCALAPPDATA%\Constant Watch\mcp-config.json. The Windows command is the installed constant-watch-service.exe with argument mcp. No Mac executable or source checkout is needed. The desktop UI currently provides daily flow, app filtering, source text, search, Markdown export, settings and model setup; the Mac's transparent desktop reveal and native day/week review interface are not ported. Review and recall remain available through the shared API and MCP.
Build on Windows with Python 3.12 and Inno Setup 6:
python -m pip install -e ".[dev]" -r windows/requirements.txt
python -m pytest -q
python -m PyInstaller --noconfirm windows/constant-watch.spec
python scripts/third-party-notices.py "dist/Constant Watch/THIRD-PARTY-NOTICES.txt"
python windows/smoke.py
& "${env:ProgramFiles(x86)}\Inno Setup 6\ISCC.exe" windows/installer.iss.github/workflows/windows.yml runs these checks on Windows. The Windows preview installer is unsigned; Windows SmartScreen may warn. Uninstalling the program preserves the journal in LocalAppData.
License and project
Released under the MIT license. Third-party dependencies retain their own licenses; packaged applications include dependency notices. Manas Vardhan is CEO of Constant Watch.
Available Tools
11 toolsask_memoryB
Answer a natural-language question with exact captured quotes and source citations. Optional YYYY-MM-DD and app/topic filters. Say evidence is missing when no relevant record exists.
| Name | Required | Description | Default |
|---|---|---|---|
| day | No | ||
| topic | No | ||
| app_id | No | ||
| question | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden, and it does add real signal: results come back as exact quotes with source citations, and the tool should report missing evidence rather than fabricate. However, it says nothing about scope limits, permissions, or retrieval boundaries.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action and output form, followed by filters and the missing-evidence rule. Nothing is padded, though the compressed 'app/topic filters' phrasing is slightly terse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema or annotations exist, so the description is the sole source of behavioral facts, and it does describe the return shape (quotes + citations) and the empty-result behavior. It is nearly complete for a four-parameter Q&A tool, lacking only scope/limit details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: it names the day format (YYYY-MM-DD) and identifies app/topic filters, covering all three optional parameters plus the required question. It stops short of explaining filter interaction or default behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (answer) and resource (natural-language question over captured memory) and specifies the output form: exact quotes plus source citations. It does not differentiate itself from the sibling search_screen_memory, which an agent might otherwise confuse this with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (ask a question about your memory) but gives no explicit when-to-use vs. when-not, and never names or contrasts with alternatives like search_screen_memory or read_observation. The 'say evidence is missing' line is an answer-behavior instruction rather than tool-selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
day_sessionsC
Read chronological app/window sessions, with summaries and source evidence. Use next_offset for the next page.
| Name | Required | Description | Default |
|---|---|---|---|
| day | Yes | ||
| limit | No | ||
| offset | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. 'Read' implies a non-mutating operation, and pagination is hinted at, but there is no disclosure of permissions, result volume, rate limits, or what 'summaries and source evidence' actually contain. The pagination hint even references a token (next_offset) that does not appear in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core purpose front-loaded and no filler. It loses a point because the second sentence conveys inaccurate information about pagination rather than earning its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No annotations, no output schema, and 0% schema coverage, so the description must do all the work for a paginated, day-scoped retrieval tool. It leaves parameter formats, pagination mechanics, and result shape unspecified, which is inadequate for this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for all three parameters. The description only obliquely implies the required 'day' input and mentions pagination, but never explains the format of 'day', the meaning of limit, or how offset works, and the pagination token it names is not a real parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Read chronological app/window sessions, with summaries and source evidence.' An agent knows it retrieves session data for a day, but the description never distinguishes this from close siblings like read_app_day, read_day_flow, or recent_activity, so it earns a clear-but-undifferentiated 4.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage guidance is 'Use next_offset for the next page', which is not even a parameter of this tool (the schema exposes limit/offset). No indication of when to prefer this over read_app_day or read_day_flow, and no prerequisites or exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_appsA
List captured applications, bundle IDs, observation counts, and available date ranges.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. 'List' implies a safe read-only operation and the description discloses what data is returned, but it omits scope (all vs filtered), pagination, sorting, auth, and rate limits, making it minimally adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence with the verb first, and each listed attribute earns its place by telling the agent what comes back. No redundant phrasing or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter list tool with no output schema, the description identifies the returned entities and their key fields. It leaves the shape of 'available date ranges' and pagination behavior unspecified, but it is largely complete for correct invocation and interpretation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so the semantic baseline is 4. The description does not need to explain parameters and adds no parameter-level detail, which is appropriate for a zero-argument list call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb 'List' and resource 'captured applications', and enumerates returned attributes (bundle IDs, observation counts, date ranges). It does not name a sibling or clarify when to use it over tools like list_topics or read_app_day, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, prerequisites, or alternatives are given. Siblings such as read_app_day and search_screen_memory are not mentioned, leaving the agent to infer when list_apps is the right choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_topicsC
List suggested project/topic groups across apps, with counts and grouping notice.
| Name | Required | Description | Default |
|---|---|---|---|
| day | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full behavioral burden, yet it only hints at output contents ('counts and grouping notice'). It says nothing about what 'suggested' implies (heuristic grouping? freshness?), whether the listing is scoped or permission-bound, or how the optional day influences results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence with no filler, and the key scope phrase ('across apps') comes early. It could have spent a few of its remaining words on the day parameter or exclusion guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and an undocumented parameter, the description is too thin for an agent to invoke this confidently. The overview of the return shape is a start, but the day semantics and the distinction from sibling topic tools are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single 'day' parameter is never mentioned in the description. The agent cannot tell whether day filters, defaults to today, or accepts a format like ISO date versus a relative label, despite the empty-string default.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'List suggested project/topic groups across apps,' plus the payload shape (counts, grouping notice). It is distinguishable from read_topic (singular detail fetch) and list_apps, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance at all. Nothing tells the agent whether to call this before read_topic, how it relates to list_apps, or what 'suggested' means in contrast to a raw topic list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_app_dayB
Read a complete app-specific daily Markdown journal. day is YYYY-MM-DD.
| Name | Required | Description | Default |
|---|---|---|---|
| day | Yes | ||
| app_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It does disclose the return shape as a 'complete' (i.e., non-truncated) Markdown journal, which hints there is no pagination, but says nothing about missing days, permissions, or error behavior for an invalid app_id.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero filler, and the resource description is front-loaded ahead of the parameter note. Slightly terse given the gaps elsewhere, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, but with no annotations and 0% parameter coverage the description still leaves app_id semantics and sibling routing unaddressed. It is thin for a tool in a crowded read-only namespace.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It usefully specifies the day format (YYYY-MM-DD), but leaves app_id entirely undefined beyond its self-evident name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Read') and a well-scoped resource ('app-specific daily Markdown journal'), which separates it from read_day_flow and day_sessions by its app dimension. It does not, however, name or contrast with any sibling explicitly, so differentiation is inferred rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance and no alternatives mentioned, despite several overlapping read siblings (read_observation, read_review, read_day_flow, day_sessions). Usage is only implied by the word 'app-specific'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_day_flowA
Read one continuous daily Markdown journal across all apps, grouped into chronological sessions. day is YYYY-MM-DD.
| Name | Required | Description | Default |
|---|---|---|---|
| day | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses the output shape (a Markdown journal grouped into chronological sessions), which is real behavioral context. However, it says nothing about permissions, pagination, or limits for what could be a large multi-app aggregate read.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, and the scoping/grouping behavior is front-loaded ahead of the parameter note. Nothing could be removed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. Scope, grouping, and parameter format are covered, making the definition sufficient to call the tool. Minor gaps remain around access requirements and any size limits for an all-app aggregate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does by specifying the required format 'day is YYYY-MM-DD' — the single most important semantic for the one required parameter. It does not address timezone or invalid-date handling, keeping it just short of a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: reading a continuous daily Markdown journal, with the distinguishing scope 'across all apps, grouped into chronological sessions'. This implicitly sets it apart from read_app_day (single app) and day_sessions, but it never names a sibling, so differentiation is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: 'across all apps' signals this is the whole-day, cross-app view rather than a per-app one. There is no explicit when-to-use statement, no when-not-to-use, and no sibling named as the alternative, so the agent must infer the routing decision.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_observationA
Read a cited observation's original Accessibility/OCR text and source URL. Does not open links or take actions.
| Name | Required | Description | Default |
|---|---|---|---|
| observation_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full load. It does disclose the key behavioral trait relevant here — read-only, no link opening, no side effects — which is genuinely useful. It is silent on permissions, error behavior when the observation_id is missing, and any rate limits, so it is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with what is read and the returned payload, followed by the boundary clause. Zero filler, every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read tool this is nearly complete: the return content (OCR/accessibility text plus source URL) is described even without an output schema, and the no-side-effects behavior is stated. The remaining gap is that the agent gets no guidance on sourcing a valid observation_id.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single parameter has no description in the schema. The description partially compensates by framing the id as belonging to a 'cited observation,' hinting at its origin, but gives no format, source, or lookup guidance. Marginal added value over the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Read a cited observation') and even names the payload ('original Accessibility/OCR text and source URL'), so the agent knows exactly what comes back. It does not visibly differentiate itself from sibling read tools like read_review or read_topic, but the resource scope is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'A cited observation's' implies the usage context: follow a citation/observation_id to its source. The closing clause 'Does not open links or take actions' sets a useful boundary against browsers/action tools. However, there is no explicit when-to-use-vs-alternative routing against the many sibling read_* tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_reviewA
Read an evidence-linked review and editable status draft for an inclusive YYYY-MM-DD range (up to 31 days). Counts are observations, not time worked. Projects are suggested groups; no completed actions are inferred.
| Name | Required | Description | Default |
|---|---|---|---|
| end | Yes | ||
| start | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden, and it delivers unusually valuable interpretive warnings: counts are observations not time worked, projects are suggested groups, and no completed actions are inferred. These prevent the agent from over-reading the output, though permission needs and return structure are not addressed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and constraint, and every sentence carries distinct information (scope, range limit, two interpretation caveats). Slightly dense phrasing, but no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a two-parameter read tool with no annotations and no output schema, the description covers the range constraint and output interpretation but says nothing about the shape of the returned review/draft or how it relates to sibling read tools. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% (only bare 'Start'/'End' titles), so the description must compensate, and it does: it defines an inclusive YYYY-MM-DD format and an upper bound of 31 days. That said, all the added meaning sits on the pair jointly rather than distinguishing start from end semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Read an evidence-linked review and editable status draft') and clarifies what the payload is, which separates it from pure-data siblings like read_observation and read_day_flow. It stops short of explicitly naming an alternative sibling or stating when this review view is preferred over raw reads.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The range constraint ('inclusive YYYY-MM-DD range, up to 31 days') gives usable invocation bounds, so usage is reasonably implied. However, nothing says when to choose this review tool versus read_observation, read_app_day, or read_day_flow, and no prerequisites or exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_topicA
Read chronological evidence across apps for an exact topic key returned by list_topics; paginate using next_offset.
| Name | Required | Description | Default |
|---|---|---|---|
| day | No | ||
| key | Yes | ||
| limit | No | ||
| offset | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses ordering (chronological), cross-app scope, and the pagination mechanism via next_offset in the response, but says nothing about permissions, result size, or behavior of the day filter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with purpose and scope, with the prerequisite and pagination note appended compactly. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter tool with no annotations, no output schema, and 0% schema description coverage, one sentence is thin. Purpose and pagination are covered, but the day parameter and any return-shape detail beyond next_offset are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It clarifies that key must be an exact value obtained from list_topics and implies pagination covers offset/limit, but leaves the day parameter entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read chronological evidence across apps') plus the narrow scope (an exact topic key), and ties the input source to the list_topics sibling. It is clear what the tool returns, though it does not contrast itself against other read tools like read_observation or read_app_day.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'for an exact topic key returned by list_topics' gives a clear precondition and prerequisite step, telling the agent to call list_topics first. It stops short of stating when another sibling read tool would be preferred instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recent_activityB
Read latest observations with source text. Pass the last id as before_id to page backward.
| Name | Required | Description | Default |
|---|---|---|---|
| day | No | ||
| limit | No | ||
| app_id | No | ||
| before_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the pagination mechanism (cursor-style backward paging via before_id), which is real behavioral information. However it omits the default limit behavior, what the day/app_id filters do to results, ordering guarantees, and any auth or rate-limit context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero waste, with the core action front-loaded and the paging hint immediately after. Nothing is padded or redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No annotations, no output schema, and four parameters at 0% schema coverage, three of which are unexplained anywhere. For a paged read tool this leaves major gaps (filter semantics, default result size, ordering) that the description should have closed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must document all four parameters, yet it only explains before_id. The meaning and format of day, limit (default 20), and app_id are left entirely to inference, so the majority of parameters remain undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Read') and resource ('latest observations with source text'), and implies a reverse-chronological feed distinct from the single-item siblings like read_observation and read_day_flow. It does not explicitly name a sibling alternative, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied (get a recent-activity feed) and the description never says when to prefer this over read_observation, day_sessions, or search_screen_memory. It does give actionable paging guidance ('pass the last id as before_id to page backward'), which is a genuine usage instruction, preventing a 2.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_screen_memoryB
Search captured text and summaries. All query words must match. Optionally filter by bundle ID and YYYY-MM-DD.
| Name | Required | Description | Default |
|---|---|---|---|
| day | No | ||
| limit | No | ||
| query | Yes | ||
| app_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose one genuinely useful behavioral trait: matching is conjunctive ('All query words must match'). It omits other behavior an agent needs, such as result ordering, the default limit of 20, or whether matching is case-sensitive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, zero filler, and the core action plus the matching rule are front-loaded before the optional filters. Nothing would be lost by trimming it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter search tool with no annotations and no output schema, the definition is serviceable but thin: it never explains the result shape or ordering, and 'limit' is undocumented. An agent can call it, but cannot predict what comes back.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It covers three of four params – query (all words must match), app_id (bundle ID filter), and day (YYYY-MM-DD format) – but leaves 'limit' completely unexplained, so the compensation is partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Search') and resource ('captured text and summaries') and adds a discriminating scope note ('All query words must match') that separates it from the semantic sibling ask_memory. It is clear what it does, though the resource label 'captured text and summaries' is slightly abstract.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description never says when to use this tool versus alternatives such as ask_memory or read_observation, nor when-not to use it. It only notes that filters are optional, which is parameter guidance rather than usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
11 tool updates
v0.1.0- First observed
ask_memory - First observed
day_sessions - First observed
list_apps - First observed
list_topics - First observed
read_app_day - First observed
read_day_flow - First observed
read_observation - First observed
read_review - First observed
read_topic - First observed
recent_activity - First observed
search_screen_memory
TDQS
Scored across 11 tools
Several tools overlap heavily around the same daily-scoped retrieval: read_day_flow, day_sessions, and read_app_day all surface chronological day data, and recent_activity is close to day_sessions. read_review also overlaps with read_day_flow, and search_screen_memory vs ask_memory share retrieval intent, though descriptions do give useful distinguishing cues.
Most tools follow a clear snake_case verb_noun pattern (list_apps, read_observation, read_review, search_screen_memory, ask_memory, list_topics). A few nouns-first names break the pattern (day_sessions, recent_activity), but naming remains readable and largely predictable.
11 tools is well-scoped for a capture-and-recall memory server. Each tool maps to a distinct retrieval entry point (apps, observations, days, sessions, search, Q&A, topics) without obvious padding.
As a read-only observation browser, the surface is fairly complete: discovery (list_apps, list_topics), retrieval (search_screen_memory, read_*), chronological views (day_sessions, recent_activity), and NL synthesis (ask_memory). Minor gaps like export or cross-day aggregation exist but core workflows are covered.
Maintenance
Related MCP Connectors
- KogniteOAuthdev.kognite
Hosted agent memory: store, search, and recall facts across sessions from any MCP client.
Private-by-default, local-first memory/context/task orchestrator for MCP apps and agents.
Personal knowledge graph as an AI memory layer over MCP - read, save, and link your memories.
- docs2mcpOAuthcom.docs2mcp
Query your own PDFs and documents from any MCP client. Every answer cites the page it came from.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables AI tools to query a user's private, locally stored memories (notes, documents) with source citations, using the MCP protocol.34 npmMIT
- AlicenseNot gradedqualityCmaintenanceProvides a local-first, source-cited memory layer for AI agents, with MCP tools to search, read, explain sources, and propose/apply memory updates.29 npm10Apache 2.0
- AlicenseNot gradedqualityAmaintenanceExposes a local SQLite-based memory and knowledge base as standard MCP tools, enabling AI clients to search content, recall memory fragments, and query entity relationships through natural language.1MIT
- AlicenseNot gradedqualityAmaintenanceEnables LLMs to search, retrieve, and store local memory captures, manage reminders, and access memory statistics via MCP.15 npmMIT