Skip to main content
Glama
michaelneale

Goose App Maker MCP

by michaelneale

구스 앱 메이커

이 MCP(Model Context Protocol) 서버를 사용하면 사용자가 Goose를 통해 웹 애플리케이션을 만들고, 관리하고, 제공할 수 있으며, API 호출, 데이터 액세스 등에 Goose를 활용할 수 있습니다.

설치하다

여기를 클릭하여 구스에 설치하세요 🪿 🪿🪿🪿🪿🪿🪿🪿🪿🪿🪿

특징

  • 기본 지침으로 새로운 웹 애플리케이션 만들기

  • ~/.config/goose/app-maker-apps 디렉토리에 앱을 저장합니다(각 앱은 자체 하위 디렉토리에 있음)

  • 로컬에서 필요에 따라 웹 애플리케이션 제공

  • 기본 브라우저에서 웹 애플리케이션을 엽니다(가능하면 크롬리스로 실행).

  • 사용 가능한 모든 웹 애플리케이션을 나열합니다

  • Goose를 일반 백엔드로 사용할 수 있는 앱을 만듭니다.

Related MCP server: mcp-toolkit

예시

goose를 통해 데이터 로드(확장 기능 재사용)

스크린샷 2025-04-28 오후 7시 38분 24초

스크린샷 2025-04-28 오후 7시 38분 53초

구스는 귀하의 앱을 추적합니다:

주문형 앱 만들기

풍부한 표 또는 목록 데이터도 표시합니다.

소스에서의 사용법

예를 들어 거위의 경우:

지엑스피1

중요: 이 MCP는 현재 goose 데스크톱 앱에서 실행해야 합니다(goose-server/goosed에 액세스하기 때문).

건축 및 출판

선택 사항: UV를 사용하여 깨끗한 환경에서 빌드

uv venv .venv
source .venv/bin/activate
uv pip install build
python -m build

출판

  1. pyproject.toml 에서 버전을 업데이트합니다:

[project]
version = "x.y.z"  # Update this
  1. 패키지를 빌드하세요:

# Clean previous builds
rm -rf dist/*
python -m build
  1. PyPI에 게시:

# Install twine if needed
uv pip install twine

# Upload to PyPI
python -m twine upload dist/*

작동 원리

이 MCP는 앱을 제공할 뿐만 아니라 앱이 goosed와 자체 세션을 통해 goose와 통신할 수 있도록 허용합니다.

개요

이 시스템은 웹 애플리케이션이 Goose에 요청을 보내고 메인 스레드를 차단하지 않고 응답을 받을 수 있도록 비차단 비동기 요청-응답 패턴을 구현합니다. 이는 다음 사항들의 조합을 통해 구현됩니다.

  1. 서버 측의 차단 엔드포인트

  2. 클라이언트 측의 비동기 JavaScript

  3. 스레드 동기화를 통한 응답 저장 메커니즘

웹 앱 구조

웹 앱은 리소스/템플릿을 기반으로 요청에 따라 제작(또는 다운로드)됩니다. 각 웹 앱은 ~/.config/goose/app-maker-apps 아래의 자체 디렉터리에 다음과 같은 구조로 저장됩니다.

app-name/
├── goose-app-manifest.json     # App metadata
├── index.html        # Main HTML file
├── style.css         # CSS styles
├── script.js         # JavaScript code
└── goose_api.js      # allows the apps to access goose(d) for deeper backend functionality
└── ...               # Other app files

goose-app-manifest.json 파일에는 다음을 포함하여 앱에 대한 메타데이터가 포함되어 있습니다.

  • name: 앱의 표시 이름

  • 유형: 앱 유형(예: "정적", "반응형" 등)

  • 설명: 앱에 대한 간략한 설명

  • created: 앱이 생성된 타임스탬프

  • 파일: 앱의 파일 목록

1. 클라이언트 측 요청 흐름

클라이언트가 Goose로부터 응답을 받고 싶어할 때:

┌─────────┐     ┌─────────────────┐     ┌───────────────┐     ┌──────────────┐
│ User    │────▶│ gooseRequestX() │────▶│ Goose API     │────▶│ waitForResp- │
│ Request │     │ (text/list/     │     │ (/reply)      │     │ onse endpoint │
└─────────┘     │  table)         │     └───────────────┘     └──────────────┘
                └─────────────────┘                                   │
                         ▲                                            │
                         │                                            │
                         └────────────────────────────────────────────┘
                                          Response
  1. 사용자가 요청을 시작합니다(예: "목록 응답 가져오기" 클릭).

  2. 클라이언트는 요청 함수 중 하나( gooseRequestText , gooseRequestList 또는 gooseRequestTable )를 호출합니다.

  3. 이 함수는 고유한 responseId 생성하고 이 ID로 app_response 호출하라는 지침과 함께 Goose에 요청을 보냅니다.

  4. 그런 다음 함수는 /wait_for_response/{responseId} 엔드포인트를 폴링하는 waitForResponse(responseId) 호출합니다.

  5. 이 엔드포인트는 응답이 가능하거나 시간 초과가 발생할 때까지 차단됩니다.

  6. 응답이 가능하면 클라이언트로 반환되어 표시됩니다.

2. 서버 측 처리

서버 측에서:

┌─────────────┐     ┌─────────────┐     ┌───────────────┐
│ HTTP Server │────▶│ app_response│────▶│ response_locks│
│ (blocking   │     │ (stores     │     │ (notifies     │
│  endpoint)  │◀────│  response)  │◀────│  waiters)     │
└─────────────┘     └─────────────┘     └───────────────┘
  1. /wait_for_response/{responseId} 엔드포인트는 응답이 가능할 때까지 차단하기 위해 조건 변수를 사용합니다.

  2. Goose가 요청을 처리할 때 응답 데이터와 responseId 사용하여 app_response 함수를 호출합니다.

  3. app_response 함수는 app_responses 사전에 응답을 저장하고 조건 변수를 사용하여 대기 중인 모든 스레드에 알립니다.

  4. 차단된 HTTP 요청은 차단 해제되고 클라이언트에게 응답이 반환됩니다.

3. 스레드 동기화

시스템은 Python의 threading.Condition 사용합니다. 스레드 동기화 조건:

  1. 클라이언트가 아직 사용할 수 없는 응답을 요청하면 해당 responseId 에 대한 조건 변수가 생성됩니다.

  2. HTTP 핸들러 스레드는 이 조건을 시간 초과(30초)로 기다립니다.

  3. 응답이 가능해지면 조건이 알려집니다.

  4. 응답이 제공되기 전에 시간 초과가 만료되면 오류가 반환됩니다.

주요 구성 요소

클라이언트 측 함수

  • gooseRequestText(query) : 텍스트 응답을 요청합니다.

  • gooseRequestList(query) : 목록 응답을 요청합니다.

  • gooseRequestTable(query, columns) : 지정된 열이 포함된 테이블 응답을 요청합니다.

  • waitForResponse(responseId) : 주어진 ID로 응답을 기다립니다.

서버 측 함수

  • app_response(response_id, string_data, list_data, table_data) : 응답을 저장하고 대기자에게 알립니다.

  • /wait_for_response/{responseId} 엔드포인트가 있는 HTTP 핸들러: 응답이 가능할 때까지 차단됨

Available Tools

9 tools
app_createB
Create a new web application directory and copy starter files.
The starter files are for you to replace with actual content, you don't have to use them as is.
the goose_api.js file is a utility you will want to keep in case you need to do api calls as part of your app via goose.

Args:
    app_name: Name of the application (will be used as directory name)
    description: Brief description of the application (default: "")

Returns:
    A dictionary containing the result of the operation

After this, consider how you want to change the app to meet the functionality, look at the examples in resources dir if you like.
Or, you can replace the content with existing html/css/js files you have (just make sure to leave the goose_api.js file in the app dir)

Use the app_error tool once it is opened and user has interacted (or has started) to check for errors you can correct the first time, this is important to know it works.
ParametersJSON Schema
NameRequiredDescriptionDefault
app_nameYes
descriptionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden. It mentions creating directories and copying files, but doesn't disclose critical behavioral traits like whether this requires specific permissions, if it overwrites existing directories, what happens on failure, or rate limits. The description adds some context about starter files but misses key operational details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is verbose and poorly structured, mixing operational instructions with usage advice. Sentences like 'After this, consider how you want to change the app...' don't belong in a tool description. It's front-loaded but then diverges into tangential guidance, reducing clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 2 parameters with 0% schema coverage and an output schema exists, the description adds some parameter semantics but lacks completeness. It doesn't explain the return dictionary structure or error conditions, and with no annotations, it should provide more behavioral context for a creation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains app_name will be used as the directory name and description is a brief description with a default, adding meaningful semantics beyond the bare schema. However, it doesn't cover constraints like character limits or validation rules.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool creates a new web application directory and copies starter files, specifying the verb (create) and resource (web application directory). It distinguishes from siblings like app_delete or app_list by focusing on creation, though it doesn't explicitly contrast with app_open or app_serve.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for creating new apps and mentions using app_error after creation, but doesn't explicitly state when to use this tool versus alternatives like app_open for existing apps or app_list for viewing. It provides some context but lacks clear when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

app_deleteC
Delete an existing web application.

Args:
    app_name: Name of the application to delete

Returns:
    A dictionary containing the result of the operation
ParametersJSON Schema
NameRequiredDescriptionDefault
app_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool deletes an app, implying a destructive mutation, but doesn't describe what 'delete' entails (e.g., irreversible removal, cleanup of resources, permissions required, or error handling). This leaves gaps in understanding the tool's behavior beyond the basic action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded: the first sentence states the purpose clearly, followed by structured sections for Args and Returns. There's no wasted text, and the structure aids readability, though the Returns section is vague ('A dictionary containing the result of the operation').

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (destructive operation with 1 parameter) and context (no annotations, schema coverage 0%, but has output schema), the description is minimally adequate. It covers the basic action and parameter, but lacks details on behavior, usage, and output specifics, even though the output schema might document return values. For a deletion tool, more guidance is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds minimal semantics: it names the parameter ('app_name') and indicates it's 'Name of the application to delete.' With 0% schema description coverage, the schema provides no details, so the description compensates slightly by clarifying the parameter's role. However, it doesn't explain format, constraints, or examples, leaving the parameter only partially documented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Delete an existing web application.' It specifies the verb ('Delete') and resource ('web application'), making the action unambiguous. However, it doesn't differentiate from sibling tools like 'app_stop_server' or 'app_error', which might also involve application management, so it doesn't fully distinguish from alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., the app must exist or be stopped), exclusions (e.g., don't use if the app is running), or refer to sibling tools like 'app_list' to check before deletion. Usage is implied by the purpose but lacks explicit context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

app_errorA
Report an error from the app or retrieve the list of errors.
This is useful while developing or debugging the app as it allows errors (or any messages) to be reported and monitored

Args:
    error_message: Optional error message to report. If None, returns the list of errors.
    clear: Optional, If True, clears the list of errors

Returns:
    A string containing the list of errors if error_message is None,
    otherwise a confirmation message.
ParametersJSON Schema
NameRequiredDescriptionDefault
error_messageNo
clearNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses the tool's dual behavior (reporting vs. retrieving) and mentions clearing errors, which adds context. However, it doesn't cover aspects like permissions needed, rate limits, or whether operations are persistent, leaving gaps in behavioral understanding.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded, with the core purpose stated first. The Args and Returns sections are structured but could be more integrated; however, every sentence adds value without redundancy, making it efficient overall.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (dual functionality, 2 parameters), no annotations, and an output schema present, the description is fairly complete. It explains parameters, return values, and usage context, though it could benefit from more behavioral details like error handling or persistence.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds significant meaning beyond the input schema, which has 0% coverage. It explains that 'error_message' is optional and determines the tool's mode (report if provided, retrieve if None), and that 'clear' is optional and clears the list if True. This fully compensates for the lack of schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the dual purpose: 'Report an error from the app or retrieve the list of errors.' It specifies the verb (report/retrieve) and resource (errors), though it doesn't explicitly differentiate from sibling tools like app_response or app_list, which might handle other app-related operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides implied usage context: 'useful while developing or debugging the app.' However, it lacks explicit guidance on when to use this tool versus alternatives (e.g., app_response for general responses or app_list for listing other items) and does not specify prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

app_listB
List all available web applications.

Returns:
    A dictionary containing the list of available apps and their details
ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full burden. It mentions the return format ('dictionary containing the list of available apps and their details'), which adds useful context beyond basic purpose. However, it doesn't disclose behavioral traits like whether this is a read-only operation, potential rate limits, or authentication requirements, leaving gaps in transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized with two sentences: one stating the purpose and one describing the return value. It's front-loaded with the core functionality and avoids unnecessary details, though it could be slightly more structured by explicitly labeling sections.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (0 parameters, no annotations, but with an output schema), the description is reasonably complete. It covers purpose and return format, and the output schema handles return values, so no major gaps exist. However, it lacks behavioral context like safety or performance considerations, preventing a perfect score.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has 0 parameters with 100% schema description coverage, so the schema fully documents the absence of inputs. The description doesn't need to add parameter semantics, and it appropriately avoids discussing parameters, earning a baseline score of 4 for this context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('List') and resource ('web applications'), making it immediately understandable. However, it doesn't differentiate from sibling tools like 'app_refresh' or 'app_serve' which might also involve listing operations, so it doesn't reach the highest score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. With siblings like 'app_refresh' and 'app_serve' that might overlap in functionality, there's no explicit or implied context for choosing this specific listing tool, leaving usage unclear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

app_openA
Open an app in the default web browser. If the app is not currently being served,
it will be served first.
Can only open one app at a time.

Args:
    app_name: Name of the application to open

Returns:
    A dictionary containing the result of the operation
ParametersJSON Schema
NameRequiredDescriptionDefault
app_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses key behavioral traits: it opens in a web browser, serves the app if not already served, and enforces a single-app-at-a-time constraint. However, it lacks details on permissions, error handling, or what 'served first' entails (e.g., time, resources). This is adequate but has gaps for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded: the first sentence states the core action, followed by important behavioral notes and parameter/return details. Every sentence adds value without redundancy, and the structure with 'Args:' and 'Returns:' sections enhances clarity efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (mutation with constraints), no annotations, and an output schema present (so return values are documented elsewhere), the description is fairly complete. It covers purpose, key behavior, and parameter semantics, but could improve by addressing error cases or interaction with sibling tools like app_serve. The presence of an output schema reduces the need for return value details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 0%, so the description must compensate. It adds meaning by explaining 'app_name' as 'Name of the application to open,' which clarifies the parameter's role beyond the schema's basic type. Since there's only one parameter, this is sufficient to understand its use, though it doesn't specify format or constraints like valid app names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb ('Open') and resource ('an app in the default web browser'), making the purpose specific and understandable. It distinguishes from siblings like app_serve (which only serves) and app_list (which lists), though it doesn't explicitly name alternatives. The mention of serving if needed adds useful context but doesn't fully differentiate from app_serve in a comparative way.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by stating 'Can only open one app at a time,' which provides some context on limitations, but it doesn't explicitly say when to use this tool versus alternatives like app_serve or app_list. No guidance on prerequisites or exclusions is given, leaving the agent to infer based on the description alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

app_refreshA
Refresh the currently open app in Chrome.
Only works on macOS with Google Chrome.

Returns:
    A dictionary containing the result of the operation
ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden. It discloses platform/software constraints (macOS/Chrome) and mentions the return format, but doesn't cover other behavioral aspects like error handling, permissions needed, or what 'refresh' entails operationally (e.g., reloads page, clears cache).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized with three concise sentences that each add value: stating the action, specifying constraints, and describing the return. It's front-loaded with the core purpose and wastes no words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 0 parameters, 100% schema coverage, and an output schema (which handles return values), the description is reasonably complete. It covers the tool's purpose, operational constraints, and mentions the return format. However, for a tool with no annotations, it could benefit from more behavioral context about what 'refresh' actually does.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description appropriately doesn't discuss parameters, maintaining focus on the tool's purpose and constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Refresh') and target ('the currently open app in Chrome'), providing a specific verb+resource combination. However, it doesn't differentiate from sibling tools like 'app_open' or 'app_stop_server' in terms of when to use one versus the other.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context about when this tool works ('Only works on macOS with Google Chrome'), which is helpful for usage decisions. However, it doesn't explicitly state when to use this versus alternatives like 'app_open' or 'app_stop_server' from the sibling list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

app_responseA
Use this to return a response to the app that has been requested.
Provide only one of string_data, list_data, or table_data.

Args:
    string_data: Optional string response
    list_data: Optional list of strings response
    table_data: Optional table response with columns and rows
                Format: {"columns": ["col1", "col2", ...], "rows": [["row1col1", "row1col2", ...], ...]}

Returns:
    True if the response was stored successfully, False otherwise
ParametersJSON Schema
NameRequiredDescriptionDefault
string_dataNo
list_dataNo
table_dataNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It describes the tool's behavior by specifying the return value ('True if the response was stored successfully, False otherwise') and the parameter constraints, but it lacks details on permissions, rate limits, or error handling. This is adequate for a basic tool but misses deeper behavioral traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded, starting with the purpose and key usage rule. Sentences are efficient, with no wasted words. However, the structure could be slightly improved by separating the 'Args' and 'Returns' sections more clearly, but overall it's concise and well-organized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (3 parameters, nested objects, no annotations, but has an output schema), the description is fairly complete. It covers purpose, parameter semantics, and return values, and the output schema handles return details, so no need to explain those further. It could benefit from more behavioral context, but it's sufficient for basic use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds significant meaning beyond the input schema, which has 0% description coverage. It explains the semantics of each parameter (string_data, list_data, table_data), provides a format example for table_data, and clarifies that only one should be provided. This fully compensates for the schema's lack of descriptions, making parameters clear and actionable.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'to return a response to the app that has been requested.' It specifies the verb ('return a response') and resource ('to the app'), making it understandable. However, it doesn't explicitly differentiate from siblings like app_error or app_list, which might also involve app interactions, so it misses full sibling distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides implied usage guidance: 'Provide only one of string_data, list_data, or table_data,' indicating mutual exclusivity among parameters. It doesn't explicitly state when to use this tool versus alternatives like app_error for errors or app_list for listing, nor does it mention prerequisites or exclusions, leaving usage context somewhat vague.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

app_serveA
Serve an existing web application on a local HTTP server.
The server will automatically find an available port.

Can only serve one app at a time

Args:
    app_name: Name of the application to serve

Returns:
    A dictionary containing the result of the operation
ParametersJSON Schema
NameRequiredDescriptionDefault
app_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses that the server 'will automatically find an available port' and 'can only serve one app at a time,' which are useful behavioral traits. However, it doesn't mention permissions, rate limits, or what happens if the app is already running, leaving gaps for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded with the main purpose. The sentences are efficient, though the 'Args' and 'Returns' sections could be integrated more smoothly into the flow, but overall it's concise with minimal waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema (returns 'A dictionary containing the result of the operation'), the description doesn't need to detail return values. However, as a mutation tool with no annotations and only basic behavioral context, it lacks information on error handling or prerequisites, making it adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds meaning beyond the input schema by explaining that 'app_name' is the 'Name of the application to serve.' Since there is only one parameter and schema description coverage is 0%, this compensates well, though it doesn't detail format or constraints like valid app names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Serve an existing web application on a local HTTP server.' It specifies the verb ('serve') and resource ('existing web application'), but doesn't explicitly differentiate from siblings like 'app_open' or 'app_stop_server' that might involve similar concepts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides some implied usage context with 'Can only serve one app at a time,' which suggests a constraint but doesn't explicitly state when to use this tool versus alternatives like 'app_open' or 'app_stop_server.' No clear alternatives or exclusions are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

app_stop_serverB
Stop the currently running HTTP server.

Returns:
    A dictionary containing the result of the operation
ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool stops a server, implying a destructive mutation, but doesn't clarify permissions needed, whether the action is reversible, side effects (e.g., interrupting active connections), or error handling. The mention of a return dictionary adds minimal value without details on its structure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is brief (two sentences) and front-loaded with the core action, but the second sentence about returns is vague and adds little value without specifics. It could be more structured by integrating return details more meaningfully or omitted if covered by an output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (a destructive operation with no annotations), the description is minimally adequate. It states what the tool does but lacks critical behavioral context. The presence of an output schema mitigates the need to explain return values, but gaps remain in usage guidelines and transparency for a mutation tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0 parameters with 100% coverage, so no parameter documentation is needed. The description appropriately doesn't discuss parameters, focusing on the tool's action. A baseline of 4 is applied since it avoids unnecessary repetition of schema information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Stop') and target ('currently running HTTP server'), providing a specific verb+resource combination. However, it doesn't differentiate this tool from its siblings like 'app_delete' or 'app_error', which might also affect server state, so it doesn't reach the highest score.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., a server must be running), exclusions, or how it relates to sibling tools like 'app_serve' (which likely starts a server). This leaves the agent with minimal context for decision-making.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 9 tool updates
    • First observedapp_create
    • First observedapp_delete
    • First observedapp_error
    • First observedapp_list
    • First observedapp_open
    • First observedapp_refresh
    • First observedapp_response
    • First observedapp_serve
    • First observedapp_stop_server

TDQS

A3.7/5.0

Scored across 9 tools

Disambiguation4/5

Most tools have distinct purposes, but there is some potential confusion between app_open and app_serve, as both involve launching an app, and app_error's dual functionality for reporting and retrieving errors could be ambiguous. However, the descriptions clarify their specific roles, making them generally distinguishable.

Naming Consistency5/5

All tool names follow a consistent 'app_' prefix with a descriptive action in snake_case, such as app_create, app_delete, and app_list. This pattern is maintained throughout all nine tools, making them predictable and easy to understand.

Tool Count5/5

With 9 tools, the count is well-scoped for managing web applications, covering creation, deletion, listing, serving, opening, error handling, refreshing, responding, and stopping the server. Each tool serves a clear purpose without redundancy, fitting the domain appropriately.

Completeness4/5

The toolset provides comprehensive coverage for basic app lifecycle management, including CRUD operations and runtime interactions. A minor gap exists in the lack of an update tool for modifying existing apps, but agents can work around this by recreating apps or using other tools like app_response for content changes.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers