Skip to main content
Glama

스틸 MCP 서버

대장간 배지

https://github.com/user-attachments/assets/25848033-40ea-4fa4-96f9-83b6153a0212

Claude와 같은 LLM이 Puppeteer 기반 도구와 Steel을 통해 웹을 탐색할 수 있도록 하는 모델 컨텍스트 프로토콜(MCP) 서버입니다. Web Voyager 프레임워크를 기반으로 클릭, 스크롤, 타이핑 등 모든 표준 웹 액션과 스크린샷 촬영을 위한 도구를 제공합니다.

Claude에게 다음과 같은 작업에 대한 도움을 요청하세요.

  • "레시피를 검색하고 재료 목록을 저장하세요"

  • "패키지 배송 상태 추적"

  • "특정 제품의 가격을 찾아 비교하세요"

  • "온라인 구직 신청서를 작성하세요"

🚀 빠른 시작

다음은 Claude Desktop에서 Steel Voyager를 실행하는 간단한 가이드입니다. Steel Cloud와 로컬/자체 호스팅 인스턴스 간에 전환하려면 환경 옵션만 조정하면 됩니다.

필수 조건

  1. 최신 버전의 Git 및 Node.js가 설치되었습니다.

  2. Claude Desktop 설치됨

  3. (선택 사항) 자체 호스팅을 계획하는 경우 로컬에서 실행되는 Steel Docker 이미지

  4. (선택 사항) Steel Cloud를 사용하는 경우 API 키를 가져오세요. 여기에서 키를 받으세요.


A) 퀵스타트(스틸클라우드)

  1. 프로젝트를 복제하고 빌드합니다.

    지엑스피1

  2. 서버 항목을 추가하여 Claude Desktop( ~/Library/Application Support/Claude/claude_desktop_config.json )을 구성합니다.

    {
      "mcpServers": {
        "steel-puppeteer": {
          "command": "node",
          "args": ["path/to/steel-voyager/dist/index.js"],
          "env": {
            "STEEL_LOCAL": "false",
            "STEEL_API_KEY": "YOUR_STEEL_API_KEY_HERE",
            "GLOBAL_WAIT_SECONDS": "1"
          }
        }
      }
    }
    • "YOUR_STEEL_API_KEY_HERE"를 유효한 Steel API 키로 바꾸세요.

    • 클라우드 모드에서는 "STEEL_LOCAL"이 "false"로 설정되어 있는지 확인하세요.

  3. Claude Desktop을 시작하세요. 자동으로 MCP 서버가 클라우드 모드로 실행됩니다.

  4. (선택 사항) 대시보드 에서 활성 Steel Browser 세션을 보거나 관리할 수 있습니다.


B) 빠른 시작(로컬/셀프 호스팅 스틸)

  1. 로컬 또는 자체 호스팅 Steel 서비스가 실행 중인지 확인하세요(예: 오픈 소스 Steel Docker 이미지 사용).

  2. 프로젝트를 복제하고 빌드합니다(아직 완료되지 않았다면 위와 동일):

    git clone https://github.com/steel-dev/steel-mcp-server.git
    cd steel-mcp-server
    npm install
    npm run build
  3. 로컬 모드에 대해 Claude Desktop( ~/Library/Application Support/Claude/claude_desktop_config.json )을 구성합니다.

    {
      "mcpServers": {
        "steel-puppeteer": {
          "command": "node",
          "args": ["path/to/steel-voyager/dist/index.js"],
          "env": {
            "STEEL_LOCAL": "true",
            "STEEL_BASE_URL": "http://localhost:3000",
            "GLOBAL_WAIT_SECONDS": "1"
          }
        }
      }
    }
    • "STEEL_LOCAL"은 "true"여야 합니다.

    • 클라우드 서버에서 셀프 호스팅하는 경우 "STEEL_BASE_URL"을 로컬/셀프 호스팅 Steel URL을 가리키도록 구성하세요.

  4. Claude Desktop을 시작하면 로컬에서 실행 중인 Steel에 연결되고 Steel Voyager가 로컬 모드로 실행됩니다.

  5. (선택 사항) 로컬에서 세션을 보려면 자체 호스팅 대시보드( localhost:5173 )를 방문하거나 Steel 런타임 환경에 대한 로그를 확인하세요.


이제 끝입니다! Claude Desktop이 시작되면 백그라운드에서 MCP 서버를 조정하고 Steel Voyager를 통해 웹 자동화 기능과 상호 작용할 수 있습니다.

설정에 대한 자세한 내용을 보거나 문제가 있는 경우 MCP 설정 문서를 확인하세요: https://modelcontextprotocol.io/quickstart/user

Related MCP server: Puppeteer MCP Server

구성 요소

도구

  • 탐색하다

    • 브라우저에서 모든 URL로 이동합니다.

    • 입력:

  • 찾다

  • 딸깍 하는 소리

    • 번호가 매겨진 레이블을 사용하여 페이지의 요소를 클릭합니다.

    • 입력:

      • label (숫자, 필수): 클릭할 요소의 레이블 번호입니다.

  • 유형

    • 번호가 매겨진 레이블을 사용하여 입력 필드에 텍스트를 입력합니다.

    • 입력:

      • label (숫자, 필수): 입력 필드의 레이블 번호입니다.

      • text (문자열, 필수): 필드에 입력할 텍스트입니다.

      • replaceText (부울, 선택 사항): true인 경우 필드에 있는 기존 텍스트를 모두 바꿉니다.

  • 아래로 스크롤

    • 페이지를 아래로 스크롤하세요

    • 입력:

      • pixels (정수, 선택 사항): 아래로 스크롤할 픽셀 수입니다. 지정하지 않으면 한 페이지 전체가 스크롤됩니다.

  • 위로 스크롤

    • 페이지를 위로 스크롤하세요

    • 입력:

      • pixels (정수, 선택 사항): 위로 스크롤할 픽셀 수입니다. 지정하지 않으면 한 페이지 전체가 스크롤됩니다.

  • 돌아가다

    • 브라우저 기록에서 이전 페이지로 이동합니다.

    • 입력이 필요하지 않습니다

  • 기다리다

    • 최대 10초 동안 기다리세요. 느리게 로드되는 페이지나 동적 콘텐츠가 나타나는 데 더 많은 시간이 필요한 페이지에 유용합니다.

    • 입력:

      • seconds (숫자, 필수): 대기할 시간(초)(0~10).

  • 저장_표시되지_않은_스크린샷

    • 테두리 상자나 강조 표시 없이 현재 페이지를 캡처하여 리소스로 저장합니다.

    • 입력:

      • resourceName (문자열, 선택 사항): 스크린샷을 저장할 이름입니다(예: "before_login"). 생략하면 일반 이름이 자동으로 생성됩니다.

자원

  • 스크린샷 : 저장된 각 스크린샷은 다음 형식의 MCP 리소스 URI를 통해 액세스할 수 있습니다. • screenshot://RESOURCE_NAME

    서버는 "save_unmarked_screenshot" 도구를 지정하거나 (대부분의 도구에서) 주석이 달린 스크린샷으로 작업이 종료될 때마다 이러한 스크린샷을 저장합니다. 이러한 이미지는 표준 MCP 리소스 검색 요청을 통해 가져올 수 있습니다.

(참고: 콘솔 로그는 분석 및 디버깅을 위해 계속 수집되지만, 이 구현에서는 검색 가능한 리소스로 노출되지 않습니다. 서버 로그에 나타나지만 MCP 리소스 URI를 통해 제공되지는 않습니다.)

주요 특징

  • Puppeteer를 사용한 브라우저 자동화

  • 브라우저 세션 관리를 위한 Steel 통합

  • 번호가 매겨진 레이블을 통한 시각적 요소 식별

  • 스크린샷 기능

  • 기본 웹 상호작용(탐색, 클릭, 양식 작성)

  • 스크롤을 통한 레이지 로딩 지원

  • 로컬 및 원격 Steel 인스턴스 지원

경계 상자 이해

Steel Puppeteer는 페이지와 상호 작용할 때 상호 작용 요소를 식별하는 데 도움이 되는 시각적 오버레이를 추가합니다.

  • 각 대화형 요소(버튼, 링크, 입력)에는 고유한 번호가 매겨진 레이블이 지정됩니다.

  • 색상 상자는 요소의 경계를 나타냅니다.

  • 레이블은 쉽게 참조할 수 있도록 요소 위나 내부에 표시됩니다.

  • 클릭 또는 유형 작업에 대한 요소를 지정할 때 이 숫자를 사용하세요.

구성

Steel Voyager는 "로컬" 또는 "클라우드" 두 가지 모드로 실행될 수 있습니다. 이 동작은 환경 변수에 의해 제어됩니다. 간략한 개요는 다음과 같습니다.

환경 변수

기본

설명

철강_지역

"거짓"

Steel Voyager가 로컬(true) 모드에서 실행되는지, 클라우드(false) 모드에서 실행되는지 결정합니다.

스틸 API 키

(없음)

STEEL_LOCAL = "false"인 경우에만 필요합니다. Steel 엔드포인트를 통한 요청을 인증하는 데 사용됩니다.

강철_기지_URL

" https://api.steel.dev "

Steel API의 기본 URL입니다. Steel 서버를 로컬 또는 자체 클라우드 환경에서 직접 호스팅하는 경우 이 값을 재정의합니다. STEEL_LOCAL = "true"이고 STEEL_BASE_URL이 설정되지 않은 경우 기본값은 " http://localhost:3000 "입니다.

글로벌 대기 시간(초)

(없음)

선택 사항. 각 도구 작업 후 대기할 시간(초)입니다(예: 로딩 속도가 느린 페이지를 허용하는 경우).

로컬 모드

  1. STEEL_LOCAL="true"로 설정합니다.

  2. (선택 사항) 사용자 지정 도메인에서 Steel 서버를 호스팅하는 경우 STEEL_BASE_URL을 Steel 서버로 설정하세요. 그렇지 않으면 Steel Voyager는 기본적으로 http://localhost:3000 으로 접속합니다.

  3. 이 모드에서는 API 키가 필요하지 않습니다.

  4. Puppeteer는 ws://0.0.0.0:3000을 통해 연결합니다.

예:

STEEL_LOCAL="true"로 내보내기

export STEEL_BASE_URL=" http://localhost:3000 " # 재정의하는 경우에만

클라우드 모드

  1. STEEL_LOCAL="false"로 설정합니다.

  2. Steel Voyager가 Steel 클라우드 서비스(또는 STEEL_BASE_URL을 변경한 경우 자체 호스팅 Steel)에 인증할 수 있도록 STEEL_API_KEY를 설정합니다.

  3. STEEL_BASE_URL의 기본값은 https://api.steel.dev 입니다. 다른 엔드포인트에서 실행되는 자체 호스팅 Steel 인스턴스가 있는 경우 이를 재정의하세요.

  4. Puppeteer는 wss://connect.steel.dev?sessionId=…&apiKey=…를 통해 연결합니다.

예:

STEEL_LOCAL="false" 내보내기

STEEL_API_KEY="YOUR_STEEL_API_KEY_HERE"를 내보내세요

클로드 데스크톱 구성

Claude Desktop과 함께 Steel Voyager를 사용하려면 다음과 같은 내용을 구성 파일(일반적으로 ~/Library/Application Support/Claude/claude_desktop_config.json에 있음)에 추가하세요.

{
  "mcpServers": {
    "steel-puppeteer": {
      "command": "node",
      "args": ["path/to/steel-puppeteer/dist/index.js"],
      "env": {
        "STEEL_LOCAL": "false",
        "STEEL_API_KEY": "your_api_key_here"
      }
    }
  }
}

원하는 모드에 맞게 환경 변수를 조정하세요.

• 로컬/자체 호스팅으로 실행하는 경우 "STEEL_LOCAL": "true" 유지하고 선택적으로 "STEEL_BASE_URL": "http://localhost:3000" 유지합니다.
• 클라우드 모드에서 실행하는 경우 "STEEL_LOCAL": "true" 제거하고 "STEEL_LOCAL": "false" 추가하고 "STEEL_API_KEY": "<YourKey>" 제공합니다. 이렇게 하면 Claude Desktop이 올바른 모드에서 Steel Voyager를 시작할 수 있습니다.

설치 및 실행

Smithery를 통해 설치

Smithery 를 통해 Claude Desktop용 Steel MCP Server를 자동으로 설치하려면:

npx -y @smithery/cli install @steel-dev/steel-mcp-server --client claude

지역 개발

  1. 저장소를 복제합니다

  2. 종속성 설치:

    npm install
  3. 프로젝트를 빌드하세요:

    npm run build
  4. 서버를 시작합니다:

    npm start

사용 예시 📹

우리는 Claude에게 새로운 기능으로 우리에게 인상을 심어달라고 요청했고, Claude는 Sora를 사용하여 최신 개발 사항을 조사한 다음 모델의 데이터와 작동 방식을 보여주는 대화형 시각화를 만들기로 결정했습니다. 🤯

https://github.com/user-attachments/assets/8d4293ea-03fc-459f-ba6b-291f5b017ad7

*화질이 좋지 않아서 죄송합니다. github에서는 영상 크기를 10MB 이하로 유지하도록 강제하고 있습니다 :/

문제 해결

일반적인 문제 및 해결 방법:

  1. 클라우드 서비스를 사용할 때 Steel API 키를 확인하고 로컬 Steel 인스턴스가 실행 중인지 확인하세요. 서비스에 대한 네트워크 연결이 원활한지 확인하세요.

  2. 페이지가 렌더링되거나 마크업되어 Claude로 전송되는 방식에 문제가 있는 경우 GLOBAL_WAIT_SECONDS 환경 변수를 통해 구성에 지연을 추가해보세요.

  3. 페이지가 완전히 로드되었는지 확인하고 뷰포트 크기 설정을 확인하세요. 스크린샷을 캡처하기에 충분한 시스템 메모리가 있는지 확인하세요.

  4. 현재 세션 정리가 최선이 아니므로 작업을 실행하기 위해 세션이 시작되면 수동으로 세션을 해제해야 할 수도 있습니다.

  5. 클로드에게 올바른 방식으로 촉구하는 것은 성과를 크게 향상시키고, 이로 인해 발생할 수 있는 어리석은 실수를 피하는 데 도움이 될 수 있습니다.

  6. 세션 뷰어를 활용하여 모델이 어디에서 중단되는지 분석하세요.

  7. 브라우저에서 15~20회 정도 작업을 수행한 후, Claude의 컨텍스트 창이 이미지로 가득 차면서 속도가 느려지기 시작합니다. 심각한 문제는 아니지만, 특히 Claude 데스크톱 클라이언트가 지연되는 등 약간의 지연 현상이 나타났습니다.

기여하다

이 프로젝트는 실험적이며 활발하게 개발 중입니다. 여러분의 참여를 환영합니다!

  1. 저장소를 포크하세요

  2. 기능 브랜치 생성

  3. 풀 리퀘스트 제출

다음을 포함하세요.

  • 변경 사항에 대한 명확한 설명

  • 동기 부여

  • 문서 업데이트

부인 성명

⚠️ 이 프로젝트는 실험적이며 Web Voyager 코드베이스를 기반으로 합니다. 프로덕션 환경에서의 사용은 사용자의 책임입니다.

Available Tools

16 tools
steel_actInteract with the pageA
Destructive

Click, type, fill a form, select an option, hover, scroll, press a key, go back, or dismiss a cookie or consent overlay. Target elements by the @eN reference from steel_snapshot or steel_find, or by a CSS selector. Always reports what actually changed, and says so plainly when nothing did.

ParametersJSON Schema
NameRequiredDescriptionDefault
valueNoText to type, option to select, key name to press, or scroll distance in pixels.
actionYesWhat to do.
fieldsNoFor fill_form: the fields to fill, in order, in one round trip.
targetNoA @eN reference or a CSS selector. Not needed for scroll, press, go_back or dismiss_overlays.
max_tokensNoCap on the text returned, in tokens. Defaults to 8000.
session_idYesLive session_id from steel_session_create.
include_snapshotNoAlso return the page structure afterwards. Off by default.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide destructiveHint=true and openWorldHint=true, so the mutation risk is flagged. The description adds value beyond that by promising that the tool always reports what actually changed and says so plainly when nothing did, and by enumerating the side-effecting actions it performs. This is useful behavioral context rather than a contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences cover the action set, targeting strategy, and output reporting behavior. Every sentence earns its place; there is no filler or repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter, 10-action tool with no output schema, the description covers the action set, target syntax, and reporting promise. Operational details like session_id and max_tokens are handled by the schema. The return format is only vaguely described, but the honesty guarantee compensates enough for a tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented in the schema. The description adds only light enrichment, such as mentioning @eN references from steel_snapshot/steel_find and cookie/consent overlays. Since the schema carries the parameter burden, a baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a concrete set of interactions—click, type, fill a form, select, hover, scroll, press, go back, dismiss overlays—and ties them to a specific resource: the current page. It also distinguishes itself from siblings like steel_navigate and steel_snapshot by making clear this is the general page-interaction tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: target elements by @eN references from steel_snapshot or steel_find, or by CSS selector. This implies the prerequisite of running snapshot/find first and orients the agent to the correct workflow. It does not explicitly state when not to use the tool or name alternatives, but the usage situation is still clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

steel_batchRun several browser steps at onceA
Destructive

Run known reversible steps whose later targets need no fresh read. Stops on failure or login/challenge; hand off the same session and resume only unrun steps. Stop before payment/final confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
stepsYesSteps to run in order.
max_tokensNoCap on the text returned, in tokens. Defaults to 8000.
session_idYesLive session_id from steel_session_create.
include_snapshotNoReturn the page structure once, after the last step.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint and openWorldHint, but the description goes further by disclosing that execution stops on failure or login/challenge, the session is handed off, and only unrun steps should be resumed. It also warns against including payment or final confirmation steps. This adds meaningful behavioral context beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with no wasted words. The first sentence front-loads the core usage condition, the second explains failure behavior, and the third gives a safety boundary. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a batch tool with no output schema, the description covers the essential selection criteria, failure behavior, and safety constraints. It does not describe the return shape or how results from completed steps are surfaced, which is a minor gap, but the critical invocation decisions are well supported.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the session_id, steps, max_tokens, and include_snapshot parameters. The description does not add parameter-specific syntax or format details, which is acceptable given full schema coverage. It provides general behavioral guidance but not deeper parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool runs multiple known reversible browser steps as a batch, which distinguishes it from the single-step sibling tools like steel_navigate, steel_act, and steel_wait_for. The title reinforces this by saying 'Run several browser steps at once.' It does not name a sibling explicitly, but the scope and specificity are strong enough to identify the tool's purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a concrete inclusion criterion: use for known reversible steps whose later targets do not need a fresh read. It also gives exclusions by warning to stop before payment/final confirmation and noting the batch stops on login/challenge. It does not explicitly name an alternative tool, so it falls just short of fully explicit routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

steel_findFind an element on the pageA
Read-only

Find labelled elements by text, safe regex or role and return their @eN refs.

ParametersJSON Schema
NameRequiredDescriptionDefault
roleNoOnly return elements with this role, e.g. button or link.
textNoCase-insensitive substring of the element label.
regexNoSafe regular expression matched against the element label.
cursorNoCursor from a previous truncated response, to continue reading where it stopped.
max_tokensNoCap on the text returned, in tokens. Defaults to 8000.
session_idYesLive session_id from steel_session_create.
max_resultsNoCap on matches returned.
interactive_onlyNoOnly return elements that can actually be clicked or typed into.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare read-only and open-world behavior. The description adds useful context by noting that only labelled elements are matched, that regex matching is safe, and that the result is @eN refs. It does not disclose truncation or cursor continuation behavior, but the schema covers those details and the safety profile is already established by annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence states the action, the matching methods, and the return type with no filler. Every phrase earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with eight parameters and no output schema, the description provides the essential return concept (@eN refs) and the matching criteria. Combined with the fully described schema and read-only annotations, this is nearly complete; only explicit sibling differentiation and pagination caveats are missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3 because the schema already documents all eight parameters. The description adds a high-level grouping of search modes (text, regex, role) but no new semantic detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Find'), a concrete resource ('labelled elements'), and explicit search criteria (text, safe regex, role), plus the return type (@eN refs). This clearly differentiates it from sibling tools like steel_snapshot or steel_scrape, which retrieve page content rather than locate specific elements.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is reasonably clear: call this when you need element references by label. However, it never explicitly states when to prefer it over alternatives such as steel_snapshot or steel_wait_for, nor does it provide any exclusions. Usage is implied rather than directly guided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

steel_navigateOpen a URL in a browser sessionA
Destructive

Navigate a live session and wait for it to settle. Reports the final URL and changes; set include_snapshot to also read the page.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesWhere to go.
max_tokensNoCap on the text returned, in tokens. Defaults to 8000.
session_idYesLive session_id from steel_session_create.
include_snapshotNoAlso return the accessibility snapshot. Off by default because it is expensive.

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations, the description discloses that the tool waits for the page to settle, reports the final URL and changes, and that the snapshot is off by default because it is expensive. These are meaningful behavioral details. It does not contradict the openWorldHint or destructiveHint annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences front-load the core behavior and then add the key option. Every clause earns its place with no repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a navigation tool with no output schema, the description covers the main behavioral contract: navigating, waiting, reporting URL and changes, and the snapshot option. It does not fully describe the shape of the returned changes, but with annotations covering safety and 100% parameter schema coverage, the gaps are acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds value by explaining include_snapshot ('also read the page', 'expensive') and by clarifying what navigation returns ('reports the final URL and changes'), which gives context to the url and session_id parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action: 'Navigate a live session and wait for it to settle.' The title 'Open a URL in a browser session' reinforces the resource and action. It does not explicitly compare against sibling tools like steel_scrape or steel_act, so it misses full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it: to navigate a live session and wait for it to settle, with include_snapshot when page content is needed. However, it offers no explicit exclusions or alternatives among siblings, such as 'use steel_scrape for extraction' or 'use steel_act for actions.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

steel_pdfRender a web page as PDFA
Read-onlyIdempotent

Render a page to PDF and return a link to the file. Starts no browser session. Use steel_scrape if you want to read the text — the PDF link is for handing a document to a person.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe page to render.
delay_msNoWait this long after load.
use_proxyNoUse a Steel residential proxy.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the read-only, idempotent, and open-world annotations, the description adds two useful behaviors: it starts no browser session and it returns a link rather than inline content. It does not contradict the annotations, so credit is given for the extra context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two succinct sentences front-load the core purpose and return format, then give the alternative and use case. There is no filler or repetition; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple three-parameter tool with a full schema and safety annotations, the description covers purpose, return format, session behavior, and the key sibling alternative. No output schema exists, but the description explicitly states the return value, so nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already explains url, delay_ms, and use_proxy. The description adds no parameter-level detail, but the baseline of 3 applies because the schema carries the full burden and the description doesn't need to compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Render') and resource ('a page to PDF') and states the deliverable ('return a link to the file'). It also distinguishes itself from steel_scrape by noting the PDF is for handing a document to a person, so an agent can tell it apart from the most similar sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly routes the agent: use steel_scrape when the goal is to read the text, and use steel_pdf when the goal is to produce a document for a person. This is clear when-to-use guidance with a named alternative, leaving little to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

steel_scrapeRead a web pageA
Read-onlyIdempotent

Read a web page as markdown or HTML through a real browser, so JavaScript-rendered pages and sites that block plain HTTP fetches still work. Starts no browser session, so there is nothing to release afterwards. Always returns the page links and metadata alongside the content. Use this first for anything you only need to read; reach for steel_session_create only when you need to click, type or move through several pages.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe page to read.
cursorNoCursor from a previous truncated response, to continue reading where it stopped.
formatNoFormats to return. Each extra format repeats the page inside the shared budget.
delay_msNoWait this long after load.
use_proxyNoRoute through a Steel-managed residential proxy. Needs a verified paid balance.
max_tokensNoCap on the text returned, in tokens. Defaults to 8000.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint, and the description contradicts none of them. Beyond that, it adds three substantive behavioral facts: the tool runs a real browser so JS rendering works, it starts no browser session so there is nothing to release, and it always returns page links and metadata alongside content. This is solid context beyond the annotations, though it stops short of disclosing failure behavior or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences with zero waste: purpose and capability, session lifecycle, return-value behavior, and usage routing each get one sentence. The core action is front-loaded in the first line, and every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one required parameter, a fully covered schema, and rich annotations, the description covers selection criteria, session handling, and return basics. Since there is no output schema, the return structure is only sketched as 'links and metadata,' but the combination of a 100%-covered schema, strong annotations, and this description leaves no critical gap for correct selection and invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters are already fully documented in the schema. The description adds only marginal parameter-level meaning (markdown/HTML correspond to the format parameter, and 'no browser session' explains why no session parameter exists). Baseline 3 is appropriate since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific action — 'Read a web page as markdown or HTML through a real browser' — and adds distinguishing capability detail: it handles JavaScript-rendered pages and sites that block plain HTTP fetches. An agent can clearly tell this apart from steel_screenshot, steel_pdf, and steel_session_create without opening their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The last sentence is an explicit routing directive: 'Use this first for anything you only need to read; reach for steel_session_create only when you need to click, type or move through several pages.' It names the alternative tool and the exact condition that selects it, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

steel_screenshotScreenshot a web pageA
Read-onlyIdempotent

Capture a page image. URL captures are user-facing PNG artifacts; session captures are model-visible JPEG evidence. Pixels are not action targets, so use steel_snapshot to click or type.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoPage to capture. Starts no browser session.
inlineNoFor URL captures only: false returns just a link. Defaults to true.
full_pageNoCapture the whole scrollable page, not just the viewport.
use_proxyNoUse a Steel residential proxy for a URL capture.
session_idNoCapture the current page of this session instead. Returns the image inline.

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint, so the safety profile is covered. The description adds behavioral context beyond annotations: output artifact formats (PNG vs JPEG) and the non-interactive nature of pixels. It does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two tight sentences with no filler. The core action is front-loaded, artifact distinctions follow, and the sibling-routing caveat is placed last. Every sentence earns its place and nothing repeats schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 100% schema coverage and read-only/idempotent annotations, the description covers the key disambiguation from steel_snapshot and the two capture modes. It does not explain what happens when neither url nor session_id is provided, but the optional parameter schema and mode descriptions give sufficient guidance for a competent agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema descriptions cover all 5 parameters, so the baseline is 3. The description adds mode-level meaning that helps select between url and session_id by specifying output format and audience. It does not re-explain full_page or use_proxy, but those are already well documented in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Capture a page image.' It further distinguishes URL captures and session captures by artifact type, and differentiates the tool from steel_snapshot by noting that pixels are not action targets. This makes the tool's purpose unambiguous relative to siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells agents when not to use this tool: 'Pixels are not action targets, so use steel_snapshot to click or type.' It also provides mode-selection guidance by contrasting user-facing PNG URL captures with model-visible JPEG session captures. This gives clear routing versus the main sibling alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

steel_session_createStart sessionC
Destructive

Billed; profiles/credentials via session_options; release.

ParametersJSON Schema
NameRequiredDescriptionDefault
guestNoFresh browser; skip profiles.
deviceNoClass.
viewportNoPixels.
block_adsNoAds.
namespaceNoName; not secret.
use_proxyNoProxy.
profile_idNoUUID; not secret.
timeout_msNoLifetime ms.
configurationNoPlan token.
solve_captchaNoCAPTCHA.

TDQS

C2.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With only openWorldHint/destructiveHint annotations, the description adds meaningful behavioral context: the operation is billed, and the created session must be released, which suggests a chargeable, long-lived resource. It doesn't go into consequences, but it reveals cost and lifecycle beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The definition is extremely short, but it is under-specified rather than economically structured; the fragments are not front-loaded with the core purpose and read as disconnected warnings rather than a purposeful summary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter, billed, destructive-flagged creator with no output schema, the description omits crucial context: what the call returns (e.g., session id), how billing is triggered, what 'release' implies, and any setup prerequisites. The cost/lifecycle hints are useful but far from complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 10 parameters. The prose description adds no parameter-level detail; 'profiles/credentials via session_options' is a routing hint, not semantics for any specific parameter. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description never states the core action; it is a fragment list ('Billed; profiles/credentials via session_options; release.') that forces the agent to infer 'create session' from the name/title. It also doesn't differentiate this creator from siblings like steel_session_live_view beyond the title.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The fragments imply some usage boundaries: credentials/profiles should go through steel_session_options, and the caller should later release the session. However, there is no explicit when-to-use, prerequisites, or exclusion guidance for the other sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

steel_session_diagnosticsExplain what a browser session didA
Read-onlyIdempotent

Read live/released activity or list live handles; never starts a browser.

ParametersJSON Schema
NameRequiredDescriptionDefault
sinceNoOnly events at or after this ISO-8601 time.
cursorNoCursor from a previous truncated response, to continue reading where it stopped.
list_liveNoList this credential's recoverable live session handles.
max_tokensNoCap on the text returned, in tokens. Defaults to 8000.
session_idNoLive id from steel_session_create.
steel_session_idNoFinished Steel dashboard UUID.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and idempotentHint, and the description reinforces safety with 'Read' and 'never starts a browser.' It adds a useful behavioral boundary that is not fully captured by the annotations alone: no browser is launched during invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short sentence delivers the core action, the scope, and a key exclusion. Every word earns its place, and the most important behavior is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only diagnostics tool with no required parameters, the description plus annotations provide enough context for selection and safe invocation. A bit more detail about what 'activity' includes or how output is returned would help, but the absence is not critical given the schema and annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameters are fully documented there. The description's 'live/released activity' and 'list live handles' loosely map to session_id, steel_session_id, and list_live, but it adds no parameter-level detail beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific operation and resource: 'Read live/released activity or list live handles.' It also distinguishes itself from sibling tools by adding 'never starts a browser,' which separates it from session-creating or browser-acting tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly indicates this tool is for reading diagnostics or listing live handles, not for launching a browser. However, it does not explicitly name siblings like steel_session_live_view or steel_session_replay, leaving some routing to the agent's inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

steel_session_handoffHand the browser to a personA
Destructive

Pause for a person to take exclusive control of this same live browser, then resume only after hand-back. Use for sensitive input, local files, review, or any manual step.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonYesWhy a person needs control.
session_idYesLive session_id from steel_session_create.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds behavior not visible in annotations: the tool pauses the current flow, grants exclusive human control, and resumes only after hand-back. This goes beyond the destructiveHint/openWorldHint annotations to explain the handoff lifecycle.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler; the core behavior is front-loaded and the usage guidance follows immediately. Every phrase earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter handoff tool with annotations covering safety, the description fully explains what happens, when to use it, and how the handoff is resolved. No output schema is needed because the behavior is the primary contract.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and already documents both parameters. The description still adds value by mapping the reason enum to concrete scenarios (sensitive input, local files, review, or any manual step), which helps an agent choose the correct value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states a clear action—'Pause for a person to take exclusive control of this same live browser'—and the title reinforces the handoff behavior. It is easily distinguished from sibling session/automation tools, which all operate under agent control.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly lists when to use it: 'sensitive input, local files, review, or any manual step.' It doesn't explicitly state when not to use it or name alternatives, but the use-case list is sufficient guidance for an agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

steel_session_live_viewLive view connectionC
Destructive

App-only viewer and control lease; no page content.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNo
session_idNoOmit to stream the newest live session this credential owns, which is what a viewer that never received the session does.
control_tokenNo

TDQS

C2.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare openWorldHint and destructiveHint, so the description does not need to restate those. It adds a little context with 'app-only' and 'no page content,' but it does not disclose the lease lifecycle, side effects of release, or what happens to a control lease across actions. There is no contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short and contains no filler, but it is under-specified to the point of being a cryptic fragment. It front-loads no actionable instruction and leaves the agent to infer meaning from the title and schema, which is not effective conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter control tool with no output schema, destructive and open-world annotations, and many related session siblings, this description is insufficient. The only substantive guidance comes from the session_id schema description; the agent still lacks action semantics, token requirements, and selection criteria.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%: only session_id has a description, while action and control_token are undocumented. The description's word 'lease' loosely relates to the action enum, but it does not explain connect, acquire, renew, or release semantics, nor does it clarify what control_token is or when it is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description communicates that this tool is about an app-only viewer/control lease and explicitly notes 'no page content,' which gives some domain context. However, it is phrased as a noun fragment rather than a clear verb+resource statement, and it does not distinguish this tool from siblings like steel_session_create, steel_session_release, or steel_session_handoff.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus its siblings. 'No page content' hints at an exclusion, but there are no explicit conditions, alternatives, or when-not-to-use signals, leaving an agent to guess whether live_view is appropriate relative to session_create, session_release, or handoff.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

steel_session_optionsPlan sessionC
Read-only

Find profiles/credentials; plan setup.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL.
goalYesMode.
needsNoUnique;protected_text=read;captcha!=read/protected_text;persist_profile=account.
cursorNoCursor.
countryNolocation ISO.
minutesNolong_running min.

TDQS

C2.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds minimal behavioral context consistent with read-only behavior (finding profiles/credentials and planning), but it does not disclose what the plan contains, whether live network checks are performed, or what the output looks like. It does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At six words, the description is under-specified rather than genuinely concise. 'Find profiles/credentials; plan setup' reads like a note-to-self and fails to earn its place as a functional tool definition—there is no front-loaded verb-object statement of effect, no scope, and no return behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 6 parameters, no output schema, and complex inter-parameter constraints (needs enums with exclusion logic), this description is far too thin. An agent cannot determine what the tool returns, how the goal enum affects the plan, or how this differs from other session tools. The cryptic schema notes are not compensated for by the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline of 3 applies. The description itself adds nothing about parameters like needs, cursor, country, or minutes, and the schema's own descriptions are cryptic (e.g., goal is just 'Mode'; needs uses terse constraint strings like 'protected_text=read'), but per the rubric the high schema coverage keeps this at baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description offers a concrete verb+resource in 'Find profiles/credentials' but 'plan setup' is vague and largely restates the title 'Plan session.' It hints at a discovery/planning operation but never states what the tool actually returns or does with the found profiles/credentials, and it does not clearly distinguish itself from siblings like steel_session_create or steel_find.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. It does not mention steel_session_create, steel_find, or any sibling, nor any condition such as 'call before creating a session.' An agent is left to infer the intended workflow entirely from the tool name and terse fragments.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

steel_session_releaseRelease a browser sessionA
DestructiveIdempotent

Shut down the current browser and stop the meter. Safe to call twice. Its current URL and session-only page state are gone afterwards. A profile is saved only when persistence was requested. Read what you need first; this reports the final URL and title.

ParametersJSON Schema
NameRequiredDescriptionDefault
session_idYesLive session_id from steel_session_create.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial behavioral detail beyond the annotations: it reports idempotency ('Safe to call twice'), destructive effects ('current URL and session-only page state are gone'), persistence rules, and return value ('reports the final URL and title'). This fully discloses the side effects and consequences of calling the tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, front-loaded with the primary action ('Shut down the current browser and stop the meter'), then covers safety, side effects, persistence, and return behavior in four short sentences. Every sentence adds meaningful information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity tool with one parameter and no output schema, the description is complete: it covers the destructive nature, idempotency, persistence behavior, return value, and usage timing. Nothing needed for correct invocation or expectation-setting is omitted.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides 100% coverage for the single parameter, including 'Live session_id from steel_session_create.' The description does not need to add parameter-level detail, and it does not. Baseline score of 3 is appropriate because the description adds no additional parameter semantics beyond what the schema already documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Shut down the current browser and stop the meter.' This is a specific verb + resource combination that makes the core purpose obvious. It does not explicitly name sibling tools for differentiation, but the terminating nature is distinct enough from the listed siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: 'Read what you need first' signals that this should be called after data extraction is complete. It also notes the persistence caveat ('A profile is saved only when persistence was requested'), which informs the user whether they can expect preserved state. It does not explicitly contrast with alternatives like handoff or replay, but the when-to-use guidance is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

steel_session_replayOpen a finished sessionA
Read-onlyIdempotent

Call only when the user explicitly asks to watch or replay a finished session. Returns its Steel dashboard link without starting a browser; use steel_session_diagnostics to inspect or explain activity.

ParametersJSON Schema
NameRequiredDescriptionDefault
steel_session_idNoFinished Steel UUID; omit for latest released.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already carry the safety profile (readOnlyHint, idempotentHint, openWorldHint), so the bar is lower, and the description adds genuinely valuable behavioral context beyond them: it clarifies the tool returns a dashboard link and does NOT start a browser, correcting the natural but wrong assumption that 'replay' spawns a browser session. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero fluff, and the most decision-critical information (usage condition) is front-loaded. Every clause earns its place: the trigger, the behavioral clarification, the return value, and the sibling routing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool of this complexity — one optional parameter, full schema coverage, strong annotations, simple return value — nothing an agent needs is missing. It covers when to call, what happens (no browser), what comes back (dashboard link), and where to go for the related task, despite having no output schema to lean on.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameter (steel_session_id with UUID pattern and 'Finished Steel UUID; omit for latest released') is fully documented in the schema. The description adds no parameter-level detail, which is appropriate — baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('watch or replay a finished session') and a concrete outcome ('Returns its Steel dashboard link without starting a browser'). Explicitly differentiates from sibling steel_session_diagnostics, and the contrast with steel_session_live_view (finished vs. live) is evident from context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit, restrictive trigger ('Call only when the user explicitly asks to watch or replay a finished session') and names the alternative tool for the neighboring case ('use steel_session_diagnostics to inspect or explain activity'). An agent needs zero inference to decide when to invoke this.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

steel_snapshotRead the page structureA
Read-only

Return the page as a compact accessibility tree with a @eN reference on every element you can click or type into. This is the read to use before acting. Elements with no reference cannot be targeted. If you already know what you are looking for, steel_find is much cheaper.

ParametersJSON Schema
NameRequiredDescriptionDefault
cursorNoCursor from a previous truncated response, to continue reading where it stopped.
max_tokensNoCap on the text returned, in tokens. Defaults to 8000.
session_idYesLive session_id from steel_session_create.
interactive_onlyNoSkip purely structural containers. On by default; turn off for the full tree.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the read-only safety profile and open-world scope, so the bar is lower. The description adds genuine value: the compact-tree output format, the critical behavioral rule that unreferenced elements cannot be targeted, and a cost/performance comparison ('much cheaper'). Truncation/pagination behavior is left to the schema's cursor parameter, but this is a minor gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences with zero filler. The core function is front-loaded first, followed by usage timing, the targeting constraint, and the cheaper alternative. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only tool with full annotation coverage, 100% schema description coverage, and no output schema, the description is complete: it explains what is returned, when to call it, what the references mean operationally, and when to prefer a sibling. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (cursor, max_tokens, session_id, interactive_only) are already documented in the schema. The description does not add parameter-specific detail, so the baseline 3 applies — the schema carries the load and does so adequately.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Return the page as a compact accessibility tree with a @eN reference on every element you can click or type into.' The distinctive @eN reference format makes the output immediately recognizable, and the mention of steel_find as an alternative helps differentiate it from at least one sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use guidance ('This is the read to use before acting'), a hard constraint that shapes usage ('Elements with no reference cannot be targeted'), and names the alternative with the condition that selects it ('If you already know what you are looking for, steel_find is much cheaper').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

steel_wait_forWait for something on the pageA
Read-only

Wait for named text, a CSS selector or URL substring; pass at least one.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoWait for the URL to contain this string.
textNoWait for this text to appear on the page.
selectorNoWait for an element matching this CSS selector.
session_idYesLive session_id from steel_session_create.
timeout_msNoGive up after this long.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description conveys the core waiting behavior and clarifies that conditions are alternatives. Annotations already mark the tool as read-only and open-world, so the safety profile is covered. It does not disclose timeout failure behavior, polling semantics, or what happens if multiple conditions are passed, but for a simple wait operation the stated behavior is reasonably transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one compact sentence with no wasted words. It front-loads the action and immediately enumerates the supported condition types, followed by the key invocation constraint.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The core invocation details are covered by the schema and the at-least-one note. However, there is no output schema and the description does not explain what happens on timeout or what a successful wait returns, which an agent may need to know for robust usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description adds meaningful parameter semantics by mapping 'named text', 'CSS selector', and 'URL substring' to the optional text, selector, and url fields. It also adds the at-least-one constraint, which the schema does not enforce.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: wait for page conditions including text, CSS selector, or URL substring. This avoids tautology and gives an agent a clear sense of what the tool does, though it does not explicitly contrast with sibling tools like steel_find or steel_snapshot.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Wait for...' implies the tool should be used when a page condition must be awaited, and 'pass at least one' gives a necessary precondition. However, it does not provide explicit guidance on when to choose this tool over alternatives or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 25 tool updatesv2.0.1
    • Removedclick
    • Removedgo_back
    • Removednavigate
    • Removedsave_unmarked_screenshot
    • Removedscroll_down
    • Removedscroll_up
    • Removedsearch
    • Addedsteel_act
    • Addedsteel_batch
    • Addedsteel_find
    • Addedsteel_navigate
    • Addedsteel_pdf
    • Addedsteel_scrape
    • Addedsteel_screenshot
    • Addedsteel_session_create
    • Addedsteel_session_diagnostics
    • Addedsteel_session_handoff
    • Addedsteel_session_live_view
    • Addedsteel_session_options
    • Addedsteel_session_release
    • Addedsteel_session_replay
    • Addedsteel_snapshot
    • Addedsteel_wait_for
    • Removedtype
    • Removedwait
  2. 9 tool updatesv1.0.0
    • First observedclick
    • First observedgo_back
    • First observednavigate
    • First observedsave_unmarked_screenshot
    • First observedscroll_down
    • First observedscroll_up
    • First observedsearch
    • First observedtype
    • First observedwait

TDQS

A3.5/5.0

Scored across 16 tools

Disambiguation4/5

Most tools have clearly distinct purposes, with stateless reading (scrape, screenshot, pdf) separated from session interaction (navigate, act, snapshot). A few pairs like steel_screenshot/steel_snapshot and steel_session_live_view/steel_session_replay have similar names, but detailed descriptions resolve the boundaries.

Naming Consistency4/5

All tools share the steel_ prefix and snake_case, but the pattern mixes bare verbs (scrape, navigate, act) with session_ prefixed noun_verb forms (session_create, session_release, session_handoff). This is consistent enough to be predictable, with minor deviations like steel_pdf and steel_wait_for.

Tool Count4/5

16 tools is slightly above the typical 3-15 range, but the breadth covers stateless captures, session lifecycle, element targeting, and diagnostics. Each tool addresses a distinct need, so the count feels justified rather than bloated.

Completeness4/5

The surface covers the core browser automation lifecycle: create/release sessions, navigate, read, interact, wait, and diagnose. Obvious gaps include a full-HTML/markdown read within a live session and session-scoped PDF rendering, but these can be worked around with snapshot or separate stateless calls.

Maintenance

ActivityMaintained
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers