Skip to main content
Glama

SohiB

Research Indonesian equities with AI agents through your own Stockbit account.

SohiB is a local MCP server. Point Claude Code, Claude Desktop, Codex, Cursor, OpenClaw or any other MCP client at it and the agent gains thirteen read-only research tools: company search, key statistics, financial statements, price history, analyst consensus, peer comparison, corporate actions, the dividend calendar and unsaved screening.

Every request runs inside a Chrome tab where you are already signed in to Stockbit — the page reads its own session cookie and sets the auth header itself, so SohiB never asks for your password, never stores a token, and its own process never sees one either. Read-only is structural here, not a setting: there is no login command, no token store and no order, watchlist or portfolio route anywhere in the codebase to turn on.

SohiB is an independent open-source project. It is not a Stockbit product. Use it with your own account under Stockbit's terms of service.

How it works

  1. sohib connect opens a dedicated Chrome profile with a loopback-only debugging port. You sign in to Stockbit yourself, including any device verification.

  2. When an agent calls a tool, SohiB attaches to that tab over the local debugging connection and runs one fixed JavaScript function that calls Stockbit's own web API from the page. The page reads its session token from Stockbit's cookie and sets the header itself. The token never leaves the browser.

  3. The response is validated, normalized into one envelope (records, provenance, pagination, warnings, error) and returned to the agent. Only the routes listed in sohib/stockbit/api_routes.py can be requested.

Related MCP server: financial-research-agent

Requirements

  1. Python 3.12 or newer

  2. Google Chrome or Chromium on the same machine as your agent client

  3. A Stockbit account

  4. macOS, Linux or Windows. Windows support is new; please report issues.

Quickstart

git clone https://github.com/yusufsiregar44/sohib.git && cd sohib
python3 -m pip install .             # not yet on PyPI; install from this checkout
sohib setup                          # private configuration in your user data directory
sohib connect                        # opens Chrome; sign in to Stockbit yourself
sohib doctor --smoke SIDO            # checks the session and runs one live search

Then print the configuration for your client and paste it where that client expects it:

sohib config claude      # Claude Code and Claude Desktop (mcpServers JSON)
sohib config codex       # Codex (TOML block)
sohib config openclaw    # mcporter JSON for OpenClaw
sohib config mcp         # generic mcpServers JSON (Cursor, Windsurf, others)

For Claude Code, one command is enough:

sohib config claude > sohib-mcp.json && claude --mcp-config sohib-mcp.json

Keep exactly one Stockbit tab open in the SohiB Chrome window. After a reboot, run sohib connect again; the profile keeps your session until Stockbit expires it, at which point tools return AUTH_EXPIRED or AUTH_CHALLENGE and you sign in again in the browser.

Try asking your agent:

Search Stockbit for SIDO, confirm the symbol, compare its PE ratio with its industry, and show the last four quarters of revenue. Cite the source URLs and keep the warnings.

Tools

Tool

What it returns

search_companies

IDX equity candidates for a symbol or name; warrants excluded, ambiguity preserved

get_company_summary

Identity, sector classification and the last provider quote

get_key_statistics

Grouped valuation, profitability, growth and balance-sheet metrics

list_metrics

Metric labels and IDs for keystats, screener, fundachart, comparison and financials

get_financial_statements

Income statement, balance sheet or cash flow cells; quarterly, annual or TTM

get_fundamental_history

Historical series for one FundaChart metric

get_price_series

LINE price points with a period summary (not OHLCV)

get_price_performance

Price change, high and low across provider windows

get_analyst_consensus

Recommendation counts, price targets and yearly estimates

get_peer_comparison

Ratios beside industry and sector aggregates, plus peer symbols

get_corporate_actions

Dividends, meetings, splits and tender offers with dates kept separate

get_dividend_calendar

Market-wide dividend calendar and today's scheduled events

screen_equities

Unsaved IHSG screen from numeric rules on screener metrics

Full contracts, input schemas and live evidence: docs/stockbit.

What the agent gets back

Every tool returns the same envelope. status is ok, partial (valid data with warnings you should keep) or error. Records carry a value, the provider's display string, unit, currency, period and a missing_reason when the provider had no value. Missing values are never turned into zero. provenance.fetched_at is retrieval time, not publication time; data_as_of is null unless the provider states it. Each record cites a source_url on stockbit.com.

Safety model

None of this is a default you could switch off — it is the entire feature set. There is no code path anywhere in this repository that stores a token or places an order.

  1. Read-only. There is no order, watchlist, portfolio or account tool, and no write route in the allowlist except one unsaved screen execution whose body is checked field by field.

  2. Your credentials stay with you. SohiB has no login step and no password file. The browser reads its own cookie; SohiB's Python process never sees the token.

  3. Fixed page function. Agents pass arguments, never JavaScript. Paths are validated against the allowlist before anything is sent to the browser.

  4. Explicit failure. Expired sessions, verification prompts, rate limits and provider blocks come back as typed errors. Nothing retries, nothing logs in for you, nothing bypasses Cloudflare.

  5. One call at a time per browser. Concurrent callers get a busy error rather than a queue.

See SECURITY.md for details and how to report a problem.

Limits

  1. Price data is the provider's LINE series only. Timezone and adjustment are unverified.

  2. Latency is dominated by Stockbit's response time: usually one to three seconds, occasionally much longer on cold endpoints. Give your client a tool timeout of at least 60 seconds.

  3. Responses are capped at about 24 KB per page. Use query filters and page to move through large results.

  4. Nothing here is investment advice. Values are observations at retrieval time.

Development

git clone https://github.com/yusufsiregar44/sohib.git && cd sohib
python3 -m venv .venv && source .venv/bin/activate
python -m pip install -e . && python -m pip install --group dev
python -m ruff check . && python -m ruff format --check .
python -m unittest discover -s tests -v

Tests run offline against synthetic examples; optional private-capture tests skip when local evidence is absent. Live verification against Stockbit is manual; a descriptive summary of the last run is in docs/stockbit/stockbit-evidence. See CONTRIBUTING.md for how to add an endpoint.

Documentation

  1. Setup and browser lifecycle

  2. MCP server, CLI and verified clients

  3. Codex and Claude · OpenClaw

  4. Architecture and decisions

  5. Stockbit API catalogue, contracts and evidence

License

Apache-2.0. See LICENSE.

Available Tools

3 tools
get_key_statisticsA
Read-onlyIdempotent

Read key statistics for a confirmed symbol, filtered by label/ID. Pages contain up to 20 records. Preserve missing periods, currency and source dates. Mode: disabled. Research is disabled; calls return AUTH_REQUIRED without retrieving data.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageYes
queryYesCase-insensitive label/ID filter; empty for all.
symbolYesStock symbol; null for screener taxonomy. Fixture coverage varies.

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorYes
statusYes
recordsYes
warningsYes
paginationYes
provenanceYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Even though annotations already mark it read-only and idempotent, the description adds meaningful context: page size cap of 20, preservation of missing periods/currency/source dates, and the disabled mode with AUTH_REQUIRED failures. No contradiction with the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and stays compact. 'Mode: disabled' and 'Research is disabled' are mildly redundant, but the information density is high and every sentence contributes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only tool with an output schema and robust annotations, the description covers purpose, filter, pagination, and the current auth/disabled state. It is sufficient for an agent to understand the call and its constraints, though it does not suggest alternative tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents query and symbol; the description adds that filtering is by label/ID and that pages hold up to 20 records, giving page some extra context. It does not deeply explain page-number semantics, so it adds only moderate value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Read key statistics for a confirmed symbol') and adds the label/ID filter, clearly identifying the operation. It does not explicitly contrast with sibling tools, but the resource and filter are distinct enough to avoid major ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Confirmed symbol' implies a prerequisite and gives an implied use case, and the disabled-mode warning tells the agent that calls will fail with AUTH_REQUIRED. There is no explicit when-to-use versus search_companies or list_metrics, so it stops short of clear routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_metricsA
Read-onlyIdempotent

Discover metric metadata without numeric values. keystats and financials require a symbol; screener, fundachart and comparison are market-wide (symbol=null). IDs are namespaced per service; never merge them. Mode: disabled. Research is disabled; calls return AUTH_REQUIRED without retrieving data.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageYes
queryYesCase-insensitive label/ID filter; empty for all.
symbolYesStock symbol; null for screener taxonomy. Fixture coverage varies.
namespaceYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorYes
statusYes
recordsYes
warningsYes
paginationYes
provenanceYes

TDQS

A3.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the readOnly/idempotent annotations by stating 'Mode: disabled' and 'calls return AUTH_REQUIRED without retrieving data,' which is critical for an agent deciding whether/how to invoke it. It also discloses that IDs are namespaced per service and must not be merged. No annotation contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The tool's purpose is front-loaded and every sentence carries information. The disabled/research sentences are slightly choppy and redundant ('Mode: disabled' plus 'Research is disabled'), but the description remains efficiently packaged.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple metadata-listing tool, the description covers purpose, symbol semantics, ID namespacing, and auth failure behavior, with an output schema covering return shape. The gap is that the valid namespace enum contradicts the broader service list in the description, which an agent would need clarified to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema coverage at 50%, the description compensates partially by explaining symbol=null for market-wide services and by stressing service-namespaced IDs. But it references service names absent from the namespace enum and stays silent on page semantics, so the added value is real but imperfect.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific action and resource: 'Discover metric metadata without numeric values,' which clearly differentiates it from a numeric statistics tool like get_key_statistics. However, the description then lists financials, fundachart, and comparison as supported service scopes even though the namespace enum only allows keystats and screener, so the scope is slightly muddled.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives useful context for when to pass a symbol vs null ('keystats and financials require a symbol; screener, fundachart and comparison are market-wide'), and it notes the tool is disabled with calls returning AUTH_REQUIRED. It never explicitly contrasts this tool with its siblings search_companies or get_key_statistics, nor says when to prefer another tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_companiesA
Read-onlyIdempotent

Search IDX equities by symbol or company name. Excludes warrants; preserves multiple candidates. Confirm the company before fetching statistics. Mode: disabled. Research is disabled; calls return AUTH_REQUIRED without retrieving data.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
errorYes
statusYes
recordsYes
warningsYes
paginationYes
provenanceYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the call as read-only, open-world, and idempotent, and the description adds meaningful behavioral detail: it excludes warrants, preserves multiple candidate matches, requires confirmation before fetching statistics, and discloses that calls currently return AUTH_REQUIRED without retrieving data. No contradiction with annotations exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded with the core search purpose, followed by important exclusions and disabled-state warning. Minor redundancy exists between 'Mode: disabled' and 'Research is disabled,' but overall every sentence adds useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple one-parameter schema, the presence of an output schema, and rich annotations, the description is complete enough for an agent to call or avoid the tool correctly. It covers what is searched, what is excluded, the candidate behavior, the need for confirmation, and the current AUTH_REQUIRED failure mode.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain the query parameter. It does so by stating the query can be a 'symbol or company name' and clarifies that multiple candidates may be preserved. It does not give examples or case-format details, but for a single parameter this is adequate compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Search IDX equities by symbol or company name.' It further distinguishes the tool by noting it 'Excludes warrants' and 'preserves multiple candidates,' and its role is clearly positioned as a lookup step before fetching statistics. This separates it from sibling tools like list_metrics and get_key_statistics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Confirm the company before fetching statistics' explicitly places this tool before downstream statistics tools, giving an agent a clear workflow cue. It also warns that 'Research is disabled; calls return AUTH_REQUIRED without retrieving data,' which tells the agent when the tool will not be usable. It does not explicitly name alternatives, but the workflow context is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observedget_key_statistics
    • First observedlist_metrics
    • First observedsearch_companies

TDQS

A4.2/5.0

Scored across 3 tools

Disambiguation5/5

Each tool targets a distinct action: searching companies, listing metric metadata, and retrieving key statistics. There is no overlap or ambiguity; even though all are disabled, their purposes are clear.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern: search_companies, list_metrics, get_key_statistics. This predictable naming makes it easy for an agent to infer functionality.

Tool Count5/5

With only 3 tools, the server is tightly scoped to a focused research workflow (search, discover metrics, retrieve stats). This is within the ideal 3-15 range and each tool earns its place.

Completeness4/5

The tools cover a basic read-only lifecycle: search for a company, discover available metrics, and retrieve specific statistics. Minor gaps exist (e.g., no full financial statements or historical data), but the core workflow is supported without dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Provides AI assistants with comprehensive access to real-time and historical financial market data including stocks, options, crypto, and forex through the InsightSentry API. Supports 28+ specialized tools for market analysis, screening, and news retrieval, functioning as both an MCP server and standalone CLI.
    29
    228 npm
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    MCP server that exposes stock research tools (fundamentals, news, technicals, analyst ratings) to AI clients, enabling autonomous generation of structured investment briefs.
    -
  • F
    license
    Not graded
    quality
    D
    maintenance
    MCP server exposing Indonesia Stock Exchange (IDX) market data as tools — fundamentals, broker flow, company profiles, and technical analysis via TA-Lib.
    1
    -
  • A
    license
    Not graded
    quality
    C
    maintenance
    A production-grade MCP server that provides AI assistants with real-time financial market data, company metrics, and historical prices using yfinance and FastMCP.
    MIT