SohiB
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@SohiBSearch Stockbit for SIDO and show its key statistics and analyst consensus."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
SohiB
Research Indonesian equities with AI agents through your own Stockbit account.
SohiB is a local MCP server. Point Claude Code, Claude Desktop, Codex, Cursor, OpenClaw or any other MCP client at it and the agent gains thirteen read-only research tools: company search, key statistics, financial statements, price history, analyst consensus, peer comparison, corporate actions, the dividend calendar and unsaved screening.
Every request runs inside a Chrome tab where you are already signed in to Stockbit — the page reads its own session cookie and sets the auth header itself, so SohiB never asks for your password, never stores a token, and its own process never sees one either. Read-only is structural here, not a setting: there is no login command, no token store and no order, watchlist or portfolio route anywhere in the codebase to turn on.
SohiB is an independent open-source project. It is not a Stockbit product. Use it with your own account under Stockbit's terms of service.
How it works
sohib connectopens a dedicated Chrome profile with a loopback-only debugging port. You sign in to Stockbit yourself, including any device verification.When an agent calls a tool, SohiB attaches to that tab over the local debugging connection and runs one fixed JavaScript function that calls Stockbit's own web API from the page. The page reads its session token from Stockbit's cookie and sets the header itself. The token never leaves the browser.
The response is validated, normalized into one envelope (records, provenance, pagination, warnings, error) and returned to the agent. Only the routes listed in
sohib/stockbit/api_routes.pycan be requested.
Related MCP server: financial-research-agent
Requirements
Python 3.12 or newer
Google Chrome or Chromium on the same machine as your agent client
A Stockbit account
macOS, Linux or Windows. Windows support is new; please report issues.
Quickstart
git clone https://github.com/yusufsiregar44/sohib.git && cd sohib
python3 -m pip install . # not yet on PyPI; install from this checkout
sohib setup # private configuration in your user data directory
sohib connect # opens Chrome; sign in to Stockbit yourself
sohib doctor --smoke SIDO # checks the session and runs one live searchThen print the configuration for your client and paste it where that client expects it:
sohib config claude # Claude Code and Claude Desktop (mcpServers JSON)
sohib config codex # Codex (TOML block)
sohib config openclaw # mcporter JSON for OpenClaw
sohib config mcp # generic mcpServers JSON (Cursor, Windsurf, others)For Claude Code, one command is enough:
sohib config claude > sohib-mcp.json && claude --mcp-config sohib-mcp.jsonKeep exactly one Stockbit tab open in the SohiB Chrome window. After a reboot, run
sohib connect again; the profile keeps your session until Stockbit expires it, at which point
tools return AUTH_EXPIRED or AUTH_CHALLENGE and you sign in again in the browser.
Try asking your agent:
Search Stockbit for SIDO, confirm the symbol, compare its PE ratio with its industry, and show the last four quarters of revenue. Cite the source URLs and keep the warnings.
Tools
Tool | What it returns |
| IDX equity candidates for a symbol or name; warrants excluded, ambiguity preserved |
| Identity, sector classification and the last provider quote |
| Grouped valuation, profitability, growth and balance-sheet metrics |
| Metric labels and IDs for keystats, screener, fundachart, comparison and financials |
| Income statement, balance sheet or cash flow cells; quarterly, annual or TTM |
| Historical series for one FundaChart metric |
| LINE price points with a period summary (not OHLCV) |
| Price change, high and low across provider windows |
| Recommendation counts, price targets and yearly estimates |
| Ratios beside industry and sector aggregates, plus peer symbols |
| Dividends, meetings, splits and tender offers with dates kept separate |
| Market-wide dividend calendar and today's scheduled events |
| Unsaved IHSG screen from numeric rules on screener metrics |
Full contracts, input schemas and live evidence: docs/stockbit.
What the agent gets back
Every tool returns the same envelope. status is ok, partial (valid data with warnings you
should keep) or error. Records carry a value, the provider's display string, unit, currency,
period and a missing_reason when the provider had no value. Missing values are never turned into
zero. provenance.fetched_at is retrieval time, not publication time; data_as_of is null unless
the provider states it. Each record cites a source_url on stockbit.com.
Safety model
None of this is a default you could switch off — it is the entire feature set. There is no code path anywhere in this repository that stores a token or places an order.
Read-only. There is no order, watchlist, portfolio or account tool, and no write route in the allowlist except one unsaved screen execution whose body is checked field by field.
Your credentials stay with you. SohiB has no login step and no password file. The browser reads its own cookie; SohiB's Python process never sees the token.
Fixed page function. Agents pass arguments, never JavaScript. Paths are validated against the allowlist before anything is sent to the browser.
Explicit failure. Expired sessions, verification prompts, rate limits and provider blocks come back as typed errors. Nothing retries, nothing logs in for you, nothing bypasses Cloudflare.
One call at a time per browser. Concurrent callers get a busy error rather than a queue.
See SECURITY.md for details and how to report a problem.
Limits
Price data is the provider's LINE series only. Timezone and adjustment are unverified.
Latency is dominated by Stockbit's response time: usually one to three seconds, occasionally much longer on cold endpoints. Give your client a tool timeout of at least 60 seconds.
Responses are capped at about 24 KB per page. Use
queryfilters andpageto move through large results.Nothing here is investment advice. Values are observations at retrieval time.
Development
git clone https://github.com/yusufsiregar44/sohib.git && cd sohib
python3 -m venv .venv && source .venv/bin/activate
python -m pip install -e . && python -m pip install --group dev
python -m ruff check . && python -m ruff format --check .
python -m unittest discover -s tests -vTests run offline against synthetic examples; optional private-capture tests skip when local evidence is absent. Live verification against Stockbit is manual; a descriptive summary of the last run is in docs/stockbit/stockbit-evidence. See CONTRIBUTING.md for how to add an endpoint.
Documentation
License
Apache-2.0. See LICENSE.
Available Tools
3 toolsget_key_statisticsARead-onlyIdempotent
Read key statistics for a confirmed symbol, filtered by label/ID. Pages contain up to 20 records. Preserve missing periods, currency and source dates. Mode: disabled. Research is disabled; calls return AUTH_REQUIRED without retrieving data.
| Name | Required | Description | Default |
|---|---|---|---|
| page | Yes | ||
| query | Yes | Case-insensitive label/ID filter; empty for all. | |
| symbol | Yes | Stock symbol; null for screener taxonomy. Fixture coverage varies. |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | Yes | |
| status | Yes | |
| records | Yes | |
| warnings | Yes | |
| pagination | Yes | |
| provenance | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Even though annotations already mark it read-only and idempotent, the description adds meaningful context: page size cap of 20, preservation of missing periods/currency/source dates, and the disabled mode with AUTH_REQUIRED failures. No contradiction with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and stays compact. 'Mode: disabled' and 'Research is disabled' are mildly redundant, but the information density is high and every sentence contributes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with an output schema and robust annotations, the description covers purpose, filter, pagination, and the current auth/disabled state. It is sufficient for an agent to understand the call and its constraints, though it does not suggest alternative tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents query and symbol; the description adds that filtering is by label/ID and that pages hold up to 20 records, giving page some extra context. It does not deeply explain page-number semantics, so it adds only moderate value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Read key statistics for a confirmed symbol') and adds the label/ID filter, clearly identifying the operation. It does not explicitly contrast with sibling tools, but the resource and filter are distinct enough to avoid major ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Confirmed symbol' implies a prerequisite and gives an implied use case, and the disabled-mode warning tells the agent that calls will fail with AUTH_REQUIRED. There is no explicit when-to-use versus search_companies or list_metrics, so it stops short of clear routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_metricsARead-onlyIdempotent
Discover metric metadata without numeric values. keystats and financials require a symbol; screener, fundachart and comparison are market-wide (symbol=null). IDs are namespaced per service; never merge them. Mode: disabled. Research is disabled; calls return AUTH_REQUIRED without retrieving data.
| Name | Required | Description | Default |
|---|---|---|---|
| page | Yes | ||
| query | Yes | Case-insensitive label/ID filter; empty for all. | |
| symbol | Yes | Stock symbol; null for screener taxonomy. Fixture coverage varies. | |
| namespace | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | Yes | |
| status | Yes | |
| records | Yes | |
| warnings | Yes | |
| pagination | Yes | |
| provenance | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the readOnly/idempotent annotations by stating 'Mode: disabled' and 'calls return AUTH_REQUIRED without retrieving data,' which is critical for an agent deciding whether/how to invoke it. It also discloses that IDs are namespaced per service and must not be merged. No annotation contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The tool's purpose is front-loaded and every sentence carries information. The disabled/research sentences are slightly choppy and redundant ('Mode: disabled' plus 'Research is disabled'), but the description remains efficiently packaged.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple metadata-listing tool, the description covers purpose, symbol semantics, ID namespacing, and auth failure behavior, with an output schema covering return shape. The gap is that the valid namespace enum contradicts the broader service list in the description, which an agent would need clarified to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema coverage at 50%, the description compensates partially by explaining symbol=null for market-wide services and by stressing service-namespaced IDs. But it references service names absent from the namespace enum and stays silent on page semantics, so the added value is real but imperfect.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific action and resource: 'Discover metric metadata without numeric values,' which clearly differentiates it from a numeric statistics tool like get_key_statistics. However, the description then lists financials, fundachart, and comparison as supported service scopes even though the namespace enum only allows keystats and screener, so the scope is slightly muddled.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives useful context for when to pass a symbol vs null ('keystats and financials require a symbol; screener, fundachart and comparison are market-wide'), and it notes the tool is disabled with calls returning AUTH_REQUIRED. It never explicitly contrasts this tool with its siblings search_companies or get_key_statistics, nor says when to prefer another tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_companiesARead-onlyIdempotent
Search IDX equities by symbol or company name. Excludes warrants; preserves multiple candidates. Confirm the company before fetching statistics. Mode: disabled. Research is disabled; calls return AUTH_REQUIRED without retrieving data.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| error | Yes | |
| status | Yes | |
| records | Yes | |
| warnings | Yes | |
| pagination | Yes | |
| provenance | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the call as read-only, open-world, and idempotent, and the description adds meaningful behavioral detail: it excludes warrants, preserves multiple candidate matches, requires confirmation before fetching statistics, and discloses that calls currently return AUTH_REQUIRED without retrieving data. No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the core search purpose, followed by important exclusions and disabled-state warning. Minor redundancy exists between 'Mode: disabled' and 'Research is disabled,' but overall every sentence adds useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple one-parameter schema, the presence of an output schema, and rich annotations, the description is complete enough for an agent to call or avoid the tool correctly. It covers what is searched, what is excluded, the candidate behavior, the need for confirmation, and the current AUTH_REQUIRED failure mode.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain the query parameter. It does so by stating the query can be a 'symbol or company name' and clarifies that multiple candidates may be preserved. It does not give examples or case-format details, but for a single parameter this is adequate compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Search IDX equities by symbol or company name.' It further distinguishes the tool by noting it 'Excludes warrants' and 'preserves multiple candidates,' and its role is clearly positioned as a lookup step before fetching statistics. This separates it from sibling tools like list_metrics and get_key_statistics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Confirm the company before fetching statistics' explicitly places this tool before downstream statistics tools, giving an agent a clear workflow cue. It also warns that 'Research is disabled; calls return AUTH_REQUIRED without retrieving data,' which tells the agent when the tool will not be usable. It does not explicitly name alternatives, but the workflow context is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v0.1.0- First observed
get_key_statistics - First observed
list_metrics - First observed
search_companies
TDQS
Scored across 3 tools
Each tool targets a distinct action: searching companies, listing metric metadata, and retrieving key statistics. There is no overlap or ambiguity; even though all are disabled, their purposes are clear.
All tool names follow a consistent verb_noun pattern: search_companies, list_metrics, get_key_statistics. This predictable naming makes it easy for an agent to infer functionality.
With only 3 tools, the server is tightly scoped to a focused research workflow (search, discover metrics, retrieve stats). This is within the ideal 3-15 range and each tool earns its place.
The tools cover a basic read-only lifecycle: search for a company, discover available metrics, and retrieve specific statistics. Minor gaps exist (e.g., no full financial statements or historical data), but the core workflow is supported without dead ends.
Maintenance
Related MCP Connectors
MCP server exposing the Backtest360 engine API as tools for AI agents.
Research-only MCP server: your AI as a quant research desk. 90 tools, no trades, no brokers.
MCP server for OpenMM — exposes market data, account, trading, and strategy tools to AI agents
A Model Context Protocol server exposing real-time and historical Colombo Stock Exchange (CSE) data to AI agents and LLM applications. Provides quotes and OHLCV price history, full financial statements (income, balance sheet, cash flow), pre-computed technicals (moving averages, RS ratings, volume signals), macroeconomic indicators, corporate actions, and rule-based screening across CSE stocks and sector indices, everything needed to build CSE-aware trading assistants, research tools, and market-analysis agents. This is the official MCP server of www.ceyloncharts.com
Related MCP Servers
- AlicenseAqualityBmaintenanceProvides AI assistants with comprehensive access to real-time and historical financial market data including stocks, options, crypto, and forex through the InsightSentry API. Supports 28+ specialized tools for market analysis, screening, and news retrieval, functioning as both an MCP server and standalone CLI.29228 npmMIT
- FlicenseNot gradedqualityDmaintenanceMCP server that exposes stock research tools (fundamentals, news, technicals, analyst ratings) to AI clients, enabling autonomous generation of structured investment briefs.-
- FlicenseNot gradedqualityDmaintenanceMCP server exposing Indonesia Stock Exchange (IDX) market data as tools — fundamentals, broker flow, company profiles, and technical analysis via TA-Lib.1-
- AlicenseNot gradedqualityCmaintenanceA production-grade MCP server that provides AI assistants with real-time financial market data, company metrics, and historical prices using yfinance and FastMCP.MIT