Skip to main content
Glama
heyhumayun

financial-research-mcp

by heyhumayun

Multi-Agent Financial Research with MCP

Abstract

This project implements a bounded, auditable financial-research workflow. A supervisor routes a question to specialist agents, tools collect evidence through a local gateway or Model Context Protocol (MCP) transport, and a critic reviews coverage before a research brief is produced. A labelled offline benchmark evaluates routing, tool selection, evidence linkage, failure handling, and latency. The system is a research-assistance prototype; it does not execute trades or establish investment alpha.

Related MCP server: otto

Research objectives

  • Evaluate whether specialist-agent orchestration selects appropriate tools and evidence for financial-research questions.

  • Compare deterministic orchestration with optional local-LLM reasoning on the same labelled cases.

  • Make provider failures, offline fallbacks, source links, and execution traces visible in the final report.

  • Separate software-workflow quality from the factual and financial validity of a research conclusion.

Data

Live adapters can request market data from yfinance, news from Yahoo Finance RSS, papers from arXiv, and reported fundamentals from SEC Company Facts. Optional market and news providers require their respective credentials. Local documents can be searched with a token-vector retriever or optional sentence-transformer/FAISS retrieval.

The reproducible offline mode uses packaged deterministic fixtures, including synthetic market, news, and paper records and demonstration fundamentals. The 30-case evaluation set contains project-authored labels for required agents, tools, evidence categories, contradictions, and failure conditions. Neither the fixtures nor the labels are a point-in-time financial dataset or independently adjudicated analyst judgments.

Methodology

The supervisor selects a bounded set of market, news, fundamentals, comparison, research, document, and risk specialists. Specialists call typed tools and attach evidence identifiers to their findings. A critic checks configured coverage, fallback, and contradiction conditions; the supervisor may make one evidence-recovery pass. Synthesis produces a brief, structured output, and an execution trace. Optional Ollama reasoning can assist planning and synthesis; deterministic behavior remains available without an LLM.

The tool gateway supports direct local calls, an in-process MCP-shaped compatibility mode, and a real MCP client/server subprocess over stdio. Only mcp-stdio crosses an MCP transport boundary. Saved runs contain report.md, report.json, and trace.json under runs/<run_id>/.

Experimental setup

The primary evaluation is a 30-case deterministic offline benchmark. It measures routing and tool selection, citation coverage, claim-to-evidence linkage, contradiction and critic detection, fallback behavior, and stage-level latency. A separate six-case, category-stratified ablation compares deterministic orchestration with local llama3.2:3b reasoning. Failure-injection tests cover unavailable providers, malformed MCP responses, unavailable Ollama, stale fixtures, and contradictory findings. See EVALUATION.md for definitions and protocol details.

Results

The recorded offline benchmark from 2026-08-31 reported:

Metric

Result

Overall benchmark score

0.9799

Routing precision / recall

1.0000 / 1.0000

Tool-selection accuracy

1.0000

Citation coverage

0.9189

Claim-to-evidence support

0.9716

Contradiction recall

1.0000

Critic detection rate

1.0000

In the six-case local-model ablation, the deterministic and Ollama configurations scored 0.9668 and 0.9407, respectively. Exact-route rate fell from 1.0000 to 0.5000 with Ollama, while mean end-to-end latency rose from 0.98 ms to 50,679 ms in that sample. These are software-workflow measurements, not estimates of trading or investment performance.

Limitations

  • Offline fixtures are synthetic, and the project-authored labels lack independent analyst adjudication.

  • Live public sources may be delayed, rate-limited, incomplete, or unavailable; fallback data is explicitly marked.

  • Evidence identifiers establish source linkage, not excerpt-level entailment or factual correctness.

  • Headline sentiment, confidence values, and deterministic query routing are heuristics rather than calibrated financial models.

  • The six-case LLM ablation is too small for a general model-quality conclusion; offline latency is not representative of live providers.

  • The system does not implement point-in-time market-data controls, portfolio construction, execution, or financial outcome validation.

Repository structure

src/financial_research_agent/
  agents/          supervisor, specialists, critic, and synthesis
  mcp/             tool contracts and stdio server
  tools/           provider adapters, retrieval, and risk calculations
  data/            packaged fixtures and evaluation cases
  benchmark.py     labelled workflow evaluation
  cli.py           command-line interface
  api.py           FastAPI interface
  ui.py            Streamlit interface
tests/             unit and integration tests
EVALUATION.md      evaluation protocol and recorded measurements
PRODUCT_READINESS.md  capability and deployment-gap assessment

Installation

Python 3.10 or newer is required.

python3 -m venv .venv
source .venv/bin/activate
python3 -m pip install -e ".[dev,api,ui,market]"

The optional rag extra installs sentence-transformer and FAISS dependencies. Live providers and Ollama are not required for the offline workflow.

Usage

Run an offline investigation through the real MCP transport:

FIN_RESEARCH_LIVE=0 FIN_RESEARCH_TOOL_RUNTIME=mcp-stdio \
  financial-research-agent "Assess NVDA AI infrastructure risk and research" --save

Run the recorded evaluation procedure and repository checks:

FIN_RESEARCH_LIVE=0 financial-research-eval
ruff check src tests
python3 -m unittest discover -s tests -v

financial-research-api, financial-research-ui, and financial-research-mcp start the API, Streamlit interface, and standalone MCP server. Live-provider mode may require API credentials; use FIN_RESEARCH_OFFLINE_FALLBACK=0 to fail when a live source is unavailable instead of accepting fixture fallback.

References

Available Tools

10 tools
calculate_max_drawdownA

Calculate maximum drawdown from prices.

ParametersJSON Schema
NameRequiredDescriptionDefault
pricesYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The tool appears to be a pure calculation with no side effects, but the description does not disclose output format, edge cases, or assumptions about input ordering. Without annotations, more detail about behavior would be helpful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence with no superfluous words. It directly conveys the tool's purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is adequate for a simple calculation, but it does not explain the expected output (e.g., number, ratio, percentage) or any relevant calculation conventions. Since the tool appears to be part of a larger finance toolkit, additional context would improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides the parameter type and required status, and the description mentions prices, but it does not explain ordering, time series requirements, or whether values should be adjusted. Basic semantics are implied but not explicit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the operation: calculate maximum drawdown from prices. It is specific enough to distinguish from sibling tools like calculate_returns and calculate_volatility.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to use this tool versus alternatives, nor any conditions or prerequisites. The description only states what it does, not when it should be chosen.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

calculate_returnsB

Calculate log returns from a price series.

ParametersJSON Schema
NameRequiredDescriptionDefault
pricesYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries full responsibility for disclosing behavior. It does not mention assumptions, edge cases (e.g., zero prices, ordering), or side effects, leaving the agent unaware of potential pitfalls.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence, highly concise and to the point. No unnecessary words or redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple calculation tool, the description gives the core function, but it lacks important context such as output format (even though an output schema exists, it's not shown) and potential numerical issues. It is minimally sufficient but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage for the 'prices' parameter. The description does not explain what the prices represent (e.g., chronological order, currency) or expected format, failing to compensate for the missing schema documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: calculating log returns from a price series. This is specific and distinguishes it from sibling tools like calculate_volatility and calculate_max_drawdown.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The usage is implied by the name and description (use when log returns are needed), but there is no explicit guidance on when to use this tool versus alternatives. No comparison to siblings is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

calculate_volatilityA

Calculate rolling volatility from prices.

ParametersJSON Schema
NameRequiredDescriptionDefault
pricesYes
windowNo
annualizeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden for behavioral transparency. It states the action but does not disclose any underlying assumptions (e.g., method of volatility calculation, handling of insufficient data, or whether annualization uses square-root-of-time rule). It is transparent about the lack of side effects but omits edge-case behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that immediately conveys the core functionality. It avoids unnecessary words and is well-structured for quick comprehension.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete for triggering a basic volatility calculation but lacks details about the output format or expected result (despite the tool having an output schema, it is not shown). It also does not specify any constraints like minimum number of prices required, which could affect proper invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides names and types but no descriptions. The description only says 'from prices' and does not explain the meaning of 'window' or 'annualize'. For example, it is unclear whether 'window' is the rolling window length in periods or something else. This leaves significant room for misinterpretation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Calculate' and the target resource 'volatility' from prices, and explicitly mentions 'rolling' which differentiates it from sibling tools like calculate_returns and calculate_max_drawdown. This makes the tool's purpose immediately identifiable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for volatility calculations but does not explicitly state when to prefer this tool over siblings, such as when to choose volatility over returns or drawdown. No exclusion criteria or context are provided, leaving some ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_company_fundamentalsC

Get reported company facts with provider provenance.

ParametersJSON Schema
NameRequiredDescriptionDefault
tickerYes
providerNoauto

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of disclosing side effects, permissions, or data access. It simply says 'Get' without mentioning whether this is a read-only operation, any rate limits, or dependencies on provider availability.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that directly states the tool's purpose with no unnecessary words or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with only two parameters and no output schema, the description is minimally adequate but lacks key details such as what output format to expect, what 'company facts' encompass, and how provider provenance is reflected in the result. This could lead to confusion in edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, and the tool description does not explain the purpose or expected format of the 'ticker' or 'provider' parameters. While 'ticker' may be self-explanatory, 'provider' is ambiguous, and the description fails to compensate for the lack of schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action ('Get') and a resource ('reported company facts'), with 'provider provenance' hinting at a differentiator from sibling tools like get_market_data. However, 'company facts' is somewhat vague without elaboration on what types of facts (e.g., financials, fundamentals).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives such as get_market_data or search_news. The description only states what the tool does, leaving the choice of tool to the agent's inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_market_dataC

Get recent market bars together with provider provenance.

ParametersJSON Schema
NameRequiredDescriptionDefault
daysNo
tickerYes
providerNoauto

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It implies a read-only operation ('Get') but does not explicitly state that it is non-destructive, nor does it mention any side effects, rate limits, or response structure beyond provenance. The description is too thin to give confidence about behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It is appropriately short, though it sacrifices necessary detail for brevity, which is a trade-off rather than a flaw in structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, no annotations, and three parameters, the description is incomplete. It fails to clarify the meaning of 'recent' (tied to the days parameter), the behavior of the provider 'auto' default, or what 'provider provenance' entails. An agent would need to inspect the schema or guess to call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate by explaining the parameters. It does not mention ticker, days, or provider at all, leaving the agent to infer their meaning from names and defaults alone. This is a critical gap for a tool with three parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Get') and the resource ('market bars'), with a specific qualifier ('recent') and a distinguishing feature ('provider provenance'). It does not explicitly differentiate from siblings like get_company_fundamentals, but the resource is distinct enough to be clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives, nor any prerequisites or exclusions. The description simply states what it does, leaving the agent to infer usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_arxivC

Search research papers together with provider provenance.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes
providerNoauto

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided. The description implies a read-only search but does not state side effects, rate limits, or whether it returns cached or live data. The behavior is only partially transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single compact sentence without redundancy. It directly states the core function and a distinguishing detail, achieving high information density in minimal space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description conveys the general purpose and mentions provider provenance, but lacks details on parameter behaviors, result format, and when to prefer this over sibling search tools. It is adequate for a basic understanding but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides no descriptions for query, limit, or provider, and the description does not explain any of them. With 0% schema coverage and no compensating text, the parameters remain completely undefined.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the action (search) and resource (research papers), and adds 'provider provenance' to indicate result attributes. It is distinct from sibling search tools (e.g., search_news, search_documents) but the provenance phrase is slightly ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives like search_news or search_documents. It does not mention any conditions that would favor this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_documentsC

Search local research documents.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It only states the action without explaining how the search works, what it returns, or any constraints. This is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single short sentence, which is concise, but it is under-specified. It lacks the structure and detail needed to be useful, making it more of an under-specification than a concise description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has two parameters and no annotations, yet the description provides no detail about return values, search behavior, or edge cases. An agent cannot confidently call this tool correctly based on the description alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not mention the parameters (query and limit) at all. The agent receives no guidance on how to format the query or what the limit controls.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (search) and the resource (local research documents). It is not a tautology. However, it does not differentiate from sibling tools like search_documents_vector and search_documents_semantic, which likely also search local research documents.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the other search tools (search_news, search_arxiv, search_documents_vector, etc.). The description gives no context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_documents_semanticB

Search local research documents with FAISS embeddings when available.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, and the description only mentions 'when available' without disclosing what happens if FAISS embeddings are not available (e.g., error, fallback, empty results). It also does not mention any auth, side effects, or output format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and to the point, with no unnecessary words or fluff. It fits in a single sentence and effectively communicates the core functionality.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool lacks crucial context such as return format, error handling, and how it differs from sibling search tools. Given the existence of multiple search tools in the sibling list, a brief note on when to prefer semantic search would significantly improve completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema defines 'query' as a required string and 'limit' as an integer with a default of 5, but the description provides no additional meaning. Parameters like 'query' are obvious, but 'limit' could have implications (e.g., max results) that are not explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Search'), the specific resource ('local research documents'), and the method ('FAISS embeddings'), which distinguishes it from a plain full-text search sibling like 'search_documents'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not provide any guidance on when to use this semantic search versus the sibling 'search_documents' or 'search_documents_vector'. The phrase 'when available' hints at a conditional behavior but does not explain fallback or selection criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_documents_vectorB

Search local research documents with vector-space ranking.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description gives no details about side effects, permissions, or return behavior. Since the tool is a search operation, it is reasonable to assume it is read-only, but this is not explicitly stated. The description also does not mention any edge cases (e.g., empty results, handling of malformed queries) which could affect the agent's expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that captures the essential functionality. There is no unnecessary verbosity, and every word adds value. It is well-structured and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description lacks crucial context about the tool's scope and expectations. It does not clarify what 'local research documents' refers to (e.g., a specific corpus, file types, or directory). It also does not hint at the output format, though an output schema exists. The missing parameter semantics and limited usage guidance leave the description incomplete for a robust agent decision.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides no descriptions for the parameters (query and limit), and the tool description also fails to explain them. The description does not mention what constitutes a valid query (e.g., natural language vs. keywords) or what the limit parameter controls (e.g., maximum number of results). With 0% schema coverage and no description, the agent has no guidance on parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: searching local research documents using vector-space ranking. This directly distinguishes it from sibling tools like search_documents (which likely uses keyword matching) and search_documents_semantic (which may use a different semantic approach). The verb 'search' and the specific ranking method make the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly state when to use this tool versus the alternatives. It implies that vector-space ranking is the differentiator, but there is no explicit guidance such as 'use this when you need semantic similarity' or 'use this instead of search_documents for fuzzy queries.' The sibling context provides some inference, but the tool description alone lacks clear usage instructions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_newsB

Search financial news together with provider provenance.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes
tickerNo
providerNoauto

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, and the description only states the action without detailing side effects, return format, or whether it's read-only. The description does not go beyond what the tool name alone suggests.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence that directly states the tool's purpose. No unnecessary words or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple search tool with no output schema, the description is adequate. It implies results include news and provider info. Minor gap: does not specify whether it returns full articles or just metadata, but this is not critical for basic invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema fully covers parameter names, types, defaults, and required status, so the baseline is 3. The description adds no additional meaning to the parameters (e.g., what 'provider' or 'ticker' entail).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool searches financial news and includes provider provenance, distinguishing it from other search tools like arxiv or document searches. The verb 'search' and resource 'financial news' are specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus siblings like search_arxiv or search_documents. The name and description imply news-specific use, but no criteria or conditions are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 10 tool updatesv0.3.0
    • First observedcalculate_max_drawdown
    • First observedcalculate_returns
    • First observedcalculate_volatility
    • First observedget_company_fundamentals
    • First observedget_market_data
    • First observedsearch_arxiv
    • First observedsearch_documents
    • First observedsearch_documents_semantic
    • First observedsearch_documents_vector
    • First observedsearch_news

TDQS

B3/5.0

Scored across 10 tools

Disambiguation3/5

Most tools map cleanly to distinct resources or calculations, but the three search_documents variants overlap considerably and an agent may struggle to choose between plain, vector, and semantic search. Descriptions help somewhat, but the boundaries remain unclear.

Naming Consistency5/5

All tools follow a consistent snake_case verb_noun pattern: get_*, search_*, and calculate_*. This makes tool selection predictable and reinforces the purpose of each tool.

Tool Count4/5

Ten tools is a reasonable size for a financial research server covering data retrieval, document search, and analytics. The three document search variants add mild redundancy but do not make the count excessive.

Completeness4/5

The tool surface covers core research workflows: retrieving market data and fundamentals, searching news/papers/local documents, and computing common return and risk statistics. Minor gaps exist, such as no explicit save/annotate workflow or more advanced analytics, but these are not critical for the apparent scope.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Enables AI agents to operate a local financial terminal, including market data, backtesting, paper portfolio management, and news digest, through safe, gated tools over MCP.
    6
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    A multi-agent financial portfolio risk analyzer that provides MCP tools for fetching market news and storing compliance-reviewed assessment reports in a local database.
    -
  • F
    license
    A
    quality
    B
    maintenance
    Provides multi-agent equity research for US markets with provenance-backed financial data from SEC EDGAR, technicals, macro, and Alpaca paper trading, enforcing risk limits and journaling theses.
    27
    -