Skip to main content
Glama

Search Compounds

pubchem_search_compounds
Read-onlyIdempotent

Search PubChem for compounds by name, SMILES, InChIKey, formula, substructure, or similarity; returns CIDs with pagination and optional properties.

Instructions

Search PubChem for chemical compounds by identifier (name, SMILES, or InChIKey, batched up to 25), molecular formula in Hill notation, substructure or superstructure containment, or 2D Tanimoto similarity. Returns a page of CIDs — reach matches past maxResults with offset. Optionally hydrate results with properties to avoid a follow-up pubchem_get_compound_details call.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
queryNoRequired for substructure/superstructure/similarity searches. A SMILES string (e.g. "CC(=O)O") or PubChem CID as a string (e.g. "2244").
offsetNoZero-based index of the first CID to return. Pass the nextOffset from a previous call to read the following page. Identifier lookups resolve every match up front, so paging them is free; formula, substructure, superstructure, and similarity searches have to ask PubChem for offset + maxResults records to reach a page, so deep pages cost progressively more upstream — hence the 10000 ceiling. Default: 0.
formulaNoRequired for formula search. Molecular formula in Hill notation (e.g. "C6H12O6", "CaH2O2").
queryTypeNoRequired for structure/similarity searches. Format of the query: "smiles" or "cid".
thresholdNoSimilarity search only. Minimum Tanimoto similarity (70-100). 90+ for close analogs, 70-80 for scaffold hops. Default: 90.
maxResultsNoMaximum CIDs to return per page (1-200). Use offset to reach matches past this page. Default: 20.
propertiesNoOptional: fetch these properties for each result, avoiding a follow-up details call. E.g. ["MolecularFormula", "MolecularWeight", "CanonicalSMILES"].
searchTypeYesSearch strategy; each mode needs its own fields. "identifier": name/SMILES/InChIKey lookup — requires identifierType and identifiers. "formula": molecular formula — requires formula. "substructure": find compounds containing the query as a substructure. "superstructure": find compounds that are themselves substructures of the query. "similarity": 2D Tanimoto similarity to the query. substructure, superstructure, and similarity require query and queryType.
identifiersNoRequired for identifier search. Array of identifiers to resolve (1-25). Examples: ["aspirin", "ibuprofen"] for name, ["CC(=O)OC1=CC=CC=C1C(=O)O"] for SMILES, ["BSYNRYMUTXBXSQ-UHFFFAOYSA-N"] for inchikey (27-char block format).
identifierTypeNoRequired for identifier search. Type of chemical identifier: "name", "smiles", or "inchikey".
allowOtherElementsNoFormula search only. When true, includes compounds with additional elements beyond the formula.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
capNoThe maxResults cap that was applied.
errorNoPresent when the call failed. Absent on success.
shownNoCIDs returned on this page.
noticeNoRecovery guidance when no compounds matched, when the offset runs past the matches observed, when identifiers had no match or could not be interpreted, when identifiers collided on one CID, or when further pages remain. Absent when this page is complete and every identifier resolved to its own CID.
offsetNoZero-based index of the first CID returned.
resultsNoMatching compounds.
truncatedNoTrue when matching CIDs remain past this page.
nextOffsetNoOffset to pass on the next call to continue past this page. Omitted when no further matches remain.
searchTypeNoSearch strategy used: identifier, formula, substructure, superstructure, or similarity.
totalFoundNoExact number of matching CIDs across all pages. Omitted when a formula, substructure, superstructure, or similarity search saturated the records it requested — PubChem returns no match count for those, so totalFoundAtLeast reports a floor instead.
totalFoundAtLeastNoLower bound on matching CIDs, reported in place of totalFound when the exact count is unavailable. At least this many match, and the true total may be higher; page further with offset to observe more.
unresolvedIdentifiersNoIdentifier-mode only: input identifiers that resolved to no CID — PubChem had no match, or could not interpret the input as identifierType (the notice says which). Omitted when every identifier resolved and for non-identifier searches.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed5 schema fields changedv0.6.5
    • changedInput schema / properties / searchType / description
      Previous value: -"Search strategy. \"identifier\": name/SMILES/InChIKey lookup. \"formula\": molecular formula. \"substructure\": find compounds containing the query as a substructure. \"superstructure\": find compounds that are themselves substructures of the query. \"similarity\": 2D Tanimoto similarity to the query."New value: +"Search strategy; each mode needs its own fields. \"identifier\": name/SMILES/InChIKey lookup — requires identifierType and identifiers. \"formula\": molecular formula — requires formula. \"substructure\": find compounds containing the query as a substructure. \"superstructure\": find compounds that are themselves substructures of the query. \"similarity\": 2D Tanimoto similarity to the query. substructure, superstructure, and similarity require query and queryType."
    • changedOutput schema / properties / error / properties / data / properties / reason / description
      Previous value: -"Machine-readable failure mode. Declared by this tool: `missing_identifier_args`: searchType is \"identifier\" but identifierType or identifiers were omitted. `missing_formula`: searchType is \"formula\" but the formula field was omitted. `missing_structure_args`: substructure/superstructure/similarity search missing query or queryType. `invalid_cid_query`: structure/similarity search with queryType \"cid\" but query is not a positive integer CID. Other values are possible when a failure originates below the handler."New value: +"Machine-readable failure mode. Declared by this tool: `missing_identifier_args`: searchType is \"identifier\" but identifierType or identifiers were omitted. `missing_formula`: searchType is \"formula\" but the formula field was omitted or blank. `missing_structure_args`: substructure/superstructure/similarity search with query or queryType omitted, or a blank query. `invalid_cid_query`: structure/similarity search with queryType \"cid\" but query is not a positive integer CID. `identifier_rejected`: identifier search where PubChem rejected every identifier in the batch as unreadable (HTTP 400) — for SMILES, a string it could not standardize into a structure. A batch with any other outcome lists rejected inputs in unresolvedIdentifiers instead. `search_query_rejected`: PubChem failed a formula, substructure, superstructure, or similarity search with HTTP 500 \"Search status indicates failure\" — its answer for a malformed SMILES or formula, a SMILES with a \"*\" wildcard atom, or a CID with no record. Other values are possible when a failure originates below the handler."
    • changedOutput schema / properties / error / properties / data / properties / reason / examples
      Previous value: -[
      -  "missing_identifier_args",
      -  "missing_formula",
      -  "missing_structure_args",
      -  "invalid_cid_query"
      -]New value: +[
      +  "missing_identifier_args",
      +  "missing_formula",
      +  "missing_structure_args",
      +  "invalid_cid_query",
      +  "identifier_rejected",
      +  "search_query_rejected"
      +]
    • changedOutput schema / properties / notice / description
      Previous value: -"Recovery guidance when no compounds matched, when the offset runs past the matches observed, when identifiers failed to resolve, or when further pages remain. Absent when this page is complete and every identifier resolved."New value: +"Recovery guidance when no compounds matched, when the offset runs past the matches observed, when identifiers had no match or could not be interpreted, when identifiers collided on one CID, or when further pages remain. Absent when this page is complete and every identifier resolved to its own CID."
    • changedOutput schema / properties / unresolvedIdentifiers / description
      Previous value: -"Identifier-mode only: input identifiers that resolved to no CID. Omitted when every identifier resolved and for non-identifier searches."New value: +"Identifier-mode only: input identifiers that resolved to no CID — PubChem had no match, or could not interpret the input as identifierType (the notice says which). Omitted when every identifier resolved and for non-identifier searches."
  2. Changed1 schema field changedv0.6.3
    • changedOutput schema / properties / error / properties / data / properties / reason / description
      Previous value: -"Machine-readable failure mode. Declared by this tool: `missing_identifier_args`: searchType is \"identifier\" but identifierType or identifiers were omitted `missing_formula`: searchType is \"formula\" but the formula field was omitted `missing_structure_args`: substructure/superstructure/similarity search missing query or queryType `invalid_cid_query`: structure/similarity search with queryType \"cid\" but query is not a positive integer CID Other values are possible when a failure originates below the handler."New value: +"Machine-readable failure mode. Declared by this tool: `missing_identifier_args`: searchType is \"identifier\" but identifierType or identifiers were omitted. `missing_formula`: searchType is \"formula\" but the formula field was omitted. `missing_structure_args`: substructure/superstructure/similarity search missing query or queryType. `invalid_cid_query`: structure/similarity search with queryType \"cid\" but query is not a positive integer CID. Other values are possible when a failure originates below the handler."
  3. Changed6 schema fields changedv0.6.1
    • changedInput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
    • addedInput schema / additionalProperties
      Added value: +false
    • changedOutput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
    • addedOutput schema / anyOf
      Added value: +[
      +  {
      +    "not": {
      +      "required": [
      +        "error"
      +      ]
      +    },
      +    "required": [
      +      "results",
      +      "searchType",
      +      "offset"
      +    ]
      +  },
      +  {
      +    "required": [
      +      "error"
      +    ]
      +  }
      +]
    • addedOutput schema / properties / error
      Added value: +{
      +  "additionalProperties": {},
      +  "description": "Present when the call failed. Absent on success.",
      +  "properties": {
      +    "code": {
      +      "description": "JSON-RPC error code for this failure.",
      +      "maximum": 9007199254740991,
      +      "minimum": -9007199254740991,
      +      "type": "integer"
      +    },
      +    "data": {
      +      "additionalProperties": {},
      +      "properties": {
      +        "reason": {
      +          "description": "Machine-readable failure mode. Declared by this tool: `missing_identifier_args`: searchType is \"identifier\" but identifierType or identifiers were omitted `missing_formula`: searchType is \"formula\" but the formula field was omitted `missing_structure_args`: substructure/superstructure/similarity search missing query or queryType `invalid_cid_query`: structure/similarity search with queryType \"cid\" but query is not a positive integer CID Other values are possible when a failure originates below the handler.",
      +          "examples": [
      +            "missing_identifier_args",
      +            "missing_formula",
      +            "missing_structure_args",
      +            "invalid_cid_query"
      +          ],
      +          "type": "string"
      +        },
      +        "recovery": {
      +          "additionalProperties": {},
      +          "description": "Actionable next step for the caller.",
      +          "properties": {
      +            "hint": {
      +              "type": "string"
      +            }
      +          },
      +          "required": [
      +            "hint"
      +          ],
      +          "type": "object"
      +        },
      +        "retryable": {
      +          "description": "Whether retrying may succeed.",
      +          "type": "boolean"
      +        }
      +      },
      +      "type": "object"
      +    },
      +    "message": {
      +      "description": "Human-readable description of what went wrong.",
      +      "type": "string"
      +    }
      +  },
      +  "required": [
      +    "code",
      +    "message"
      +  ],
      +  "type": "object"
      +}
    • removedOutput schema / required
      Removed value: -[
      -  "results",
      -  "searchType",
      -  "offset"
      -]
  4. Changed11 schema fields changedv0.6.0
    • changedInput schema / properties / maxResults / description
      Previous value: -"Maximum CIDs to return (1-200). Default: 20."New value: +"Maximum CIDs to return per page (1-200). Use offset to reach matches past this page. Default: 20."
    • changedInput schema / properties / maxResults / type
      Previous value: -"number"New value: +"integer"
    • addedInput schema / properties / offset
      Added value: +{
      +  "default": 0,
      +  "description": "Zero-based index of the first CID to return. Pass the nextOffset from a previous call to read the following page. Identifier lookups resolve every match up front, so paging them is free; formula, substructure, superstructure, and similarity searches have to ask PubChem for offset + maxResults records to reach a page, so deep pages cost progressively more upstream — hence the 10000 ceiling. Default: 0.",
      +  "maximum": 10000,
      +  "minimum": 0,
      +  "type": "integer"
      +}
    • addedOutput schema / properties / nextOffset
      Added value: +{
      +  "description": "Offset to pass on the next call to continue past this page. Omitted when no further matches remain.",
      +  "type": "number"
      +}
    • changedOutput schema / properties / notice / description
      Previous value: -"Recovery guidance when no compounds matched — echoes search strategy and suggests how to broaden. Absent when results were returned."New value: +"Recovery guidance when no compounds matched, when the offset runs past the matches observed, when identifiers failed to resolve, or when further pages remain. Absent when this page is complete and every identifier resolved."
    • addedOutput schema / properties / offset
      Added value: +{
      +  "description": "Zero-based index of the first CID returned.",
      +  "type": "number"
      +}
    • changedOutput schema / properties / shown / description
      Previous value: -"CIDs returned after the maxResults cap."New value: +"CIDs returned on this page."
    • changedOutput schema / properties / totalFound / description
      Previous value: -"Total CIDs found before the maxResults cap."New value: +"Exact number of matching CIDs across all pages. Omitted when a formula, substructure, superstructure, or similarity search saturated the records it requested — PubChem returns no match count for those, so totalFoundAtLeast reports a floor instead."
    • addedOutput schema / properties / totalFoundAtLeast
      Added value: +{
      +  "description": "Lower bound on matching CIDs, reported in place of totalFound when the exact count is unavailable. At least this many match, and the true total may be higher; page further with offset to observe more.",
      +  "type": "number"
      +}
    • changedOutput schema / properties / truncated / description
      Previous value: -"True when CIDs were capped at maxResults — more matches exist than returned."New value: +"True when matching CIDs remain past this page."
    • changedOutput schema / required
      Previous value: -[
      -  "results",
      -  "searchType",
      -  "totalFound"
      -]New value: +[
      +  "results",
      +  "searchType",
      +  "offset"
      +]
  5. Changed1 schema field changedv0.2.5
    • addedOutput schema / properties / unresolvedIdentifiers
      Added value: +{
      +  "description": "Identifier-mode only: input identifiers that resolved to no CID. Omitted when every identifier resolved and for non-identifier searches.",
      +  "items": {
      +    "type": "string"
      +  },
      +  "type": "array"
      +}
  6. Changed3 schema fields changedv0.2.4
    • addedOutput schema / properties / cap
      Added value: +{
      +  "description": "The maxResults cap that was applied.",
      +  "type": "number"
      +}
    • addedOutput schema / properties / shown
      Added value: +{
      +  "description": "CIDs returned after the maxResults cap.",
      +  "type": "number"
      +}
    • addedOutput schema / properties / truncated
      Added value: +{
      +  "description": "True when CIDs were capped at maxResults — more matches exist than returned.",
      +  "type": "boolean"
      +}
  7. Changed9 schema fields changedv0.1.22
    • changedInput schema / properties / identifiers / description
      Previous value: -"Required for identifier search. Array of identifiers to resolve (1-25). Examples: [\"aspirin\", \"ibuprofen\"] for name, [\"CC(=O)OC1=CC=CC=C1C(=O)O\"] for SMILES."New value: +"Required for identifier search. Array of identifiers to resolve (1-25). Examples: [\"aspirin\", \"ibuprofen\"] for name, [\"CC(=O)OC1=CC=CC=C1C(=O)O\"] for SMILES, [\"BSYNRYMUTXBXSQ-UHFFFAOYSA-N\"] for inchikey (27-char block format)."
    • changedInput schema / properties / query / description
      Previous value: -"Required for substructure/superstructure/similarity searches. A SMILES string or PubChem CID (as string) for the query structure."New value: +"Required for substructure/superstructure/similarity searches. A SMILES string (e.g. \"CC(=O)O\") or PubChem CID as a string (e.g. \"2244\")."
    • changedInput schema / properties / searchType / description
      Previous value: -"Search strategy: \"identifier\" (name/SMILES/InChIKey lookup), \"formula\", \"substructure\", \"superstructure\", or \"similarity\"."New value: +"Search strategy. \"identifier\": name/SMILES/InChIKey lookup. \"formula\": molecular formula. \"substructure\": find compounds containing the query as a substructure. \"superstructure\": find compounds that are themselves substructures of the query. \"similarity\": 2D Tanimoto similarity to the query."
    • addedOutput schema / properties / notice
      Added value: +{
      +  "description": "Recovery guidance when no compounds matched — echoes search strategy and suggests how to broaden. Absent when results were returned.",
      +  "type": "string"
      +}
    • addedOutput schema / properties / results / items / description
      Added value: +"Matching compound entry."
    • changedOutput schema / properties / results / items / properties / properties / description
      Previous value: -"Compound properties when requested."New value: +"Compound properties keyed by name (echoes input.properties; only present when requested)."
    • changedOutput schema / properties / searchType / description
      Previous value: -"The search strategy used."New value: +"Search strategy used: identifier, formula, substructure, superstructure, or similarity."
    • changedOutput schema / properties / totalFound / description
      Previous value: -"Total CIDs found (before maxResults cap)."New value: +"Total CIDs found before the maxResults cap."
    • changedOutput schema / required
      Previous value: -[
      -  "searchType",
      -  "totalFound",
      -  "results"
      -]New value: +[
      +  "results",
      +  "searchType",
      +  "totalFound"
      +]
  8. First observedv0.1.11

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, openWorldHint, and idempotentHint, so the description need not restate these. It adds behavioral context by mentioning pagination via offset and the option to hydrate properties to avoid a follow-up call. No contradictions with annotations exist. It does not discuss rate limits or error conditions, but these are not essential given the read-only, idempotent nature.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is exactly two sentences with no filler. It front-loads the core purpose and search types in the first sentence, then adds pagination and property hydration details in the second. Every phrase contributes to understanding the tool, and there is no redundancy or tangential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with five search modes and 11 parameters, the description provides a high-level overview of all major capabilities, including pagination and optional property hydration. The presence of an output schema and a fully descriptive input schema means the description does not need to repeat parameter details. It covers the essential aspects an agent needs to know to decide when and how to invoke the tool, though it does not delve into edge cases or error scenarios.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all parameters are already fully documented in the input schema. The description adds high-level context like 'batched up to 25' and 'Hill notation' and '2D Tanimoto similarity', but these details are also present in the schema's parameter descriptions. The description adds minimal additional meaning beyond what the schema provides, so a baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states it searches PubChem for chemical compounds and enumerates all five search modes (identifier, formula, substructure, superstructure, similarity). It clearly distinguishes itself from sibling tools like pubchem_get_compound_details by noting it returns CIDs and can optionally hydrate properties to avoid a follow-up call. The verb 'search' plus resource 'PubChem' and specific search types make the purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly mentions that hydrating properties avoids a follow-up pubchem_get_compound_details call, which provides clear guidance on when to use this tool over that sibling. It also implies usage scenarios for each search type through its enumeration, though it does not explicitly state exclusions or when not to use it. The schema further details each search mode's requirements, so the guidance is largely sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.