Skip to main content
Glama

Cell Annotation MCP

New to the project? Start with the 34-slide English introduction, which explains single-cell annotation, existing methods, and why an evidence-aware decision layer is useful. The bilingual speaker script is a separate Markdown document and is not embedded in the slides. See presentation setup and verification for local preview and publishing details.

Cell Annotation MCP is an evidence-traceable Model Context Protocol server for human and mouse single-cell annotation. It resolves Cell Ontology and UBERON terms, searches versioned marker and reference-atlas indexes, ranks candidate cell types, audits proposed labels, adjudicates outputs from multiple methods, plans follow-up marker panels, controls the strongest defensible claim, and exports sealed decision bundles with source and policy provenance.

The project is an operational beta with an end-to-end server, reproducible knowledge-base builds, and layered validation. The default installation packages 215,282 primary marker assertions across 40 normalized source layers. These include Cell Marker Accordion, HRA ASCT+B, Tabula Muris Senis, HPA, HCL, Mouse Cell Atlas, tissue-specific CELLxGENE and GEO branches, and manually reviewed primary-study panels. The newest layer adds the five-component DPP4/PI16/SEMA3C/MFAP5/FBN1 Fibro_4 panel from Heimli/GSE207206/PMID36741401 for pediatric human thymus. A complete panel returns fibroblast (CL:0000057) conservatively as parent_only, three of five components meet the source-local coverage rule, two or fewer components abstain, and adding the source-explicit CD45-depleted stromal-context contradiction PTPRC forces unknown. Every single-marker query abstains. The preceding Henry prostate and Van Zyl eye layers provide similarly bounded positive panels and context-specific contradictions.

A separate 5,451-assertion conservative pan-tissue fallback gives 220,733 records in the provenance-validation view. These assertions are served from a packaged read-only SQLite index whose manifest pins all 39 normalized TSV inputs by SHA-256; the TSV loaders remain available as a compatibility and rebuild path. The package also contains 135,807 official HGNC/MGI gene records, 4,480 CELLxGENE Census organ/cell-type references, 80 PubMed citation records with 1,091 reviewed claims, 1,055 approved panel components across 41 direct evidence families, 36 context/caution claims across 16 families, a 7,628-record citation-only PubMed review queue, a supplementary 266-panel ScTypeDB index, 170 human/mouse Garnett classifier panels, and 49,437 exact-deduplicated human/mouse CellMatch marker-reference rows. Annotation-tool indexes are never inserted into annotation scoring or counted as independent evidence.

These broad indexes are not presented as complete biological evidence. The machine-readable 52-organ release specification now finds organ-specific marker data and literature discovery coverage for all 102 required human/mouse contexts. This closes the presence gate, not the biological-validation gates: many contexts still lack a healthy reference atlas, balanced negative evidence, independent-study replication, held-out evaluation, or calibrated probabilities. For an anatomy outside the required scope with no organ panel, the default annotation path may use an explicitly downgraded pan-tissue fallback and still abstains when support is weak; the fallback never satisfies an organ release gate and can be disabled per call.

Why this project exists

Cell annotation is rarely just a marker lookup. A defensible answer needs biological context, canonical identifiers, positive and negative evidence, source versions, and an explicit way to abstain. This project keeps those concerns in the data model and the MCP response instead of leaving them to an unconstrained LLM.

Current capabilities include:

  • a strict registry of 165 annotation databases, atlases, repositories, models, and literature sources, with machine-readable separation between 40 loaded marker-evidence sources, 16 loaded reference-index sources, and 109 catalog-only candidates;

  • a packaged Cell Ontology 2026-06-08 index with 3,540 terms and 5,063 normalized edges;

  • a packaged UBERON 2026-06-19 index with 16,071 terms and 33,946 normalized relationships;

  • a packaged official gene-nomenclature index with 45,031 HGNC and 90,776 MGI Gene/Pseudogene records, stable IDs, aliases, previous human symbols, and available Ensembl/Entrez cross-references;

  • a pinned full healthy Cell Marker Accordion corpus with 178,702 positive and 1,278 negative assertions;

  • an official HRA ASCT+B v2.5 human corpus with 10,617 positive gene/protein panel components from 31 biomarker-bearing organ tables, normalized to 755 UBERON tissues, 536 CL cell types, and 1,311 HGNC genes;

  • an official Tabula Muris Senis v1 mouse FACS/3m layer with 3,330 quantitative positive panel components across 21 normalized tissues and 111 tissue/cell-type panels, all tied to MGI, CL, UBERON, raw-file checksums, the source paper, and one study evidence family;

  • an official HPA v25.1 human single-cell layer with 11,523 quantitative positive panel components across 35 normalized tissues and 401 tissue/cell-type panels, preserving 23 upstream publication families rather than counting HPA as independent replication;

  • a pinned CC BY 4.0 HCL v4 human layer with 224 quantitative positive panel components across adult vermiform appendix and gallbladder, nine tissue/cell-type panels, exact HGNC/CL/UBERON provenance, and explicit one-donor/one-batch limitations;

  • pinned CC BY 4.0 Mouse Cell Atlas v8 Figure 2 and detailed adult-testis branches with 2,178 positive mouse panel components across 18 normalized tissues and 76 tissue/cell-type panels, with strict anatomical quarantine, two-batch testis support, label-leakage exclusion, and one-study accounting;

  • a pinned healthy adult-mouse gingiva CELLxGENE dataset-version layer with 240 positive panel components across eight exact author/CL cell types, raw-count derivation, exact MGI/CL/UBERON provenance, and explicit one-pooled-replicate/three-animal limitations;

  • a pinned mouse caecum CELLxGENE dataset-version layer with 420 positive immune/stromal panel components across 14 CL cell types, raw-count derivation, semantic quarantine, exact MGI/CL/UBERON provenance, and explicit pooled H. hepaticus-colonized experimental limitations;

  • a pinned human oral CELLxGENE dataset-version layer restricted to 27,289 healthy GSE164241 buccal-mucosa cells from seven donors, with 394 positive donor-consistent panel components across 14 CL cell types, exact HGNC/CL/UBERON provenance, semantic quarantine, and a runtime license-review boundary;

  • a pinned healthy human nasal-cavity CELLxGENE dataset-version layer restricted to 102,060 primary cells from 37 donors and 38 samples, with 863 donor-consistent positive components across 29 CL panels, 12,130 state-like or unsafe-mapping cells quarantined, post-COVID cells excluded, and a runtime license-review boundary;

  • a pinned healthy human nasopharynx CELLxGENE dataset-version layer restricted to 8,874 exact normal, PCR-negative, WHO-0 cells from 15 participants, with 240 participant-recurrent positive components across eight stable CL panels, 2,767 disease/state/blood/low-support cells quarantined, and final-article/preprint dependency recorded as one evidence family;

  • a pinned healthy naive-WT GSE245074 mouse nasal-cavity layer with 1,020 positive components across 34 CL panels, raw-count derivation, recurrence in at least three of four mice, 2,194 quarantined cells, exact MGI/CL/UBERON provenance, FACS frequency-bias disclosure, and a runtime license-review boundary;

  • a pinned physiological-estrus mouse oviduct CELLxGENE dataset-version layer with 180 positive panel components across six admitted CL cell types, raw-count derivation, exact MGI/CL/UBERON provenance, one-sample/five-pooled-mouse limitations, and a runtime license-review boundary;

  • a pinned CC BY 4.0 mouse seminal-vesicle primary-study layer with 313 positive panel components across 12 conservative CL mappings, 20 manually reviewed row-located claims, four author-cluster quarantines, and explicit mixed-cohort and one-family limitations;

  • a pinned CC BY 4.0 mouse-gallbladder GSE179524 layer with 32 positive author-marker figure components across five conservative CL panels, 32 figure- or Results-located reviewed claims, and explicit mixed-condition, one-family, non-quantitative boundaries;

  • a pinned GSE254855 healthy adult-mouse esophagus layer with 260 batch-aware quantitative components across 13 conservative CL panels, 39 reviewed project-recomputed panel claims, a stomach-cell quarantine, and explicit one-family, positive-only, license-review boundaries;

  • a pinned MIT-licensed GSE150327 adult-mouse submandibular-gland layer with 226 quantitative components across eight canonical CL panels, 24 reviewed project-recomputed panel claims, exact UBERON ancestry, and explicit one-sample, female-only, one-family boundaries;

  • a pinned CC BY 4.0 adult-mouse adrenal-gland layer from PMID 32743779, with 246 cross-platform quantitative components across ten conservative CL panels, 30 reviewed author-supplied differential-expression claims, exact UBERON anatomy, and explicit three-batch/one-family boundaries;

  • a pinned CC BY 4.0 GSE239316 mouse endocrine-system atlas from PMID 40899455, with 2,433 author-DE components across six organs, 84 exact tissue/cell-type panels, 42 CL identities, 252 reviewed Table S2 claims, ten quarantined author panels, and explicit pooled-age/sex/animal/library limitations;

  • a pinned GSE145443 adult-mouse epididymis layer from PMID 32729827, with 180 recomputed positive components across six conservative CL panels, 18 reviewed project-recomputed claims, 84 semantically quarantined cells, vas-deferens exclusion, and explicit one-family, positive-only, license-review boundaries;

  • a pinned CC BY 4.0 adult-mouse vaginal-epithelium layer from PMID 37007744, with ten author-text panel components across basal, later suprabasal, and vagina-squamous identities, ten reviewed claims, three quarantined state or single-marker signals, and explicit no-public-matrix and one-family boundaries;

  • a pinned CC BY 4.0 adult-mouse endometrium layer from PMID 32998907, with 12 positive and three explicitly negative components across luminal and glandular epithelial CL panels, exact Table S1 and Results/Figure locators, and Mouse Cell Atlas family deduplication;

  • a pinned CC BY 4.0 adult-mouse corpus-cavernosum layer from PMID 38856719, with 27 author-marker components across nine conservative CL panels, exact KAP230548 and article provenance, mixed-condition pooled-library disclosure, an LBH panel-only rule, and explicit raw-data-license, recurrence, and validation boundaries;

  • a pinned CC BY 4.0 normal adult-mouse rectum layer from PMID 33980830, with four manually reviewed Figure 2 components for CL:1001595, exact GSE163394/CL/UBERON provenance, wounded and non-rectum branch quarantine, no raw-count reanalysis, and explicit unreported-replicate, license, and validation boundaries;

  • a pinned CC BY-NC 4.0 normal-human parathyroid layer from PMID 38471398, with 30 manually reviewed Figure 1/Table S3 components across seven author major populations, exact OMIX005492/PRJCA020004/HGNC/CL/UBERON provenance, adenoma/subtype/mixed-cluster quarantine, and explicit unavailable-file, academic-use, recurrence, and validation boundaries;

  • a pinned experimental mouse-parathyroid layer from PMID 37379351, GSE232601, GSE232602, and PRJNA972984, with 47 exact Dataset S1 author-DE components across eight broad CL identities, a 50% multi-marker panel gate, chimeric Gcm2-null/mESC-complemented scope, unresolved Cluster 7 quarantine, and an independently published non-cell-resolved spatial corroboration of Pth and Gcm2 from PMID 41025096/GSE298115;

  • a pinned GSE241556 cultured human seminal-vesicle epithelial layer from PMID 37984408, with 15 positive author-DE components recurrent across SVEC clusters 2, 5, and 6, detection in both public seminal-vesicle libraries, exact CL:1001597/UBERON:0000998 provenance, a 50% multi-marker gate, and explicit single-patient, culture-selection, adjacent-cancer-cohort, same-family qPCR, and license-review boundaries;

  • a pinned CC BY 4.0 LungMAP CellCards mouse-bronchus layer from PMID 34936882, with 13 exact-locator expert panel components across basal, club, multiciliated, pulmonary neuroendocrine, bronchial smooth-muscle, and tracheobronchial chondrocyte identities; a 50% same-source multi-marker gate; a context-only goblet-cell signal; and explicit derived-database, one-family, positive-only, non-quantitative, non-independent, non-held-out, and uncalibrated boundaries;

  • a bounded mouse-nasopharynx layer from PMIDs 38200313, 30300386, and 28461487, with four checksum-audited GSE227311 lymphatic-endothelial components, 14 naive-NALT immune components, nine author-explicit negative flow-gate components, eight multi-marker CL panels, three publication families, exact nasopharyngeal-mucosa/NALT UBERON contexts, and explicit subregion, pooling, ontology-gap, license-review, non-held-out, and uncalibrated boundaries;

  • a pinned CC BY 4.0 Human Lung Cell Atlas literature layer from PMID 37291214/PMC10287567, with the official PMC OA package fixed by size and SHA-256, two exact Supplementary Table 6 claims for MFSD2A/AT2 and SOSTDC1/aerocyte identity, one dependent 14-dataset integration family, and no new scoring markers, negative evidence, held-out result, or calibrated confidence;

  • a pinned Muraro healthy adult-human pancreas literature layer from PMID 27693023/PMC5092539/GSE85241, with nine Results-explicit markers verified against exact Supplementary Table S3 workbook rows for alpha, beta, delta, PP, epsilon, acinar, ductal, mesenchymal, and endothelial identities; all four donors and nine claims remain one publication family, the CC BY-NC-ND source files stay local, and the layer adds no scoring marker rows;

  • a pinned CC BY 4.0 Muto healthy adult-human kidney-cortex literature layer from PMID 33850129/PMC8044133/GSE151302, with 13 Results-linked markers verified against exact Supplementary Data 1 rows for major nephron, collecting-duct, podocyte, vascular, stromal, and leukocyte identities; all five samples and 13 claims remain one dependent publication family, including three reused GSE131882 samples, and the layer adds no scoring rows;

  • a pinned Aizarani healthy adult-human liver literature layer from PMID 31292543/PMC6687507/GSE124395, with ten Figure-1-linked markers verified against exact Supplementary Table 1 rows for hepatocyte, cholangiocyte, hepatic sinusoidal and general endothelial, stellate, Kupffer, B, NK, T, and erythrocyte identities; all nine donors and 10,372 cells remain one publication family, publisher and GEO reuse terms remain under review, source payloads stay local, and the layer adds no scoring rows;

  • a pinned Bandyopadhyay healthy adult-human-bone-marrow literature layer from PMID 38714197/PMC11162340/GSE253355/S-EPMC11162340, with twelve exact Supplementary Table S2 differential-expression rows for mesenchymal stem, endothelial, hematopoietic stem, macrophage, B, megakaryocyte, monocyte, neutrophil, osteoblast, plasmacytoid dendritic, plasma, and erythrocyte identities; all assay, donor, and cell subsets remain one publication family, unsafe subtypes are quarantined, publisher and repository reuse limits remain explicit, and the layer adds no scoring rows;

  • a pinned Madissoon healthy adult-human-spleen literature layer from PMID 31892341/PMC6937944/PRJEB31843, with five exact Supplementary Figure S9c labels for macrophage, platelet, B-cell, plasma-cell, and T-cell panel components; all donors, preservation time points, and tissues remain one publication family, figure dot values are not reconstructed, and the layer adds no scoring rows;

  • a pinned Abe metastasis-free adult-human-lymph-node literature layer from PMID 35332263/PMC9033586/EGAD00001008311, with eleven exact Supplementary Tables 3, 4, and 10 differential-expression panel components and one context-only PDPN heterogeneity caution; all samples, tables, cohorts, and validation branches remain one publication family, the nine nodes came from patients with neoplasms rather than healthy volunteers, EGA data remain controlled access and were not downloaded, and the layer adds no scoring rows;

  • pinned Park and Ransick healthy-adult-mouse-kidney layers from PMID 29622724/GSE107585 and PMID 31689386/GSE129798, with 35 exact supplementary-table positive grade-C components across 28 CL identities; every scoring row has a one-to-one reviewed claim, the two studies count as two evidence families, within-study batches/animals/sexes/zones do not create extra votes, and low or missing expression is never converted into a negative marker;

  • a pinned CC BY 4.0 He healthy-adult-mouse-kidney layer from PMID 33837218/PMC8035407/GSE160048, with four exact Supplementary Data 8 mesangial panel components, one explicit CD45-sort Ptprc contradiction, and one quarantined CD31/Pecam1 conflict; all rows, cohorts, assays, and reporter experiments remain one study family, and the source-specific 50% panel guard does not restrict Park/Ransick or other kidney candidates;

  • a bounded CC BY-NC-ND 4.0 Van Zyl nondiseased-adult-human-eye layer from PMID 35858321/PMC9303934/GSE199013, with exact Results-sentence support for conjunctival epithelial PDGFC and LGR6, a context-only NTRK2 exclusion, one publication family across six donors, and explicit no-offline-full-text, no-supplement-reproduction, no-independent-validation, and no-calibration boundaries;

  • a bounded normal-adult-human-prostate Henry layer from PMID 30566875/PMC6411034/GSE120716, with exact Results-sentence support for basal prostate epithelial KRT5 and KRT14, an author-explicit context-only KRT13-low contradiction, canonical HGNC/CL/UBERON provenance, one publication family, publisher-copyright/NIH-public-access-manuscript status, and explicit no-offline-full-text, no-independent-validation, no-held-out, and no-calibration boundaries;

  • a bounded CC BY 4.0 pediatric-human-thymus Heimli layer from PMID 36741401/PMC9895842/GSE207206, with exact Results-sentence support for five Fibro_4 panel components, an author-explicit CD45-depleted stromal-context PTPRC contradiction, canonical HGNC/CL/UBERON provenance, one publication family, a 60% source-local panel rule, hard conflict abstention, and explicit no-offline-full-text, no-independent-validation, no-held-out, and no-calibration boundaries;

  • a pinned healthy P60 mouse-lung Wang layer from PMID 29463737/PMC5877944/GSE106960, with 30 exact Dataset S5 AT1 biomarkers, 30 exact AT2 biomarkers, one article-explicit Aqp5 AT1 component, and two article-explicit Pecam1/Foxj1 contamination exclusions; all 63 scoring rows and 64 reviewed claims remain one study family, and the Dataset S5 finding that Sftpc is detected in 62% of P60 AT1 cells is retained as a quarantined context conflict rather than an AT1 scoring negative;

  • a pinned adult-human-brain Siletti/Allen layer from PMID 37824663, author-rule commit 02e092b4a3f0d294d4bcd8ab574dfc98f51d7981, adult-analysis commit 2b5aa12dffb8cc1bcb3979c409cb6ac907546b2e, and WHB taxonomy release 20240330, with 40 positive and six explicit-negative grade-C components across 16 major CL classes; the layer preserves rule presence/absence semantics, quarantines three unused rules, packages no raw expression or patient data, and shares one upstream publication family with the existing HPA brain branch;

  • a pinned La Manno developing-mouse-brain layer from PMID 34321664, PRJNA637987, author-rule commit c398c285861ac714ec1e8827d1f5ac3661d4b8da, and final Supplementary Table 2, with 31 positive and three explicit-negative grade-C components across 12 exact CL classes; all 215 source rules are audited, six stable ambiguous or unmapped rules are quarantined, and every admitted row shares one upstream publication family;

  • a separate grade-C pan-tissue fallback containing 5,451 positive assertions recurrent across at least three source tissues and two upstream evidence families;

  • a 52-organ human/mouse coverage specification with explicit marker, negative-evidence, evidence-family, and reviewed-literature gates;

  • an offline CELLxGENE snapshot containing 2,102 human/mouse dataset records from 387 public collections, with payload licensing kept per_dataset;

  • a pinned CELLxGENE Census 2025-11-08 index containing 4,480 normal-primary organ/cell-type aggregates, including 1,574 major references, with explicit descendant exclusions that prevent epididymal fat pad from satisfying epididymis and uterine/vaginal tissues from satisfying oviduct;

  • a commit-pinned GPL-3.0 ScTypeDB reference index containing 266 unique human HGNC-normalized positive/negative tool panels across 16 UBERON-mapped source tissues; 136 panels have safe CL mappings, while 130 retain source labels without guessed ontology IDs; the index is supplementary, non-scoring, non-independent, and excluded from release gates because upstream marker-level citation lineage is unavailable;

  • a commit-pinned GPL-3.0-or-later CellMatch/scCATCH 3.2.2 reference index containing 49,437 exact-deduplicated positive marker rows for human and mouse across 184 source tissues and 353 source cell types; it preserves cancer, condition, three subtype fields, 49,369 species-aware HGNC/MGI mappings, and 2,096 source reference identifiers, including 2,095 numeric PMIDs, while keeping all source labels and unresolved mappings auditable;

  • a commit-pinned MIT Garnett reference index containing 170 hierarchical human/mouse classifier panels for lung, PBMC/blood, and brain/spinal cord, with 405 species-aware HGNC/MGI components, exact file-line locators, 40 exact CL mappings, and 123 source-hierarchy-derived nearest CL ancestors;

  • context-aware, specificity- and frequency-weighted marker ranking;

  • ontology-lineage-aware batch annotation with parent_only, confidence abstention, unrelated near-tie refusal, source-specific panel guards, and a pan-tissue lineage-conflict audit that can force abstention without promoting a pan-tissue label;

  • schema, ontology, and source-provenance validation of annotation results;

  • an offline 80-record PubMed citation index with 1,091 reviewed claims, including 1,055 approved panel components across 41 direct evidence families and 36 context/caution claims across 16 families, plus immediate-source atlas citations and 23 HPA upstream publication families without automatic gene-claim promotion;

  • a deduplicated 7,628-record PubMed citation review queue from 102 non-truncated organ/species queries, with exact query ranks and no abstracts or full text;

  • literature cross-checking that distinguishes metadata, context, tissue-scope mismatches, panel support, and independent evidence families;

  • validation of 10x-style CSC HDF5 matrices and matching metadata;

  • reproducible marker extraction, evidence-family-isolated benchmarks, deterministic open-set/dropout/mixed-profile stress tests, empirical risk-coverage reports, and validation-only conformal prediction sets;

  • a read-only SQLite marker backend with immutable input-manifest verification and a versioned source-panel policy table;

  • PostgreSQL schema and transactional source-registry import tooling.

Related MCP server: SCMCP

MCP interface

Tool

Purpose

list_sources

Search the registry by category, species, scope, priority, verification status, or runtime integration status, and report exact loaded-record counts.

resolve_cell_type

Resolve a CL identifier, label, or synonym against the packaged Cell Ontology release.

resolve_tissue

Resolve an anatomy identifier, label, or synonym against packaged UBERON.

resolve_gene

Resolve an HGNC/MGI stable ID, approved symbol, alias, or previous human symbol without fuzzy matching.

search_cell_markers

Retrieve normalized marker assertions for a cell type and context.

search_annotation_tool_panels

Search supplementary ScTypeDB or Garnett panels by source, species, tissue, cell type, gene, source reference, and polarity without treating them as scoring or independent evidence.

search_annotation_tool_markers

Search supplementary CellMatch rows by species, tissue, cell type, subtype, gene, cancer, condition, resource, or source reference without promoting them to evidence.

get_knowledge_coverage

Audit organ/species coverage counts, gaps, and release gates.

search_literature

Search offline PubMed metadata, discovery tags, and reviewed claims.

search_literature_candidates

Search the organ-stratified citation-only PubMed review queue without promoting candidates to evidence.

search_reference_datasets

Search the offline CELLxGENE collection/dataset metadata snapshot by species, tissue, disease, or text.

search_reference_cell_types

Search normal-primary CELLxGENE Census organ/cell-type aggregates and their dataset and collection support.

rank_candidate_cell_types

Rank candidates from an observed marker list.

annotate_clusters

Annotate one or more clusters and validate every result.

audit_annotation

Audit a proposed label against ranked support, explicit contradictions, ontology-related alternatives, source-panel coverage, and deduplicated evidence families.

explain_abstention

Report the conservative gates that withheld a call and the next evidence to collect.

adapt_annotation_run

Convert bounded CellTypist, SingleR, Azimuth, manual, or generic JSON records into a checksummed, CL-normalized exchange with an auditable row-level mapping trace.

compare_annotation_runs

Align cluster calls, distinguish exact agreement from parent/child compatibility and true conflict, and report evidence dependence as independent, dependent, or unknown.

benchmark_annotation_decision_layer

Compare ontology- and dependency-aware adjudication with naive exact-ID majority voting under explicit external-test qualification gates.

suggest_discriminating_genes

Return source-attributed genes that separate two to five candidates without treating missing assertions as negative evidence.

plan_disambiguation

Select a compact marker panel with positive, explicit-negative, assay-support, and unresolved-pair output.

get_annotation_readiness

Report whether a species-tissue-cell-type context permits an exact uncalibrated label, a parent-only label, abstention, or a scope-qualified calibrated claim.

build_annotation_decision_bundle

Run annotation, merge optional normalized external calls, and return a checksummed evidence, adjudication, readiness, abstention, and follow-up bundle.

validate_annotation_decision_bundle

Validate bundle structure, ontology identities, cluster alignment, internal annotation provenance, and canonical checksum.

validate_annotation_result

Validate an external or LLM-generated annotation against the output contract.

cross_check_annotation_literature

Revalidate an annotation, verify citation metadata, and match reviewed literature claims.

The server also exposes:

  • source://{source_id} for source metadata;

  • schema://annotation-result for the versioned annotation JSON Schema;

  • schema://annotation-decision-bundle for decision bundle schema 1.0;

  • schema://annotation-run-exchange for external-run exchange schema 1.0.

list_sources and source://{source_id} never equate registry membership with runtime evidence. Each source includes an integration object whose status is one of marker_evidence_integrated, reference_index_integrated, or catalog_only. The object reports runtime roles, primary positive and negative marker counts, a derived-fallback count, reference-index record counts, and an evidence boundary. Use the integration_status argument to select one group. The current build has 56 runtime-integrated sources and 109 catalog-only entries. Catalog-only entries cannot support annotation, and the 5,451 derived fallback rows are explicitly not independent evidence.

The audit tools use evidence_family, not source count, as the primary independence unit. compare_annotation_runs also detects a shared declared reference. Missing family or reference metadata remains unknown, never independent confirmation. External confidence values are preserved but are not averaged because tools do not share a calibrated scale. suggest_discriminating_genes quarantines polarity conflicts and labels missing assertions as not_asserted_for, never as negative evidence.

Annotation Decision Layer v1

adapt_annotation_run accepts bounded JSON records exported by CellTypist, SingleR, Azimuth, manual review, or a generic caller. It selects documented default columns or caller-supplied column names, resolves labels exactly to the packaged Cell Ontology, preserves source labels and scores in adapter_trace, converts unresolved labels to explicit abstentions by default, and seals the result with a canonical SHA-256 checksum. The exchange records method, model, reference version, evidence families, column mapping, and audit counts. Its confidence values remain method-local.

The server neither imports nor executes third-party annotators and does not read arbitrary server-side paths. A client runs each tool in its own environment and sends its exported records to the MCP. A sealed exchange can then be supplied directly wherever a normalized run is accepted, including compare_annotation_runs, build_annotation_decision_bundle, and benchmark_annotation_decision_layer.

build_annotation_decision_bundle accepts marker clusters plus zero or more normalized runs or sealed exchanges. An external run declares a run_id, method and optional model version, optional reference identity/version/evidence families, and cluster annotations containing a CL ID or resolvable label.

The returned schema 1.0 bundle contains:

  • original markers and gene-normalization audit;

  • candidate ranking, positive and explicit-negative evidence, source releases, PMIDs, accessions, and reviewed literature checks;

  • normalized run calls and ontology-aware adjudication;

  • an abstention explanation when the internal engine withholds a label;

  • context readiness and the strongest allowed claim;

  • a compact disambiguation plan with assay evidence and unresolved pairs;

  • package, ontology, runtime-schema, source-manifest, policy, and calibrator metadata;

  • a canonical SHA-256 checksum that fails validation after payload changes.

Exact agreement, parent/child compatibility, incomplete run alignment, and unrelated conflicts lead to different recommendations. If any method omits a cluster, that cluster remains incomplete and cannot be promoted to consensus. Parent/child calls report the safest term already called by a method; unrelated calls remain unresolved. The disambiguation panel uses deterministic greedy set cover. It is an evidence heuristic, not measured expected information gain or guaranteed assay performance.

Readiness separates data availability from claim qualification. Multiple exact positive markers can permit exact_uncalibrated; related-only evidence permits parent_only; inadequate evidence requires abstain. exact_calibrated additionally requires a supported calibrator contract, independent-study holdout, an explicitly separate final test cohort, production qualification, and an exact species/tissue/cell-type scope match. External cohorts and reviewed negatives therefore remain release evidence: they do not become MCP features, but they control whether the MCP may upgrade its claim. The committed multi-method example is a synthetic contract fixture, not a real CellTypist or SingleR benchmark.

The decision-layer benchmark reports exact, safe-parent, over-specific, unrelated, abstention, coverage, unsafe-claim-rate, and selective-safe-accuracy counts for both adjudication and naive exact-ID majority voting. A report is external_test_qualified only when its reference metadata declares independent audit, study-level holdout, and separation from policy development. Otherwise it is engineering_only, even if its deterministic results are favorable. The committed adapter and benchmark example exercises exact consensus, safe-parent fallback, and conflict abstention; it is deliberately synthetic and establishes contract behavior, not biological accuracy.

Annotation-tool reference indexes

search_annotation_tool_panels currently exposes the official ScTypeDB workbook at repository commit 630e15cf1e51f2612eda4ad0406dfb17503fa8c9. The reproducible adapter verifies the workbook and GPL-3.0 license checksums, maps all 16 source tissues to UBERON, normalizes genes only through the pinned human HGNC release, collapses two exact duplicate source rows, and preserves every unresolved gene or cell label. It emits 266 unique panels with 4,228 resolved positive and 27 resolved negative components. The source workbook has no species column, so the runtime index is human-only; it does not uppercase, ortholog-map, or otherwise transfer panels to mouse.

This index is intentionally separate from search_cell_markers. ScTypeDB is a derived annotation-tool compendium and does not provide per-marker upstream citation lineage. Its result therefore states that annotation scoring, independent-evidence credit, release-gate credit, single-marker annotation, and mouse transfer are disabled. Use it to inspect whether a tool panel agrees with primary tissue-matched evidence, never to raise confidence by counting the same upstream knowledge twice.

The same panel tool exposes source="garnett" for the four human/mouse marker files distributed with Garnett's pre-trained classifier catalog. The acquisition pins marker commit 50954f24e99fb2e553e0200280c280fc05b1e152, package/license commit ad3e1f7aa913b74899afb8aab7e1e55c224a378e, the exact catalog, source files, package metadata, citation file, and MIT notice. It emits 170 panels and preserves all 406 input positive components: 405 resolve through the appropriate HGNC or MGI release, while ND3 remains an unresolved mouse source symbol rather than being coerced. Forty source cell labels map by a unique exact CL label or synonym. For 123 panels, the response also reports the nearest CL-mapped ancestor reached solely by following Garnett's declared subtype of chain; this ancestor is never presented as an exact subtype mapping.

Garnett rules remain derived classifier knowledge. Marker-file links, training-data links, and the method paper are typed by scope, marked as source-provided but not project-reviewed, and accompanied by exact source-file line locators. The index has no negative markers and cannot change scoring, confidence, evidence-family counts, release gates, or an unknown result. Some rules contain a single gene, but the MCP still requires multi-marker and exclusion-marker corroboration before annotation.

search_annotation_tool_markers exposes the official CellMatch object bundled with scCATCH 3.2.2 at repository commit 06c6ffb960795d028c675643f0d22fbc78080a97. The immutable acquisition pins the RDA, package description, documentation, and GPL-3.0-or-later license. The deterministic adapter validates 49,560 source rows, collapses 123 exact duplicates while retaining every original row number, and emits 49,437 records: 29,711 human and 19,726 mouse. It preserves 184 source tissues, 353 source cell types, cancer and condition context, three subtype fields, resource type, and all 2,096 source reference labels. Of those labels, 2,095 are numeric PMIDs; the remaining CD Handbook label is returned as source text without a fabricated PubMed URL.

CellMatch gene identities are resolved separately for human and mouse through the pinned HGNC/MGI store, giving 49,369 mapped records. Only unique normalized exact UBERON or CL labels and synonyms are admitted; ambiguous or unresolved labels remain source-only. CellMatch is a derived, positive-only annotation-tool database. Its source-provided reference identifiers have not been reviewed against exact passages by this project, so its records provide discovery and comparison only: no scoring, independent-family credit, reviewed-literature credit, negative-marker inference, release-gate contribution, single-marker decision, or confidence change is allowed.

CellTypist remains catalog-only. Its software repository is MIT-licensed, but the official model payloads do not state a separate model-data or derivative-redistribution license, and the hosting institute's default content terms are not suitable for repackaging model-derived marker coefficients. No CellTypist model weight or derivative is distributed until a source-specific permission is recorded.

CellGuide also remains catalog-only. The official portal code is pinned at f3473f0cd8edf4ca78eb44ce3ee55a29de50a904; its frontend reads versioned public snapshots and the 2026-08-18 audit observed pointer 1764612212, but the official OpenAPI comments out the marker GET route. The repository MIT license covers software, while the current site terms grant limited access/use without a general marker-payload redistribution license. CellGuide canonical markers are compiled from HRA ASCT+B, already integrated here as the upstream family, and its computational markers derive from CELLxGENE/Census data. The project can use the live view for source verification under its terms, but it will not bundle the payload or count it as independent evidence.

See Annotation-tool source audit for the distinction between a tool-owned knowledge payload, a tutorial fixture, a user-supplied marker interface, and an upstream-derived database with unresolved redistribution terms.

Install from source

Python 3.11 or newer is required and is the tested runtime baseline.

git clone https://github.com/haoyunLi/CellAnnotation-MCP.git
cd CellAnnotation-MCP
python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e .
cell-annotation-admin build-runtime-database

The Git checkout includes the normalized evidence tables and their manifests. The generated SQLite index is ignored by Git; the final command builds it locally from those tables without downloading raw datasets. Run it before starting the server, running make check or make full-check, or building a distribution wheel. If the index already exists and needs rebuilding, use cell-annotation-admin build-runtime-database --overwrite.

The evidence files retain their exact bytes through .gitattributes so Git line-ending conversion cannot invalidate their recorded checksums.

For a source checkout with the runtime dependencies installed, use make docs-check mcp-smoke to verify documentation and MCP communication after building the index. make check and make full-check also audit original source acquisitions and require the corresponding local data/raw/ snapshots. Those snapshots are not needed to run the MCP server and are not distributed through Git.

Install optional data-validation or PostgreSQL support when needed:

python -m pip install -e ".[validation]"
python -m pip install -e ".[database]"
python -m pip install -e ".[adapters]"
python -m pip install -e ".[census]"  # maintainers rebuilding the Census index

On the development HPC, load the matching Python module before using the virtual environment:

module load python/3.11.11
source .venv/bin/activate

Configure an MCP client

The default transport is stdio.

{
  "mcpServers": {
    "cell-annotation": {
      "command": "/absolute/path/to/CellAnnotation-MCP/.venv/bin/python",
      "args": ["-m", "cell_annotation_mcp"],
      "cwd": "/absolute/path/to/CellAnnotation-MCP"
    }
  }
}

You can also start the server directly:

cell-annotation-mcp

For development over Streamable HTTP:

CELL_ANNOTATION_TRANSPORT=streamable-http cell-annotation-mcp

Marker evidence

The primary marker store combines 38 separately versioned normalized source layers. The first is the deterministic admissible healthy portion of the Cell Marker Accordion v1.0.0 workbook:

  • species: human and mouse;

  • contexts: 319 human/mouse species-tissue combinations represented by 270 canonical UBERON terms;

  • 179,980 assertions, including 1,278 negative markers;

  • immediate source: cell_marker_accordion;

  • upstream dependency recorded per assertion as evidence_family;

  • pinned commit, input/output checksums, filters, counts, and limitations recorded in marker_assertions.manifest.json.

The adapter updates source labels to the packaged CL release and excludes obsolete CL terms and disease sheets. At runtime, source and query symbols are resolved conservatively against pinned HGNC and MGI releases and scored by stable gene ID when resolution is unique. Exact case-sensitive symbols take precedence, unique aliases and previous symbols are accepted, fuzzy matching is disabled, and ambiguous inputs are quarantined. Unresolved values can match only the same literal case-folded symbol and are reported as unresolved. The runtime also validates every marker tissue against packaged UBERON and expands an organ query only through is_a and part_of ancestry. It does not silently use an unrelated organ panel.

The second primary artifact is HRA ASCT+B v2.5, released 2026-06-15 under CC BY 4.0. Its reproducible adapter acquires 36 official versioned organ tables and admits 10,617 human positive gene or protein biomarker rows from the 31 tables containing usable biomarkers. Every retained record has a canonical non-obsolete CL term, a canonical non-obsolete UBERON context, a uniquely resolved HGNC stable ID, an HBM digital-object identifier, table DOI, table version, and source-row locator. The adapter quarantines 18 ambiguous gene values, 296 source-ID/symbol conflicts, 422 unresolved genes, 618 rows whose deepest cell type is not an admitted CL term, and 138 rows without an admitted UBERON context. HRA supplies panel components, not single-marker proofs, negative markers, mouse evidence, direct primary-literature claims, or 36 independent studies. All rows conservatively share one evidence family.

For an organ context supported only by HRA, annotation requires the normal multi-gene threshold, at least 5% of the proposed reference panel, and prefers a supported CL ancestor whose score is at least 90% of a more specific proposal. The response exposes hra_asctb_expert_panel, the shared evidence family, and these decision thresholds. Contexts that also contain Cell Marker Accordion evidence use the ordinary combined-source policy.

The third primary artifact is derived from the official Tabula Muris Senis Figshare item 12654728, version 1, released under the MIT license and linked to PMID 32669714. The acquisition adapter pins and verifies all 23 official FACS H5AD files (5.06 GB locally); the normalized build uses only 3-month cells, raw integer counts, canonical CL/UBERON terms, and unique MGI identities. It admitted 38,357 of 44,518 age-selected cells and quarantined 6,161 cells with unresolved, ambiguous, or ontology-unrelated source ID/label combinations. The distributed artifact contains no cell-level matrix or donor identifiers.

For every admitted tissue/cell-type group, marker derivation requires at least 30 cells, two eligible mice, marker detection in at least two mice, 25% within-group sensitivity, 60% specificity against other admitted cells in the same tissue, a 0.20 detection-fraction difference, log2 fold change of at least 1.5, and group pseudobulk CPM of at least 1. Up to 30 markers are retained as grade-C panel_component assertions. All files and tissues share tabula_muris_senis_figshare_12654728_v1; they are one study, not 23 independent studies. The adapter provides no negative-marker semantics and does not turn a computational marker into a reviewed paper claim.

When a mouse organ is supported only by this atlas layer, annotation requires the normal multi-gene threshold, at least 10% of the 30-marker panel, and a 0.90 supported-ancestor preference. Responses expose the age, technology, evidence grade, evidence family, donor-consistency rule, and non-independence boundary. Tissue routing is species-aware, so the presence of a mouse panel cannot weaken an HRA-only human policy for the same UBERON organ.

The fourth primary artifact is built from the official HPA v25.1 single-cell cluster downloads. Its immutable acquisition contains five versioned HPA files (213 MB), 36 upstream tissue datasets, and 23 unique source-paper families. The adapter reads 23,542,642 gene/cluster expression rows, retains only clusters that HPA marks as high-reliability and included in aggregation, maps source tissues, cell types, and genes to canonical UBERON, CL, and HGNC identities, and emits 11,523 positive grade-C panel components across 35 normalized tissues, 139 CL types, 4,233 genes, and 401 tissue/cell-type panels.

HPA marker admission requires at least 30 target and background cells, target nCPM of 5, expression in at least 50% of target-cluster cells at nCPM 1, log2 fold change of 2 against the strongest other mapped type in the same tissue, and specificity of 0.8. Up to 30 genes are retained per panel. HPA is an integration and presentation layer, not an additional independent experiment: each assertion uses the tissue dataset's upstream PubMed family, HPA itself is never counted as another family, and no negative marker or reviewed gene-specific literature claim is created automatically. The distributed package contains normalized marker rows and citation metadata, not cell-level data or raw upstream matrices. HPA content is CC BY 4.0 where HPA owns the copyright, while upstream terms still apply.

Whenever HPA contributes to an organ result, annotate_clusters performs a second, non-promoting pan-tissue lineage audit. If at least three query genes support an ontology-unrelated pan-tissue lineage and its score is at least 1.10 times the organ-specific proposal, the final decision is forced to unknown. The organ proposal remains an alternative and its evidence is retained with context direction. This protects against missing lineages in a tissue atlas—for example, the provided skin regression no longer reports an NK cluster as T cell merely because the HPA skin dataset lacks a matching NK panel. The audit never substitutes a pan-tissue label and never changes organ release coverage.

The fifth primary artifact is built from official HCL Figshare item 7235471.v4, DOI 10.6084/m9.figshare.7235471.v4, released under CC BY 4.0 and linked to GSE134355 and PMID 32214235. The immutable snapshot contains the 599,926-cell by 27,341-gene Figure 1 H5AD, exact author cell-information workbook, and author cluster-marker workbook. The builder verifies all cell identities and their H5AD concatenation suffixes, all 102 cluster labels, 59 tissue categories, 106 batches, raw non-negative integer UMI values, and the exact cell order between the H5AD and workbook.

The current conservative HCL branch normalizes adult vermiform appendix and gallbladder only. It uses 4,486 appendix cells and 9,769 gallbladder cells, a closed author-label mapping, canonical CL and UBERON identities, and unique official HGNC identities. Fetal, state-only, anatomically misleading, or unsupported labels are quarantined rather than guessed; for example, the one gallbladder cell carrying the global cluster label Sinusoidal endothelial cell is not coerced into hepatocyte or hepatic endothelial CL terms. Marker derivation compares each mapped cell type with every other author-labeled cell in the same organ and requires at least 30 target and background cells, 25% sensitivity, 60% specificity, a 0.20 detection-fraction difference, log2 fold change of 1.5, CPM of 1, and a five-gene panel. The released artifact contains 224 positive grade-C panel components in nine panels: appendix T, B, plasma, and dendritic cells, plus gallbladder endothelial, goblet, smooth-muscle, macrophage, and fibroblast cells.

Each selected HCL organ contains one author donor and one source batch. The runtime therefore exposes one donor, one batch, positive-only status, and the single human_cell_landscape_gse134355 evidence family; it never describes these panels as donor-replicated, negative-marker evidence, or external validation. HCL-only annotation requires multiple matching genes, 10% reference-panel coverage, and the 0.90 supported-ancestor preference. The pan-tissue sibling-conflict audit is not used to force HCL abstention because tissue-specialized CL terms and lineage-subtype terms can occupy compatible classification axes; ordinary schema, ontology, active-record, confidence, and citation checks remain mandatory. The 11-case same-artifact regression covers all nine released panels and proves one-marker abstention in both organs.

The sixth primary artifact is derived from the official MCA Figshare item 5435866.v8, released under CC BY 4.0 and linked to PMID 29474909. Its Figure 2 branch uses the published batch-background-removed expression matrix and exact 98-cluster annotation workbook. It validates the order of all 61,637 matrix cells, processes 25,133 gene rows, admits 22,897 of 25,465 cells in selected current-schema tissues, and quarantines 1,502 cells whose published label is anatomically incompatible with its tissue. The later 333,778-cell H5AD is not assigned Figure 2 labels because its 104 recomputed Louvain IDs are a different label space.

Figure 2 marker admission requires at least 30 target and background cells, 25% target detection, 60% specificity, a 0.20 detection-fraction difference, log2 fold change of 1.5, CPM of 1, and at least five retained genes. For a tissue represented by multiple Figure 2 batches, a gene must also pass detection in at least two batches; single-batch tissues remain explicitly grade C rather than being described as independently replicated. This branch contains 2,102 positive panel components across 17 UBERON tissues, 35 CL types, 1,161 MGI genes, and 72 panels. Fifty-one panels have multi-batch support, 21 are single-batch, and two use an explicit all-other-tissues background.

The detailed adult-testis branch pins the author MCA_CellAssignments.csv, paired per-batch matrices, and age/sex workbook. It joins every one of 14,005 assigned testis cells to expression, keeps a closed mapping for all 19 author annotations, and coarsens marker-named subclusters only to defensible stable CL identities. A label-defining gene such as Tnp1 or Lyz2 is excluded from the corresponding derived panel. The adapter preserves the other Figure 2 thresholds except for a predeclared 0.15 detection-difference threshold selected by sensitivity analysis: 0.20 yielded only 62 assertions across three types, while 0.15 was the least permissive tested value that retained a two-batch spermatid panel. The packaged result adds 76 rows in four panels: spermatocyte, spermatid, Leydig cell, and macrophage. Spermatogonium, Sertoli cell, and erythroblast remain quarantined because they did not reach the five-marker gate.

Together the MCA branches contain 2,178 rows across 18 UBERON tissues and 76 tissue/cell-type panels. They share mouse_cell_atlas_figshare_5435866_v8, so the runtime never counts the branches, tissues, clusters, or batches as independent studies. Neither branch infers negative markers or creates automatic gene-specific paper claims.

An organ context supported only by MCA requires multiple query markers, at least 10% reference-panel coverage, and the 0.90 supported-ancestor preference. Responses expose both source branches, positive-only grade-C status, batch boundary, shared evidence family, source publication, and quantitative provenance. Known-panel and runtime regressions cover prostate, stomach, small intestine, ovary, placenta, uterus, and three testis identities, and prove that one-marker queries abstain. These are same-artifact integration checks, not held-out accuracy or confidence calibration.

The seventh source layer is derived from the exact CELLxGENE dataset version 513bf5e9-b2d8-43b0-99f5-ccf4b821f9d3, “Healthy gingiva of adult mouse,” linked to GSE267511, PMID 38876998, and DOI 10.1038/s41467-024-49037-y. The immutable local snapshot contains 12,118 cells and 30,198 features. The builder uses only the raw integer-count matrix and eight exact author labels already mapped to current CL terms; it never uses the scaled X matrix. Ensembl feature identifiers are resolved through the pinned MGI layer, unresolved or ambiguous genes and many-to-one MGI mappings are quarantined, and the source anatomy remains exact gingiva (UBERON:0001828). UBERON ancestry permits this panel to satisfy a mouth mucosa query without relabeling the source tissue.

Each of the eight panels has 30 retained positive markers after the same conservative within-dataset thresholds used for HCL: at least 30 target and background cells, 25% sensitivity, 60% specificity, a 0.20 detection difference, log2 fold change of 1.5, and target CPM of 1. The panels cover fibroblast, glial, muscle, vascular endothelial, gingival epithelial, mural, inflammatory, and salivary gland cells. All 240 rows share publication:PMID38876998 because the source is one biological replicate pooled from three mice. The runtime requires multiple matching genes and 10% panel coverage, exposes the pooling and positive-only grade-C boundary, and returns the publication as verified citation metadata rather than an automatically reviewed gene claim. It provides neither donor replication, negative evidence, an independent study, held-out validation, nor calibrated confidence.

The eighth source layer pins CELLxGENE dataset version 56496507-a8ca-4544-b364-a169ca71d60b, linked to ENA PRJEB57700, PMID 38570678, PMCID PMC11041794, and DOI 10.1038/s41586-024-07251-0. The immutable 134,373,984-byte H5AD contains 10,831 cells across caecum, small-intestine lamina propria, mesenteric lymph node, and Peyer's patch. The builder selects only the 3,509 exact caecum cells, reads raw.X integer UMI counts, and excludes every other tissue. It also quarantines 408 cells from three semantically incompatible author subclusters, including 367 GC.BC_DZ.1 cells that the portal mapped to a B-cell-zone reticular CL term even though the author label denotes germinal-center B cells.

Fourteen caecum immune or stromal cell types pass the 30-cell and five-marker gates, producing 420 positive grade-C components. Each panel uses every other selected caecum cell as background and requires 25% sensitivity, 60% specificity, a 0.20 detection difference, log2 fold change of 1.5, and target CPM of 1. All rows share publication:PMID38570678. The source exposes only a pooled donor label and comes from an H. hepaticus-colonized niche experiment; it is not a general untreated colon atlas. Only three epithelial cells are present, so no epithelial panel is released. The runtime can answer a broader colon request through validated UBERON ancestry, but it exposes the experimental condition, pooling, semantic quarantine, positive-only evidence, absent epithelial coverage, and lack of negative markers or held-out validation. A multi-marker macrophage regression resolves C1qa/C1qb/C1qc/Ms4a7/Csf1r; a single-marker query abstains.

The ninth source layer pins CELLxGENE dataset version 0e0d41d5-5403-4aac-a140-c1a4f6fb11c7 and retains only the healthy GSE164241/Williams buccal-mucosa branch. The source H5AD has 29,478 cells and 35,377 features; the builder selects 27,289 exact buccal-mucosa cells from seven donors, reads 95,449,700 raw integer UMIs, and excludes 2,189 labial-mucosa cells from the other upstream study. It resolves features to unique official HGNC identities and uses exact portal CL and UBERON terms. A 686-cell author label combining monocytes and macrophages is quarantined before target aggregation because it cannot support a macrophage-only panel.

Marker admission requires at least 30 target and background cells, 25% sensitivity, 60% specificity, a 0.20 detection difference, log2 fold change of 1.5, target CPM of 1, and at least five retained genes. A donor is eligible only with at least three target cells, and each marker must reach 20% detection in at least 60% of eligible donors. Fourteen cell types pass, producing 394 positive grade-C panel components. Every row shares publication:PMID34129837; seven donors strengthen within-study recurrence but remain one study family. The integrated atlas publication, PMID 42147490, is retained as dependent processing provenance rather than independent evidence. No negative marker or gene-specific paper claim is inferred. Upstream data terms apply, raw human cell-level data remain local, and runtime evidence emits LICENSE_REVIEW. A five-gene KRT5/KRT14/KRT15/S100A2/S100A14 regression returns keratinocyte for a mouth-mucosa query through validated UBERON ancestry; a one-marker query abstains.

The tenth source layer pins CELLxGENE dataset version 1fc5cf5e-9e89-445d-9898-aa5245d7324d, dataset ID 0f788857-3015-4fb9-912d-2545f2779c07, GSE164291, PMID 33818810, PMCID PMC8189321, and DOI 10.1096/fj.202002747R. The 538,198,472-byte H5AD contains 5,440 normal primary oviduct cells and 27,064 features. Its raw.X matrix has 14,242,606 nonzero integer values and 41,384,072 total UMIs. The source is one physiological-estrus sample made by pooling oviducts from five adult female mice.

Six cell types pass the 30-cell and five-marker gates, producing 180 positive grade-C components: fibroblast, endothelial, smooth-muscle, macrophage, secretory epithelial, and multiciliated epithelial panels. The mixed T/NK group has only 17 cells and is excluded. Every panel compares its target with all other author-labeled cells, contains 30 markers, and shares publication:PMID33818810. Pooled animals are not independent replicates; no negative marker, gene-specific paper claim, independent study, or held-out validation is inferred. Public CELLxGENE/GEO access and the PMC author manuscript do not establish an aggregate-data redistribution license, so the raw H5AD stays local and runtime evidence emits LICENSE_REVIEW. Multi-marker secretory and multiciliated regressions return the expected CL terms with citation metadata, dependent-evidence, and license warnings; a one-marker query abstains.

The eleventh source layer is the CC BY 4.0 primary study “A single cell atlas of the mouse seminal vesicle,” linked to GSE267191, PMID 40036847, PMCID PMC12060236, and DOI 10.1093/g3journal/jkaf045. Its immutable local snapshot retains the two source workbooks and GEO series metadata, while the package contains only normalized aggregate assertions and citation records. Supplementary Table S2 supplies 4,077 differential-expression rows across 21 author clusters. The closed curation map admits 12 conservative CL panels and quarantines four ambiguous, contaminating, doublet, or prostate-like clusters.

The quantitative build applies adjusted-P-value, log2-fold-change, sensitivity, specificity, and detection-difference floors, retains at most 30 genes per panel, and emits 313 positive grade-C components representing 243 MGI genes. Secretory, basal, and fibroblast identities additionally require recurrence across author subclusters. All rows share publication:PMID40036847: the 23 mice, multiple clusters, and same-study flow cytometry are within-study support, not independent evidence families or held-out validation. Twenty exact Table S2 gene/cluster rows were manually admitted as tissue-specific panel_component claims. A five-gene Nkg7/Cd3g/Gzma/Ccl5/Xcl1 regression returns the paper's presumptive NKT mapping with three reviewed claims and a dependent-evidence warning; one marker abstains. The result remains uncalibrated, positive-only, and without negative-marker or independent-study support.

The twelfth source layer is the CC BY 4.0 GSE179524 mouse-gallbladder study, linked to PRJNA744145, SRP327140, PMID 34650971, PMCID PMC8505819, and DOI 10.3389/fcell.2021.714271. Its immutable source snapshot pins GEO metadata, the main cell-type figure, and four supplementary marker figures by checksum. The study reports 35,654 cells from four pooled 15-mouse groups spanning basal chow, lithogenic diet, and TUDCA. Because GEO does not release the author's cell-to-cluster assignments or a quantitative differential-expression table, the project does not re-cluster the count matrices or present circularly inferred statistics as author evidence.

Manual review admits 32 positive grade-C author-marker components in five conservative panels: gallbladder epithelial cell, endothelial cell, smooth-muscle cell, macrophage, and gallbladder fibroblast. Each component has an exact Results or figure locator and a matching reviewed literature claim, but all rows share publication:PMID34650971. The proliferative state cluster and unresolved cluster 22 remain excluded. The machine-readable six-case regression reproduces all five multi-marker decisions and one-marker abstention with direct citation validation. It establishes deterministic policy behavior only: there is no negative evidence, per-animal recurrence, independent study, held-out validation, or calibrated probability.

The thirteenth source layer pins GSE254855, PMID 39426382, PRJNA1072091, and DOI 10.1016/j.devcel.2024.09.025. Its immutable local snapshot contains the author annotation table, the 44,008-gene by 15,949-cell intron-plus-exon UMI matrix, and GEO metadata. The builder applies the author QC flag, removes 1,125 stomach cells, retains 12,952 esophagus cells, reconciles matrix barcodes independently of column order, and recomputes exact library totals from the released matrix before calculating CPM.

Thirteen stable CL mappings pass the declared pooled and batch-aware gates, producing 260 positive grade-C components: basal and suprabasal epithelial, fibroblast, macrophage, dendritic, Langerhans, NK, T, ILC2, mast, endothelial, pericyte, and smooth-muscle panels. State-only, broad, low-cell, and stomach labels are quarantined. Thirty-nine predeclared genes were manually selected from the derived panels and admitted as reviewed panel_component claims with exact project-computation provenance. They are not author-supplied DE rows, independently discriminative markers, or 39 independent experiments; all claims and marker rows share publication:PMID39426382. The 14-case regression recovers all 13 multi-marker panels and preserves a COL17A1-only abstention. Pericyte and smooth-muscle markers pass only one eligible batch, no negative markers or animal-level recurrence are available, and the raw source remains local because its redistribution license is unresolved.

The fourteenth source layer pins the MIT-licensed Figshare item 13157726.v2, GSE150327, PRJNA631774, PMID 33305192, and DOI 10.1016/j.isci.2020.101838. The immutable local RDS contains 19,951 genes and 6,995 cells. The production branch selects the 2,324 adult S3 cells from one 10-month female mouse, admits 2,231 cells, and quarantines 62 semantically unresolved Ascl3-positive duct cells plus 31 NK, stromal, or erythroid cells below the 30-cell floor.

Eight canonical CL panels pass the quantitative gates, producing 226 positive grade-C components: basal duct, myoepithelial, serous acinar, seromucous acinar, intercalated duct, striated duct, endothelial, and macrophage. Every panel uses all other admitted adult cells as background and retains at most 30 markers after target/background count, CPM, detection, specificity, detection-difference, and fold-change checks. Twenty-four canonical panel members are separately represented as reviewed project-recomputed claims; they are not author-supplied DE rows or adaptations of the CC BY-NC-ND article. All evidence shares publication:PMID33305192. The nine-case regression recovers all eight panels and preserves a KRT14-only abstention. This source is a one-sample, female-only, same-source engineering validation without negative markers, independent replication, held-out accuracy, or calibrated confidence.

The fifteenth source layer is the CC BY 4.0 adult-mouse adrenal-gland study linked to GSE108097, GSE134355, PMID 32743779, PMCID PMC7396412, and DOI 10.1186/s13619-020-00042-8. The immutable local acquisition pins the article metadata and three official supplementary workbooks. The marker builder reads the two author differential-expression workbooks, requires a gene to pass the declared adjusted-P-value, fold-change, detection, specificity, and detection-difference gates on both BGI and Illumina outputs, and applies a closed author-cluster-to-CL map. Broad myeloid clusters, a proliferating state, and the anatomically unsafe Centrocyte label remain quarantined.

Ten conservative CL panels contribute 246 positive grade-C components: adrenal-cortex fasciculata, broad endothelial, macrophage, stromal, NK, dendritic, vascular endothelial, chromaffin, smooth-muscle, and neutrophil cells. Thirty predeclared genes are represented as reviewed, exact-row, author-supplied differential-expression claims. All marker and claim rows share publication:PMID32743779; the two sequencing platforms and three batches do not become independent studies. The 11-case regression recovers every multi-marker panel with exact-organ evidence and reviewed support, while a Cyp11b1-only query abstains. The artifact has no negative markers, sex metadata, held-out cohort, or calibrated probability, and the paper reports age only as adult.

The sixteenth source layer pins GSE239316, PRJNA998802, PMID 40899455, PMCID PMC12888919, DOI 10.1093/procel/pwaf074, and the CC BY 4.0 article supplements. Its immutable snapshot contains the official sample workbook, 91,078-row author Table S2 differential-expression workbook, supplementary methods PDF, and GEO series metadata. The study reports 169,859 high-quality cells or nuclei from 24 C57BL/6J mice: twelve 6-month and twelve 24-month animals, balanced by sex within age. Five organs use scRNA-seq and hypothalamus uses snRNA-seq.

A closed mapping admits 84 panels across adrenal gland, hypothalamus, islet of Langerhans, pineal body, pituitary gland, and thyroid gland. The build retains 2,433 positive grade-C components after adjusted-P-value, log2-fold-change, target detection, background detection, detection-difference, five-gene panel-floor, MGI identity, and within-tissue shared-gene ownership gates. Double-negative and proliferating lymphocytes, broad NonThy-Epe, and ambiguous islet stellate states are quarantined. Three deterministic released markers per panel become 252 reviewed exact-tissue Table S2 claims. All rows remain one publication family; Table S2 is pooled within organ and is not stratified by age, sex, animal, or library. The paper compares external public datasets for adrenal, hypothalamus, islet, and pituitary, but not thyroid or pineal, and those comparisons are not counted as new claims or independent evidence families.

The 90-case regression covers every admitted panel plus one single-marker abstention per organ. Annotation coverage is measured within each source artifact panel, so a compact organ panel is not diluted by hundreds of same-CL markers from another database. The gate still requires at least two genes from one source panel, preserves combined cross-source coverage as an audit metric, and permits ontology ancestor preference only when the ancestor passes the same coverage gate. This is deterministic same-artifact validation, not held-out accuracy, confidence calibration, or external generalization.

The seventeenth source layer pins GSE145443, PRJNA607227, SRP249803, PMID 32729827, PMCID PMC7426093, DOI 10.7554/eLife.55474, and the CC BY 4.0 article. The immutable audit records the official GEO metadata and adjusted count matrix, the author cell metadata, the article marker supplement, and exact byte sizes and SHA-256 checksums. Production excludes all 3,506 vas-deferens cells, admits 5,290 epididymis cells, and quarantines 84 mixed, state-only, or unresolved cells.

Six conservative CL panels contribute 180 positive grade-C components: epididymal secretory epithelial, basal epithelial, fibroblast, endothelial, macrophage, and smooth-muscle cells. The builder found that the released adjusted matrix has 16,892 unique gene rows rather than the 16,878 reported in the paper and that its cell library sums are 976,746 UMIs below the pre-SoupX UMI_sum metadata total; it therefore derives quantitative statistics only from the adjusted matrix after an explicit two-pass audit. Eighteen canonical components are reviewed as project-recomputed claims in the single publication:PMID32729827 family. The seven-case regression recovers all six multi-marker panels and preserves a Krt14-only abstention. The result is positive-only, same-study engineering validation with no negative markers, held-out cohort, independent family, or calibrated confidence. GEO supplies public access but no explicit content redistribution license, so raw files stay local and runtime evidence remains LICENSE_REVIEW; the article CC BY 4.0 license is not applied to the data.

The eighteenth source layer pins PMID 37007744, PMCID PMC10065133, DOI 10.1093/pcmedi/pbad006, and the CC BY 4.0 Europe PMC full-text and supplementary-methods responses. The study reports five adult virgin female mice, 7,823 captured vaginal-epithelium cells, 6,187 post-QC cells, and six author clusters. No public expression-matrix accession or author cell-label payload was identified, so the build performs no quantitative reanalysis.

Ten author-defined positive components form three conservative multi-gene panels: basal cells (Ngfr, Esr1, Fzd10, Tcf7), later Ifitm3-high/Notch-high suprabasal cells, and vagina squamous cells (Sprr1b, Krt1). Proliferating Mki67/Birc5 is quarantined as a state; Krt8-only columnar and Fabp5-only suprabasal identities remain below the multi-marker floor. All marker and reviewed-claim rows share publication:PMID37007744. The four-case regression recovers all three released panels and preserves an NGFR-only abstention. This layer adds exact marker presence for mouse vagina but does not satisfy the six-cell-type, negative-evidence, two-family, or validation release gates.

The nineteenth source layer pins CELLxGENE dataset edc8d3fe-153c-4e3d-8be0-2108d30f8d70, dataset version f9efb73e-f116-46b5-a775-d4233e758024, PMID 34937051, PMCID PMC8828466, DOI 10.1038/s41586-021-04345-x, and EGA accession EGAD00001007718. The 3,573,361,310-byte H5AD contains 236,977 primary airway cells and 32,383 features. Production selects exactly 102,060 healthy nasal-cavity cells from 37 donors and 38 samples; COVID-19, post-COVID, bronchial, and tracheal cells are not admitted into the nasal panels.

The adapter resolves HGNC, CL, and UBERON identities, requires marker recurrence in at least 60% of eligible donors, and emits 863 positive grade-C components across 29 multi-marker panels. It quarantines 12,130 cycling, activated, inflammatory, unknown, generic-state, or portal-misassigned cells. Safe trachea-specific portal terms are mapped to generic goblet, basal, or tuft identities only after author-label checks; unsupported secretory, melanocyte, neuroendocrine, and NKT panels abstain at the build gates. Eighty-seven canonical components are manually reviewed as project-recomputed claims, not author-supplied differential-expression rows or article adaptations. The five-case regression recovers ionocyte, tuft, basal, and goblet identities and preserves a CFTR-only abstention. All evidence remains one publication family, positive only, uncalibrated, without held-out or independent-study validation, and under upstream participant-data terms.

The twentieth source layer pins GSE245074, BioProject PRJNA1026904, exact naive-WT samples GSM7835749 through GSM7835752, PMID 38306414, PMCID PMC11127180, and DOI 10.1126/sciimmunol.abq4341. The immutable local acquisition preserves the official GEO SOFT file and doubly compressed Seurat RDS with exact size, MD5, and SHA-256 checks. The source object contains 50,448 cells, 31,053 genes, 126,460,493 nonzero values, and 415,348,869 raw UMIs. Production selects 19,275 untreated wild-type nasal cells from four mice, two female and two male, and reconstructs each animal from the author metadata.

The adapter uses the raw RNA@counts matrix, closed CL mappings, exact MGI and UBERON identities, pooled differential-expression gates, and a recurrence requirement of at least three mice. It admits 17,081 cells into 34 panels, quarantines 2,194 developmental-state, inflammatory-state, undersized, or unsafe-mapping cells, rejects 43 gene-panel pairs that pass pooled gates but fail animal recurrence, and emits 1,020 positive grade-C components. The literature layer adds 102 manually selected canonical project-recomputed claims, three per panel, in the single publication:PMID38306414 family. A five-case regression recovers ionocyte, olfactory receptor, tuft, and nasal serous identities and preserves a Foxi1-only abstention. FACS equalized epithelial and immune compartments and olfactory and respiratory material, so the source cannot support unbiased cell-frequency estimates. GEO states no explicit dataset license and the article is not in the PMC Open Access subset; raw files remain local, derived aggregates require license review, and the runtime reports no negative, held-out, independent-study, or calibrated evidence.

The adult-mouse corpus-cavernosum layer pins PMID 38856719, PMCID PMC11164535, DOI 10.7554/eLife.88942, and KAP230548. Manual source review admits 27 positive grade-C components across chondrocyte, fibroblast, myofibroblast, lymphatic-endothelial, vascular-endothelial, smooth-muscle, pericyte, Schwann-cell, and macrophage panels. The ten mice were pooled into one normal and one diabetic library; the source therefore supplies neither per-animal recurrence nor a healthy-only marker analysis. LBH is retained only beside established pericyte markers, and valve-related, differentiation-stage, single-marker, and state-defined signals are quarantined. The ten-case regression recovers all nine complete panels and keeps LBH alone unknown. The article is CC BY 4.0, but KAP230548 redistribution terms were not stated, so the raw reads are not packaged or reused.

The separate pan_tissue_marker_assertions.tsv artifact is derived only from the packaged Cell Marker Accordion records. A positive (species, CL term, gene) assertion is retained when it occurs in at least three distinct source tissues and two distinct upstream evidence families. Every derived row is assigned the root anatomy context UBERON:0000061, downgraded to grade C, and marked as non-independent. It is used only when an organ-specific marker context is absent. Fallback annotation additionally requires at least 10% of the candidate reference panel, prefers a supported CL ancestor within 90% of a more specific candidate's score, and preserves near-tie and confidence abstention. search_cell_markers, rank_candidate_cell_types, and annotate_clusters expose allow_pan_tissue_fallback; set it to false for strict organ-only behavior. See Third-party data notices.

Coverage can be inspected before annotation:

get_knowledge_coverage(species="human", organ="UBERON:0002107")

The response reports assertion, marker-panel cell-type, Census reference-cell-type, gene, negative-marker, evidence-family, healthy-atlas collection, unreviewed literature-candidate, combined literature-discovery, approved-literature, and pan-tissue-fallback counts together with each failed release gate. Its summary includes pass and failure totals for all six gates, coverage-level totals, the incomplete-context count, and the exact contexts that pass every gate. In the current 102-context release, the cell-type, major-reference, evidence-family, negative-evidence, reviewed-literature, and healthy-atlas gates pass in 76, 38, 40, 20, 39, and 71 contexts respectively. Human blood, bone marrow, lymph node, spleen, brain, eye, kidney, liver, lung, pancreas, prostate, and thymus; the parent human bone-element context reached through exact descendant evidence; mouse brain, lung, and kidney pass all six. This produces 16 complete contexts and 86 incomplete contexts. Combined discovery is the union of candidate records and exact-tissue approved records, preventing completed review from being misreported as a discovery gap. Fallback availability is reported separately and does not change assertion_count, coverage_level, complete, or any release gate. The scope and thresholds are versioned in organ_scope.tsv; passing a record-count threshold alone never implies biological validation.

Reference atlas metadata

search_reference_datasets queries a local snapshot of the official CELLxGENE curation API. The packaged 2026-08-10 snapshot contains 1,611 human and 491 mouse dataset records; 1,754 records include the portal's normal disease term. Results preserve dataset and collection version IDs, DOI, publication year, UBERON tissues, diseases, assays, suspension types, raw-data links, and collection-level dependency. Runtime queries do not contact CELLxGENE.

This is a discovery and coverage index, not a redistribution grant or biological validation result. The package contains no expression matrix or cell-level metadata. Each dataset payload remains subject to its own license, consent restrictions, and source terms; those must be reviewed before download or reuse.

search_reference_cell_types queries a second offline artifact built from the pinned CELLxGENE Census LTS release 2025-11-08. The builder selects portal-normal, primary observations whose tissue_type is tissue, maps their UBERON anatomy to the 52-organ scope, validates cell labels against the packaged Cell Ontology, and emits only organ/cell-type aggregates. The packaged index contains 3,717 human and 763 mouse records. A record is marked as a major reference only when it contains at least 100 cells and spans at least two collection IDs. Scope-specific descendant exclusions prevent anatomically adjacent but out-of-scope tissues from entering an organ aggregate; for example, epididymal fat pad is not epididymis evidence.

The Census index contains no expression values, donor attributes, or cell-level rows. Cell-type presence helps establish which labels occur in a reference atlas; it is not marker-gene evidence, and collection count is only a dependency proxy. Consequently, annotate_clusters does not use these counts as evidence or increase confidence from them.

To use another reviewed normalized store:

export CELL_ANNOTATION_MARKER_ASSERTIONS=/absolute/path/to/marker_assertions.tsv
export CELL_ANNOTATION_PAN_TISSUE_MARKERS=/absolute/path/to/pan_tissue_marker_assertions.tsv
export CELL_ANNOTATION_GENE_NOMENCLATURE=/absolute/path/to/gene_nomenclature.tsv
export CELL_ANNOTATION_SCTYPE_PANELS=/absolute/path/to/sctype_panels.tsv
export CELL_ANNOTATION_CELLMATCH_MARKERS=/absolute/path/to/cellmatch_markers.tsv
export CELL_ANNOTATION_GARNETT_PANELS=/absolute/path/to/garnett_panels.tsv
export CELL_ANNOTATION_HRA_ASCTB_MARKERS=/absolute/path/to/hra_asctb_marker_assertions.tsv
export CELL_ANNOTATION_TABULA_MURIS_SENIS_MARKERS=/absolute/path/to/tabula_muris_senis_marker_assertions.tsv
export CELL_ANNOTATION_HPA_SINGLE_CELL_MARKERS=/absolute/path/to/hpa_single_cell_marker_assertions.tsv
export CELL_ANNOTATION_HUMAN_BRAIN_SILETTI_MARKERS=/absolute/path/to/human_brain_siletti_marker_assertions.tsv
export CELL_ANNOTATION_HUMAN_CELL_LANDSCAPE_MARKERS=/absolute/path/to/human_cell_landscape_marker_assertions.tsv
export CELL_ANNOTATION_HUMAN_ORAL_MUCOSA_MARKERS=/absolute/path/to/human_oral_mucosa_marker_assertions.tsv
export CELL_ANNOTATION_HUMAN_NASAL_CAVITY_MARKERS=/absolute/path/to/human_nasal_cavity_marker_assertions.tsv
export CELL_ANNOTATION_HUMAN_NASOPHARYNX_MARKERS=/absolute/path/to/human_nasopharynx_marker_assertions.tsv
export CELL_ANNOTATION_HUMAN_PARATHYROID_MARKERS=/absolute/path/to/human_parathyroid_marker_assertions.tsv
export CELL_ANNOTATION_HUMAN_SEMINAL_VESICLE_MARKERS=/absolute/path/to/human_seminal_vesicle_marker_assertions.tsv
export CELL_ANNOTATION_HUMAN_EYE_VANZYL_MARKERS=/absolute/path/to/human_eye_vanzyl_marker_assertions.tsv
export CELL_ANNOTATION_HUMAN_PROSTATE_HENRY_MARKERS=/absolute/path/to/human_prostate_henry_marker_assertions.tsv
export CELL_ANNOTATION_HUMAN_THYMUS_HEIMLI_MARKERS=/absolute/path/to/human_thymus_heimli_marker_assertions.tsv
export CELL_ANNOTATION_MOUSE_PARATHYROID_MARKERS=/absolute/path/to/mouse_parathyroid_marker_assertions.tsv
export CELL_ANNOTATION_MOUSE_OVIDUCT_MARKERS=/absolute/path/to/mouse_oviduct_marker_assertions.tsv
export CELL_ANNOTATION_MOUSE_SEMINAL_VESICLE_MARKERS=/absolute/path/to/mouse_seminal_vesicle_marker_assertions.tsv
export CELL_ANNOTATION_MOUSE_GALLBLADDER_MARKERS=/absolute/path/to/mouse_gallbladder_marker_assertions.tsv
export CELL_ANNOTATION_MOUSE_ESOPHAGUS_MARKERS=/absolute/path/to/mouse_esophagus_marker_assertions.tsv
export CELL_ANNOTATION_MOUSE_SALIVARY_GLAND_MARKERS=/absolute/path/to/mouse_salivary_gland_marker_assertions.tsv
export CELL_ANNOTATION_MOUSE_ADRENAL_MARKERS=/absolute/path/to/mouse_adrenal_marker_assertions.tsv
export CELL_ANNOTATION_MOUSE_ENDOCRINE_ATLAS_MARKERS=/absolute/path/to/mouse_endocrine_atlas_marker_assertions.tsv
export CELL_ANNOTATION_MOUSE_EPIDIDYMIS_MARKERS=/absolute/path/to/mouse_epididymis_marker_assertions.tsv
export CELL_ANNOTATION_MOUSE_VAGINA_MARKERS=/absolute/path/to/mouse_vagina_marker_assertions.tsv
export CELL_ANNOTATION_MOUSE_ENDOMETRIUM_MARKERS=/absolute/path/to/mouse_endometrium_marker_assertions.tsv
export CELL_ANNOTATION_MOUSE_PENIS_MARKERS=/absolute/path/to/mouse_penis_marker_assertions.tsv
export CELL_ANNOTATION_MOUSE_RECTUM_MARKERS=/absolute/path/to/mouse_rectum_marker_assertions.tsv
export CELL_ANNOTATION_MOUSE_CELL_ATLAS_MARKERS=/absolute/path/to/mouse_cell_atlas_marker_assertions.tsv
export CELL_ANNOTATION_MOUSE_CELL_ATLAS_TESTIS_MARKERS=/absolute/path/to/mouse_cell_atlas_testis_marker_assertions.tsv
export CELL_ANNOTATION_MOUSE_GINGIVA_MARKERS=/absolute/path/to/mouse_gingiva_marker_assertions.tsv
export CELL_ANNOTATION_MOUSE_CAECUM_MARKERS=/absolute/path/to/mouse_caecum_marker_assertions.tsv
export CELL_ANNOTATION_MOUSE_BRONCHUS_CELLCARDS_MARKERS=/absolute/path/to/mouse_bronchus_cellcards_marker_assertions.tsv
export CELL_ANNOTATION_RUNTIME_BACKEND=sqlite
export CELL_ANNOTATION_RUNTIME_DATABASE=/absolute/path/to/runtime_evidence.sqlite3
export CELL_ANNOTATION_SOURCE_PANEL_POLICIES=/absolute/path/to/source_panel_policies.tsv
cell-annotation-mcp

The default sqlite backend opens the runtime database with read-only, immutable, query-only connections and verifies every configured marker TSV against the database manifest. After changing any marker path, rebuild the index with cell-annotation-admin build-runtime-database --overwrite; use CELL_ANNOTATION_RUNTIME_BACKEND=tsv only when the compatibility loader is intentionally required.

The required columns are defined in marker_store.py. Every assertion must include a stable assertion ID, gene, canonical CL term, species, polarity, evidence grade, registered source ID, source release, and curation status. The loader fails on malformed rows, duplicate assertion IDs, and unsupported values.

Gene nomenclature and query normalization

resolve_gene queries the packaged gene_nomenclature.tsv artifact. The human component is the complete HGNC download last modified 2026-08-07; the mouse component is the MGI 6.24 marker report and gene-model cross-reference report dated 2026-08-03, restricted to Gene and Pseudogene. The artifact contains 135,807 records and its manifest records input URLs, HTTP modification dates, raw and normalized SHA-256 checksums, record counts, source releases, licenses, and exclusions.

Marker ranking returns a gene_normalization audit for every input, including the resolution status, stable ID, approved symbol, whether the input was used for scoring, and any ambiguity candidates. Exact HGNC/MGI IDs and unique official Ensembl or Entrez cross-references are accepted in addition to symbols and aliases; for example, ENSG00000000003 resolves to TSPAN6. The packaged Cell Marker Accordion audit found 179,107 assertions with an exact approved symbol, 564 resolved through an alias or previous symbol, two mouse case-normalized assertions, 306 unresolved assertions, and one ambiguous human assertion. These categories describe nomenclature resolution only; they do not change evidence grade or establish marker specificity.

Literature evidence

search_literature_candidates covers the discovery side of the workflow. The packaged 2026-08-10 queue was generated from 102 UBERON organ/species queries through the official NCBI Entrez wrapper with a per-query retmax of 2,000. No query was truncated. After PMID deduplication and removal of 28 records already in the citation store, it contains 7,628 citation records across 99 non-empty candidate contexts. Mouse nasopharynx returned no result under that strict query design, but three subsequently reviewed exact-subregion publications now supply approved evidence. Reviewed mouse-brain, mouse kidney, mouse seminal-vesicle, gallbladder, human seminal-vesicle, human-pancreas, human-prostate, human-thymus, human-kidney, human-liver, human-bone-marrow, human-eye, and mouse-nasopharynx papers are not reintroduced as unreviewed candidates. The coverage API reports discovery as the union of candidate records and exact-tissue or exact-descendant reviewed records; all 102 required contexts now have discovery coverage.

Each candidate retains exact (UBERON tissue, species taxon, query ID, retrieval rank) tuples. These tuples are not separated or recombined, so a paper retrieved by human-brain and mouse-colon queries cannot be mislabeled as human colon. Query association remains a discovery signal rather than proof that the paper studied that exact context. Candidate search defaults to the original PubMed relevance rank and can instead sort by publication year.

Candidate records never enter marker ranking, annotation confidence, or direct literature support. Promotion requires manual source-level review of a specific figure, table, Results passage, method, or supplement plus CL, species, tissue, dataset accession, evidence-family, and approval-scope assignment.

The packaged citation layer contains 80 PubMed records and 1,091 reviewed claims. Of these, 1,055 are exact-tissue approved_identity panel components across 41 direct evidence families and 36 are context or specificity cautions across 16 families. The source-specific reviewed sets include human brain, eye, lung, pancreas, prostate, thymus, kidney cortex, liver, bone marrow, lymph node, spleen, nasal cavity, nasopharynx, parathyroid, cultured human seminal-vesicle epithelium, mouse brain, lung, kidney, seminal vesicle, bronchus, gallbladder, esophagus, submandibular gland, adrenal, the six-organ endocrine atlas, epididymis, vagina, endometrium, corpus cavernosum, rectum, parathyroid, and nasopharynx. Other atlas and source-paper records remain metadata-only unless a claim row explicitly says otherwise.

The adult-human-brain addition replaces the metadata-only Siletti citation with 40 approved identity components and six reviewed negative-context components. Every claim is exact to human brain (UBERON:0000955) and points to the pinned class-rule line, cluster workbook columns, repository commits, and Allen taxonomy release. All 46 claims and all HPA brain rows derived from the same study share upstream_pubmed:37824663; they are one family, not independent corroboration. A seven-gene oligodendrocyte panel returns CL:0000128, cites reviewed PLP1, MBP, and MOG claims, and passes the six release gates. RBFOX3 alone abstains. Adding PDGFRA produces an explicit negative conflict and lowers the uncalibrated ranking score. This deterministic self-check is not a raw-expression reanalysis, held-out benchmark, external replication, or calibrated probability.

The developing-mouse-brain addition pins La Manno PMID 34321664, DOI 10.1038/s41586-021-03775-x, PRJNA637987, author-rule commit c398c285861ac714ec1e8827d1f5ac3661d4b8da, and the final 798-cluster, 292,495-cell Supplementary Table 2. It audits all 215 YAML rules but admits only 12 stable rules with exact current CL mappings and observed final-cluster use, producing 31 positive and three author-explicit negative components. Six stable compound, regional, unresolved, or unmapped rules remain quarantined. A five-gene OPC query returns CL:0002453 with reviewed PDGFRA and CSPG4 evidence; PDGFRA alone abstains, and adding LUM triggers the author-explicit OPC negative rule and abstention. Every row shares upstream_pubmed:34321664. This is a developmental E7–E18 same-artifact provenance and policy regression, not an adult atlas, per-embryo recurrence analysis, independent cohort, held-out benchmark, or calibrated probability.

The adult-human-bone-marrow addition pins Bandyopadhyay et al. PMID 38714197, PMCID PMC11162340, DOI 10.1016/j.cell.2024.04.013, GSE253355, PRJNA1065394, BioStudies S-EPMC11162340, twelve GEO samples, and the 111,315-row Supplementary Table S2 workbook by size and SHA-256. Twelve project-reviewed claims preserve the exact worksheet row, P value, log2 fold change, target and background detection, and adjusted P value. Every claim must match an existing active marker in the same species, tissue, and CL identity. A 25-gene osteoblast panel returns CL:0000062, links the exact IBSP row and PMID/DOI, and validates all 25 scoring evidence objects; IBSP alone remains unknown. All claims remain one publication family and add no scoring markers, negative markers, independent replication, held-out evaluation, or calibrated probability.

The adult-human-spleen addition pins Madissoon et al. PMID 31892341, PMCID PMC6937944, DOI 10.1186/s13059-019-1906-x, and PRJEB31843. Five project-reviewed claims preserve the exact Supplementary Figure S9c marker label, broad CL identity, embedded-panel checksum, human-spleen scope, and one shared study family. A ten-gene B-cell panel returns the conservative parent CL:0000236, attaches the MS4A1 claim and citation, and validates all scoring evidence; MS4A1 alone remains unknown. Figure dot values are not digitized, and the literature layer supplies no scoring rows, negatives, independent replication, held-out evaluation, or calibrated probability.

The adult-human-lymph-node addition pins Abe et al. PMID 35332263, PMCID PMC9033586, DOI 10.1038/s41556-022-00866-3, EGAD00001008311, and the exact 23-sheet supplementary workbook. Eleven project-reviewed claims preserve workbook, sheet, row, author population, effect size, detection fractions, adjusted P value, conservative CL mapping, and human-lymph-node scope. A ten-gene lymphatic-endothelial panel returns CL:0002138, attaches the exact PROX1 claim and citation, and validates all scoring evidence; PROX1 alone remains unknown. A separate PDPN heterogeneity statement is context only, not a negative scoring marker. The nine metastasis-free nodes were obtained from patients with neoplasms, all claims share one family, and controlled EGA data were not downloaded.

The healthy-adult-mouse-kidney addition independently reviews Park et al. PMID 29622724/GSE107585 and Ransick et al. PMID 31689386/GSE129798. Thirty-five exact supplementary-table rows become both positive grade-C scoring components and one-to-one literature claims across 28 exact CL identities. Each study is one evidence family; its batches, animals, sexes, zones, tables, and cell populations cannot create extra votes. A six-gene podocyte panel (Nphs1, Nphs2, Podxl, Synpo, Wt1, Ptpro) returns CL:0000653, links reviewed Park and Ransick evidence, and reports two independent literature families; Nphs1 alone remains unknown.

He et al. PMID 33837218/GSE160048 add a third kidney family. Four exact Supplementary Data 8 genes (Pdgfra, Gata3, Itga8, and MGI-normalized Septin4) support the conservative mesangial parent CL:0000650; Pdgfra alone remains unknown. The explicit CD45-positive-cell exclusion admits Ptprc as a context-specific contradiction, closing the organ-level negative-evidence gate without creating a universal kidney negative. The failed CD31 gate is preserved as a Pecam1 caution and never scores. The source-specific 50% panel policy is applied only to He-supported candidates, so it does not reject a podocyte panel corroborated across Park and Ransick. Mouse kidney now passes all six formal gates, while the regressions remain deterministic same-artifact provenance and policy checks rather than held-out accuracy, independent validation of the He mesangial call, or calibrated probability.

The Human Lung Cell Atlas addition pins PMID 37291214, PMCID PMC10287567, DOI 10.1038/s41591-023-02327-2, CELLxGENE collection 6f6d381a-7701-4781-935c-db10d30de293, the 118,146,619-byte PMC OA package, and its SHA-256. It admits only two exact healthy-lung panel components from Supplementary Table 6: MFSD2A for pulmonary alveolar type 2 cell and SOSTDC1 for alveolar capillary type 2 endothelial cell. Both must match an existing active scoring marker row. The article's 14 core datasets and 61 final types are one integrated dependency family; the HLCA adapter adds literature provenance, not marker scores or fourteen independent votes.

The Muraro pancreas addition pins PMID 27693023, PMCID PMC5092539.1, DOI 10.1016/j.cels.2016.09.002, GSE85241, the PMC article version, and the exact 525,036-byte Supplementary Table S3 workbook checksum. It admits nine canonical panel_component claims only after matching the Results marker passage, the exact workbook row, HGNC/CL/UBERON context, and an existing active positive marker. A seven-gene beta-cell panel returns CL:0000169 with the INS claim, PMID/DOI/GSE accession, and beta!A2:H2 locator; INS alone abstains. This is deterministic evidence linkage and output self-validation, not independent-study replication, held-out accuracy, or calibrated probability.

The Muto kidney addition pins PMID 33850129, PMCID PMC8044133, DOI 10.1038/s41467-021-22368-w, GSE151302 and reused GSE131882 samples, the PMC OA update, and the exact 677,255-byte Supplementary Data 1 workbook checksum. It admits 13 panel_component claims only after the Results context, exact worksheet row and quantitative values, HGNC/CL/UBERON mapping, and an existing positive scoring marker all agree. A seven-gene podocyte panel returns CL:0000653 with the PTPRO claim, PMID/DOI/GSE accession, and PODO!A3:F3 locator; PTPRO alone abstains. Five samples, including three reused samples, remain one evidence family. PT_VCAM1 and unsafe state or over-specific mappings remain quarantined. The source adds provenance and self-validation, not scoring rows, independent replication, held-out accuracy, or calibrated probability.

The human seminal-vesicle addition contributes 15 approved panel-component claims from PMID 37984408 and one SIM1 context-only claim. SIM1 is not a production marker because it is absent from the admitted cluster-DE rows and is detected in fewer than 20% of cells in each public seminal-vesicle library. The three additional-patient qPCR results belong to the same publication family and therefore do not increase the independent-family count.

The mouse-bronchus addition contributes 13 approved expert-panel components from the LungMAP CellCards publication, PMID 34936882. Every claim has an exact article text locator and shares derived_database:lungmap-cellcards-2022. The normal mouse-airway goblet-cell discussion is retained as context only: Muc5ac is not used as a negative marker or an annotation vote, because the source describes goblet cells as rare or absent at baseline and induced by injury or allergen exposure. The source is a static expert synthesis without a released raw matrix, quantitative differential-expression table, negative-marker set, held-out cohort, or calibrated confidence model.

The mouse-nasopharynx addition contributes 18 approved panel components and nine reviewed explicit-negative flow-gate claims from three independent publication families. GSE227311 provides four lymphatic-endothelial components in nasopharyngeal mucosa, backed by a checksum-pinned 842-cell normalized matrix from 30 adult mice. Two NALT studies provide broad B, CD4 T, CD8 T, gamma-delta T, myeloid, dendritic, and memory-T panels. The layer is deliberately subregion-bounded: it is not a whole-organ epithelial, stromal, vascular, neuronal, and immune atlas. TCRbeta and MHC class II remain assay-level features rather than guessed gene mappings; five source LEC subclusters and an M-cell branch remain quarantined where Cell Ontology lacks a safe non-gut term.

The mouse-parathyroid addition contributes 47 exact author-DE claims from PMID 37379351 in one experimental single-cell family and two PTH/GCM2 spatial claims from PMID 41025096 in a second publication family. Only the chief-cell panel can therefore report independently replicated literature support, and the second family is spatial rather than cell-resolved single-cell evidence. Every approval is limited to panel_component scope; no individual gene is declared independently discriminative, and duplicate tissue rows, panels, clusters, donors, participants, pooled animals, conditions, assays, libraries, preprint/final versions, or recomputed marker rows from one study still count as one evidence family.

The first direct evidence batch is deliberately bounded:

  • human blood B, T, and NK claims from SRP073767;

  • human blood B and NK protein-panel claims from GSE164378;

  • mouse spleen and liver NK claims from GSE189807;

  • mouse seminal-vesicle secretory, basal, stromal, vascular, immune, pericyte, lymphatic-endothelial, and Schwann-cell claims from GSE267191 Supplementary Table S2;

  • mouse bronchus basal, club, multiciliated, pulmonary-neuroendocrine, smooth-muscle, and chondrocyte expert-panel claims from LungMAP CellCards;

  • mouse nasopharyngeal-mucosa lymphatic-endothelial and naive-NALT immune panels from GSE227311 and two flow/immunophenotyping studies, including explicit exclusion markers that are used only for contradiction checks;

  • mouse gallbladder epithelial, endothelial, smooth-muscle, macrophage, and fibroblast claims from GSE179524 Results and author figures;

  • adult mouse esophagus epithelial, stromal, vascular, and immune panel claims selected from the reproducible GSE254855 computation, explicitly distinguished from author-supplied DE rows;

  • adult mouse submandibular-gland ductal, acinar, myoepithelial, endothelial, and macrophage claims selected from the reproducible GSE150327 computation, explicitly distinguished from author-supplied DE rows and article adaptation;

  • adult mouse vaginal basal, later suprabasal, and vagina-squamous panel claims manually located to PMID37007744 Results and Figure 1, with state-only and source-only single-marker signals excluded;

  • adult mouse endometrial luminal and glandular epithelial panels from PMID32998907 Table S1 and Results/Figure evidence, with Foxa2, Foxj1, and Tacstd2 non-specificity observations retained as context or explicit negative evidence rather than independent identity support;

  • adult mouse corpus-cavernosum chondrocyte, stromal, vascular, smooth-muscle, pericyte, Schwann-cell, and macrophage panels from PMID38856719 Methods, Results, and Figures 1-2, with LBH restricted to exact-tissue panel use and the pooled normal/diabetic design retained as a limitation;

  • healthy human nasal ionocyte, tuft, basal, goblet, immune, and other epithelial panel components selected from the reproducible PMID34937051 donor-consistent computation, with disease/state exclusions and without claiming author-supplied differential-expression rows;

  • healthy human nasopharynx ciliated, squamous, secretory, macrophage, CD8 T, basal, goblet, and ionocyte panel components selected from the reproducible PMID34352228 participant-recurrent computation, with exact healthy WHO-0/PCR-negative scope and without treating its preprint as independent evidence;

  • adult mouse adrenal cortical, medullary, vascular, stromal, smooth-muscle, and immune panel claims from exact rows in the paired author differential-expression workbooks;

  • mouse adrenal, hypothalamus, pancreatic-islet, pineal, pituitary, and thyroid claims from three exact GSE239316 Table S2 rows per released tissue/cell panel;

  • Tabula Muris and Tabula Muris Senis metadata without paper-level gene approval; the Senis marker layer remains explicitly computational primary-atlas evidence;

  • no approved mouse blood B- or T-cell claim yet; the approved seminal-vesicle T-cell and presumptive NKT claims are organ-specific and do not transfer to blood.

search_literature searches the local index. Every annotate_clusters result includes an automatic literature_validation block, and cross_check_annotation_literature provides the same check for an external annotation. Both paths require the complete annotation validator to pass before verifying immediate-source PMID/DOI metadata and matching reviewed gene, cell type, species, and tissue scope. Every returned reviewed claim includes its PMID/DOI source citation and exact evidence locator. The result distinguishes:

  • citation_metadata_only: the publication metadata resolves, but no reviewed biological claim matches;

  • context_only: a reviewed contextual or non-specificity claim matches;

  • reviewed_identity_supported: at least one manually approved identity-marker claim matches;

  • none: no packaged literature evidence matches.

Direct support requires the reviewed UBERON tissue to match exactly either the requested annotation tissue or the source assertion's exact tissue. This lets a parent-organ request retain correctly scoped descendant evidence without relabeling it as parent-organ evidence. A gene/cell/species match from any other tissue is returned under approved_identity_scope_mismatches and cannot increase the direct-support count. The summary separately reports claim counts, unique evidence-family counts, and independently_replicated_identity_support. The literature layer never changes the current confidence value and performs no network request at runtime. Automated PubMed discovery enters a review queue; it does not write production markers. See Literature evidence for the workflow, source set, review rules, and maintenance contract.

Custom reviewed artifacts can be selected with:

export CELL_ANNOTATION_LITERATURE_RECORDS=/absolute/path/to/literature_records.tsv
export CELL_ANNOTATION_LITERATURE_CLAIMS=/absolute/path/to/literature_claims.tsv
export CELL_ANNOTATION_LITERATURE_CANDIDATES=/absolute/path/to/literature_candidates.tsv
cell-annotation-mcp

Validated annotation output

annotate_clusters and validate_annotation_result use schema version 1.0. An accepted annotation contains:

  • a canonical, non-obsolete CL ID whose label matches the packaged ontology;

  • a decision of annotated, parent_only, or unknown;

  • a confidence value explicitly described as uncalibrated when generated by the current marker ranker;

  • same-source panel coverage requiring the configured number of distinct markers and coverage fraction within at least one immediate-source panel;

  • source ID, release, record ID, evidence family, direction, statement, and genes for every evidence item;

  • active-store verification that marker record IDs match their source, release, family, species, gene, polarity/direction, and predicted CL term;

  • canonical UBERON validation of tissue IDs and labels plus rejection of marker evidence from unrelated tissues;

  • structured validation errors, warnings, and a provenance summary.

Ranked candidates expose combined positive_marker_coverage for audit, per-source source_panel_coverages, and maximum_source_panel_coverage. The decision policy reports positive_marker_coverage_basis=same_source_marker_panel_with_minimum_supporting_genes; neither a candidate nor a preferred CL ancestor can qualify from a single marker or from coverage assembled across unrelated sources.

An abbreviated result looks like this:

{
  "annotation": {
    "schema_version": "1.0",
    "dataset_id": "example",
    "cluster_id": "2",
    "species_taxon_id": 9606,
    "tissue": {"id": "UBERON:0000178", "label": "blood"},
    "decision": "parent_only",
    "predicted_cell_type": {"id": "CL:0000236", "label": "B cell"},
    "confidence": 0.63,
    "evidence": [
      {
        "direction": "supporting",
        "evidence_type": "marker",
        "source_id": "cell_marker_accordion",
        "source_release": "1.0.0@b506dd515d42",
        "record_id": "cma:1.0.0@b506dd515d42:b0a606ba685a4ac646b38d7e",
        "statement": "cell_marker_accordion records CD79A as a positive marker for B cell with evidence grade B. Immediate-source citation: PMID 40623970; DOI 10.1038/s41467-025-60900-4.",
        "genes": ["CD79A"],
        "evidence_family": "cma_upstream:bdbiosciences"
      }
    ],
    "alternatives": [],
    "warnings": ["Confidence is an uncalibrated marker-overlap ranking score."]
  },
  "validation": {
    "valid": true,
    "errors": [],
    "provenance_summary": {
      "source_ids": ["cell_marker_accordion"]
    }
  }
}

See Annotation output contract for the complete semantics and failure rules.

Validate a single-cell dataset

Install the validation extra, then run:

cell-annotation-admin validate-dataset \
  --dataset-id GSE123813 \
  --source-id geo \
  --source-record-id GSE123813 \
  --species-taxon-id 9606 \
  --tissue-id UBERON:0002097 \
  --tissue-label skin \
  --matrix expression.h5 \
  --metadata cell_metadata.tsv \
  --full-scan \
  --manifest-output validation/dataset_manifest.json

The validator checks the 10x CSC structure, dimensions, sparse pointers and indices, numeric values, barcode alignment, feature metadata, checksums, and expression scale. It writes a manifest only when an output path is requested and will not overwrite one without --overwrite.

Development checks

make check
make full-check
make wheel-audit

make check compiles the code, validates the source registry, both packaged ontology artifacts, official HGNC/MGI nomenclature, every packaged marker layer through the Heimli pediatric-human-thymus branch, every staged literature build through final schema 43, atlas publication metadata, derived marker artifacts, CELLxGENE metadata and Census indexes, checks documentation, and runs the unit suite. make full-check also launches the real stdio server through the official MCP client and exercises all tools and resources in Python 3.11. make wheel-audit builds an isolated no-dependency wheel, compares its runtime package inventory with the source tree, checks the Python requirement and console entry points, loads the service only from the extracted wheel, and pins all release counts.

The current final schema-43 audit covers 366 unit tests and a 224-file wheel containing all 165 registered sources, 56 runtime-integrated sources, 215,282 primary marker assertions, 1,091 reviewed literature claims, and 16 release-gate-complete species-organ contexts. Exact release wheel size and checksum are recorded in Validation so the README does not create a self-referential package hash.

The local GSE123813 files are not part of the repository. If they are present under the expected filenames, run the full data and benchmark workflow with:

make bcc-validation

The same-study regression, earlier five-tissue Cell Marker Accordion benchmark, pan-tissue fallback benchmark, standalone HRA source-isolation benchmark, Tabula Muris Senis policy regression, HCL all-panel regression, MCA multi-organ service regression, source-specific organ regressions, and current mixed-source tissue-scope regression remain engineering baselines, not independent-cohort claims. The fallback benchmark evaluates 22 mapped post-treatment skin clusters: 13 are annotated, all 13 are hierarchy-compatible with the broad reviewed mapping, and nine abstain. The historical HRA-only benchmark annotates 14 of 22 clusters; five are exact and all 14 are hierarchy-compatible, while eight abstain. Each organ regression freezes complete-panel calls, one-marker abstention, exact source/tissue/ontology provenance, literature linkage, and applicable dependency or license boundaries. validation/mouse_nasopharynx_panel_regression.json adds eight complete-panel cases, one Ptprc-only abstention, and one mixed B/myeloid negative-evidence redirect. validation/human_brain_siletti_regression.json adds a clean oligodendrocyte panel, RBFOX3-only abstention, and a PDGFRA negative-conflict penalty. These thresholds and scores are engineering fixtures, not calibrated probabilities or recommended universal cutoffs. Exact and hierarchy-compatible metrics must not be conflated. External benchmarks must be rerun and reinterpreted whenever a primary marker artifact or decision policy changes; current release qualification is tracked in Validation.

validation/mouse_rectum_panel_regression.json adds the two-case normal mouse rectum baseline: the complete four-gene panel must return CL:1001595, while EPCAM alone must remain unknown. It also freezes reviewed-claim linkage, wound/non-rectum quarantine, unreported replicate and pooling fields, the GEO license boundary, and one-family evidence accounting. This is same-artifact policy validation, not an independent biological benchmark.

validation/human_parathyroid_panel_regression.json adds four normal-human parathyroid cases: complete chief-cell, endothelial, and fibroblast panels must return their exact CL identities, while PTH alone must remain unknown at 10% coverage. It freezes reviewed-claim linkage, the normal-gland boundary, adenoma/subtype/mixed-cluster quarantine, the unavailable OMIX file state, CC BY-NC and academic-use terms, and one-family evidence accounting. This is same-artifact policy validation, not held-out accuracy or confidence calibration.

validation/mouse_parathyroid_panel_regression.json adds ten mouse-parathyroid cases: all eight complete author-DE panels must return their exact broad CL identities, while Pth alone and Nkg7 alone must remain unknown. The chief-cell case links eight experimental single-cell claims plus two independently published spatial PTH/GCM2 claims and therefore reports two evidence families; all other panels remain one-family evidence. The regression preserves the chimeric-source warning, 50% coverage gate, Cluster 7 quarantine, pooled-library design, non-cell-resolved spatial boundary, licenses, absent negative evidence, absent held-out validation, and uncalibrated confidence.

The tissue-scope regression reuses six B/T/NK clusters from the provided GSE123813 skin data with allow_pan_tissue_fallback=false. Skin now resolves to combined HRA and HPA organ-specific contexts. At the conservative 0.45 threshold, four broad T-lineage clusters are called and two clusters abstain. Most importantly, an NK cluster whose organ-only HPA score favored T cell is forced to unknown because the non-promoting pan-tissue audit finds a stronger ontology-incompatible NK signature. All six outputs validate, organ evidence remains source-traceable, and otherwise matching reviewed blood claims remain explicit tissue-scope mismatches with zero direct literature support. The compact regression is stored in validation/literature_bcc_tissue_gate_validation.json; the standalone HRA and fallback artifacts remain source-isolation baselines.

PostgreSQL

Apply migrations in order:

psql "$CELL_ANNOTATION_DATABASE_URL" -v ON_ERROR_STOP=1 \
  -f sql/001_initial_schema.sql
psql "$CELL_ANNOTATION_DATABASE_URL" -v ON_ERROR_STOP=1 \
  -f sql/002_ontology_identifiers.sql
psql "$CELL_ANNOTATION_DATABASE_URL" -v ON_ERROR_STOP=1 \
  -f sql/003_marker_evidence_family.sql

Validate or import the source registry:

cell-annotation-admin import-registry --dry-run
cell-annotation-admin import-registry

The committed import is an idempotent upsert. It does not delete database rows that are absent from the TSV.

Documentation

Data, attribution, and release status

The packaged Cell Ontology term and edge artifacts are derived from the CL basic OBO release dated 2026-06-08 and are distributed under CC BY 4.0. Their URL, source checksum, normalized checksums, release, and record counts are stored in cl_manifest.json.

The packaged UBERON term and relationship artifacts are derived from the UBERON basic release dated 2026-06-19 and are distributed under CC BY 3.0. Their source and normalized checksums, relationship counts, and release are stored in uberon_manifest.json.

Cell Marker Accordion is pinned to an official repository commit and its repository MIT license is included with the package. Cell Marker Accordion aggregates other resources, so every normalized assertion preserves the upstream resource as an evidence family. Source-registry entries remain discovery metadata, not a statement that every catalogued payload may be redistributed.

HRA ASCT+B is pinned to the v2.5 release and distributed under CC BY 4.0. Each normalized row preserves its versioned table release, HBM identifier, table DOI, and source-row locator. Full attribution and evidence boundaries are recorded in Third-party data notices.

Tabula Muris Senis is pinned to Figshare item 12654728.v1 under the MIT license. Raw H5AD files remain local and immutable; the package contains only aggregate grade-C marker assertions, file checksums, derivation thresholds, aggregate cell counts, publication identifiers, and evidence boundaries. See Third-party data notices.

Human Protein Atlas is pinned to v25.1. The package contains only normalized grade-C marker assertions and citation metadata derived from official downloads; it does not redistribute cell-level observations, raw sequencing files, or article full text. HPA copyrightable database content is CC BY 4.0, but HPA incorporates third-party datasets and their applicable terms remain in force. Every assertion therefore preserves its upstream PubMed family and source dataset lineage. See Third-party data notices.

Mouse Cell Atlas is pinned to official Figshare item 5435866.v8 under CC BY 4.0. One immutable snapshot holds the published Figure 2 matrix, exact 98-cluster mapping, and batch metadata; a second holds the detailed author assignments and paired per-batch matrices used only for the adult-testis extension. The recomputed H5AD is not assigned incompatible Figure 2 labels. Raw cell data are not packaged. The runtime contains 2,178 ontology-normalized aggregate grade-C rows, all branches remain one study family, single-batch limitations stay visible, and absence is never interpreted as a negative marker. See Third-party data notices.

The healthy-mouse gingiva layer pins one official CELLxGENE dataset version and the corresponding CC BY 4.0 primary publication. Its 1.14 GB H5AD remains an immutable local build input and is not packaged. The distributed artifact contains only 240 aggregate marker assertions, provenance, derivation statistics, and citation metadata; all panels remain one pooled biological replicate and one evidence family. See Third-party data notices.

The mouse-caecum layer pins one official CELLxGENE dataset version and the corresponding CC BY primary publication. Its 134 MB H5AD remains an immutable local build input and is not packaged. The distributed artifact contains 420 aggregate caecum marker assertions, semantic-quarantine records, derivation statistics, and citation metadata. All panels remain one pooled experimental dataset and one publication family; they are not presented as a healthy epithelial colon atlas. See Third-party data notices.

The human oral layer pins one official CELLxGENE dataset version but uses it only as a local processing surface for the exact GSE164241 buccal branch. The distributed package contains aggregate marker rows and citation metadata, not the human cell-level H5AD. The primary DOI is 10.1016/j.cell.2021.05.013; the integrated processing DOI is 10.1016/j.cpblue.2026.100007. Upstream data terms remain applicable, so this source is registered as license_review and every returned evidence record retains that warning. See Third-party data notices.

The human nasal layer likewise distributes only normalized aggregate marker rows, manifests, and citation/claim metadata. Its 3.57 GB CELLxGENE H5AD remains an ignored local build input; the EGA participant-data accession and upstream terms are retained in provenance. The article is CC BY 4.0, but that article license is not treated as a blanket redistribution license for participant-derived data. Runtime evidence therefore remains license_review. See Third-party data notices.

The mouse nasal layer distributes only normalized aggregate marker rows, manifests, and citation/claim metadata. The official GSE245074 SOFT and nested Seurat RDS remain ignored immutable local build inputs. GEO does not state an explicit dataset license and the author manuscript is not in the PMC Open Access subset, so no article license is extended to the data. Runtime evidence remains license_review; FACS frequency bias and the one-family, no-negative, no-held-out, no-independent, and uncalibrated boundaries stay explicit. See Third-party data notices.

The mouse-oviduct layer pins one official CELLxGENE dataset version and GSE164291. Its 538 MB H5AD remains an immutable local build input and is not packaged. The distributed artifact contains 180 aggregate physiological-estrus oviduct marker assertions, derivation statistics, and PMID 33818810 citation metadata. All panels remain one sample pooled from five mice and one publication family. Because an explicit aggregate-data redistribution license was not established, the source is registered as license_review and every returned evidence record retains that warning. See Third-party data notices.

The mouse-gallbladder layer uses the CC BY 4.0 article and figures for GSE179524. Selected source figures and GEO metadata remain immutable local audit inputs; no article full text or raw patient data are packaged. The distributed artifact contains 32 manually reviewed aggregate marker rows and matching claim records with exact figure locators. All panels remain one mixed-condition publication family and are not quantitative, independent, negative-marker, held-out, or calibrated evidence. See Third-party data notices.

The mouse-bronchus layer uses the CC BY 4.0 LungMAP CellCards article as a static expert synthesis. The package distributes only 13 normalized panel-component rows, the citation record, reviewed claim rows, and checksummed manifests; the local BioC text and publisher metadata snapshots are not packaged. All rows remain one derived CellCards family and are not raw-expression, differential-expression, negative-marker, independent-study, held-out, or calibrated evidence. See Third-party data notices.

The mouse-nasopharynx layer distributes only 27 normalized assertion rows, six citation records, 27 project-authored claim rows, and checksummed manifests. GSE227311 normalized data, article representations, and source snapshots remain immutable local build inputs and are not packaged. The PLoS article is CC BY; the Nature article is CC BY 4.0, while GSE227311 data reuse and the PNAS manuscript remain license review. See Third-party data notices.

The La Manno developing-mouse-brain layer distributes only 34 normalized author-rule components, one citation record, 34 project-authored claim rows, and checksummed manifests. The rule repository archive, PubMed XML, Supplementary Table 2, cells, expression data, and article content remain immutable local build inputs and are not packaged. The rule repository has no stated license and the Nature article and supplement remain under publisher terms. The runtime exposes the E7-E18 developmental scope, one-family accounting, six quarantined stable rules, and absent adult, per-embryo, independent, held-out, and calibrated evidence. See Third-party data notices.

The Bandyopadhyay human-bone-marrow layer distributes only one citation record, twelve project-authored exact-row claims, and checksummed manifests. The PubMed and GEO records, BioStudies inventory, and Cell Supplementary Table S2 workbook remain immutable local build inputs and are not packaged. The article and supplement require publisher-license review, the BioStudies mirror inherits source terms, and GEO states no separate redistribution license. The runtime exposes one-family accounting and the absence of negative, independent, held-out, and calibrated evidence. See Third-party data notices.

The Madissoon human-spleen layer distributes only one existing citation record, five project-authored Figure S9c claim rows, and checksummed manifests. PubMed, Europe PMC, PMC OA, ENA, the CC BY 4.0 supplement, and its embedded panel remain immutable local build inputs and are not packaged. ENA exposes no separate license field. The runtime exposes one-family accounting, qualitative-figure boundaries, and the absence of scoring, negative, independent, held-out, and calibrated evidence. See Third-party data notices.

The Abe human-lymph-node layer distributes one citation record, eleven project-authored exact-row claims, one project-authored context caution, and checksummed manifests. PubMed, Europe PMC, PMC OA, and the CC BY 4.0 workbook remain immutable local build inputs and are not packaged. EGAD00001008311 is controlled access, and no patient-level or expression data were acquired. The runtime exposes exact table provenance, one-family accounting, the non-healthy-volunteer sampling boundary, and the absence of scoring additions, independent replication, held-out evaluation, or calibrated probability. See Third-party data notices.

The Park/Ransick mouse-kidney layer distributes 35 normalized positive grade-C rows, two citation records, 35 project-authored one-to-one claims, and checksummed manifests. The eleven PubMed/PMC/article/supplement source files remain immutable local inputs and are not packaged; the PMC OA API records no open redistribution license for either NIH author manuscript. Park and Ransick count as two studies, while batches, animals, sexes, anatomical zones, assays, tables, and genes within a study remain dependent. No background, low, missing, or unreported expression is converted into a negative marker. Runtime output exposes exact table provenance and two-family podocyte support. The pair remains positive-only; the separate He layer supplies the organ-level negative-evidence gate. See Third-party data notices.

The He mouse-kidney layer distributes four normalized positive mesangial components, the source-explicit Ptprc exclusion, one citation record, six reviewed claims, and checksummed manifests. The official article and Supplementary Data 8 are CC BY 4.0 immutable local audit inputs and are not packaged; no human patient-level or expression data were acquired. Source Sept4 is normalized to current MGI Septin4, while the failed-CD31-sort Pecam1 observation is retained as a non-scoring conflict. All rows remain one GSE160048/PMID33837218 family. Runtime output exposes candidate-local 50% panel coverage, conservative mesangial-parent selection, single-marker abstention, explicit contradiction output, formal mouse-kidney gate completion, and the absence of independent validation, held-out accuracy, or calibrated probability. See Third-party data notices.

The mouse-esophagus layer uses GSE254855 annotations and UMI counts as immutable local inputs. The package contains only 260 aggregate marker assertions, provenance statistics, citation metadata, and 39 selected project-recomputed claims; it does not redistribute the matrix, cell metadata, or article full text. The source is marked license_review, all rows remain one publication family, and no claim is presented as author-supplied differential expression, independent replication, negative evidence, held-out accuracy, or calibrated confidence. See Third-party data notices.

The mouse-submandibular-gland layer uses the MIT-licensed Figshare 13157726.v2 RDS as an immutable local input. The package contains only 226 aggregate marker assertions, provenance statistics, citation metadata, and 24 project-recomputed claims; it does not redistribute the RDS or adapt the CC BY-NC-ND article. All rows remain one female sample and one publication family, and no claim is presented as author-supplied differential expression, independent replication, negative evidence, held-out accuracy, or calibrated confidence. See Third-party data notices.

The ScTypeDB lookup index pins the official repository at commit 630e15cf1e51f2612eda4ad0406dfb17503fa8c9 and includes the upstream GPL-3.0 license text in the wheel. The source workbook remains an immutable local acquisition; the package distributes only the deterministic normalized panel TSV and its manifest. Because the source is a derived database without per-marker upstream citation lineage, these panels are reference-tool output rather than annotation evidence and cannot affect ranking, validation confidence, evidence-family counts, or organ release gates. See Third-party data notices.

No repository software license has been selected yet. Do not assume permission to redistribute the code until a license file is added by the project owner.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables deep probabilistic analysis of single-cell omics data using scvi-tools through natural language. Supports SCVI for scRNA-seq analysis, SCANVI for cell type annotation, TOTALVI for multi-modal RNA/protein data, and PEAKVI for scATAC-seq analysis.
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    An MCP server that enables scRNA-Seq analysis through natural language, providing tools for data preprocessing, clustering, and biological visualization. It supports both predefined function execution and a flexible code mode powered by a Jupyter backend for automated single-cell transcriptomics workflows.
    16
    BSD 3-Clause
  • A
    license
    B
    quality
    A
    maintenance
    Enables LLM agents to query the CZ CELLxGENE Census single-cell atlas with ontology-aware filters, cost caps, and full provenance, allowing natural language questions about cell types, tissues, and gene expression.
    13
    MIT