noisefloor
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| ab_testA | Can you call a winner on an A/B test yet? Uses anytime-valid confidence sequences, so it is SAFE TO RUN AFTER EVERY OBSERVATION — peeking does not inflate the false-positive rate the way a t-test or z-test does. |
| did_it_changeA | Did a metric actually change, or is the move noise? Detects both rises AND collapses against the metric's own history, with a stated false-alarm rate and no assumption about the distribution. |
| forecast_nextC | What should the next reading be, and within what range? Range adapts to the metric's recent volatility and stays valid even when the metric shifts. |
| real_or_samplingA | Did the metric move, or did the sample size underneath it move? Run this before reporting any RATE as a change — conversion rates, error rates and click-through all shift when the denominator shifts, for reasons that have nothing to do with the thing being measured. |
| score_forecastsA | How good would these forecasts actually have been? Grades every prediction the tool would have made over the history, using only what was known at the time, and reports calibration plus the worst misses. |
| which_metrics_matterB | You watch many metrics; which genuinely stand out? Controls the false discovery rate across all of them at once, which per-metric thresholds do not: forty metrics each alerting wrongly 5% of the time means two false alarms every round. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 6 tools
Each tool targets a distinct statistical question: A/B testing, change detection, forecasting, denominator shift, forecast evaluation, and multiple testing correction. No overlap in purposes.
All tool names follow a consistent verb_noun pattern in lowercase snake_case, e.g., did_it_change, forecast_next, which_metrics_matter. No deviations.
6 tools is an ideal number for a focused statistics toolkit, covering essential operations without being overwhelming or sparse.
The set covers key statistical tasks (A/B testing, change detection, forecasting, multiple testing), but lacks a sample size/power analysis tool, which is a minor gap.