run_prompt
Run prompts through Gemma 2 2B and report active SAE features per token and layer, with Neuronpedia labels and activation rankings.
Instructions
Run a prompt through the model and report which SAE features are active.
Returns, per layer:
- top_features: the 10 features most strongly active anywhere in the text,
each at its peak token, with its Neuronpedia label, density (fraction of
all tokens it fires on) and the output tokens it promotes.
- tokens: for each token, its top_k features.
`act` is the raw activation. `rel` is act / the feature's typical max
activation, so rel near 1 means firing about as hard as it ever does.
Rankings use rel. Features active on >10% of all tokens carry little
meaning and are hidden unless include_dense is true. The <bos> token is
omitted. The first real token often shows position artifacts: features
with very large activations unrelated to its meaning.
Args:
prompt: Text to run. Gemma 2 2B is a base model, not a chat model.
layers: Residual-stream layers to read (0-25). Default [12].
top_k: Features to list per token.
max_new_tokens: If > 0, greedily generate this many tokens first and
analyze prompt + completion.
include_dense: Include features active on >10% of tokens.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| top_k | No | ||
| layers | No | ||
| prompt | Yes | ||
| include_dense | No | ||
| max_new_tokens | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||