Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses output semantics - that figures are tagged with evidence provenance and that multi-band score ranges omit a letter grade - which is real value beyond structure. However, it says nothing about invalid-slug handling, auth needs, or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.