Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), and the description adds substantial non-obvious behavior: the 0-100 scale is computed only over criteria actually measured, total_score is null with a reason when no criteria are given, and unscored signals (question-term overlap, avg sentence length, markdown) are also returned. This is context the annotations cannot convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.