Skip to main content
Glama
michalhron

Scopus Plus MCP

by michalhron

coding_agreement

Measure inter-coder agreement on filled coding sheets: compute Cohen's kappa with 95% interval, per-label agreement, confusion matrix, and disagreements; validate draft labels against coders.

Instructions

Inter-coder agreement on a coding sheet from path_transmission (or citation_network with edge_contexts) once two coders have filled their columns: Cohen's kappa with a 95% interval and its Landis & Koch reading, agreement per label, the confusion matrix and the disagreeing edges. Also scores the draft labels against each coder and against the coders' consensus, i.e. how far the heuristic can be trusted. Reads .csv (comma, semicolon or tab), as saved from Excel.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
pathYesThe filled coding sheet (.csv).
coder_columnsNoThe two coder columns (default coder_1_label, coder_2_label).
reference_columnNoLabels to validate against the coders (default draft_label; '' for none).draft_label

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full behavioral burden. It discloses that the tool is a read/analysis operation on an existing file and describes its outputs, but it does not state whether the file is modified, whether missing or partial coder columns cause errors, or what happens with fewer than two coders. For an analysis tool with no annotations, this is adequate but incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two moderately dense sentences front-load the core action and then list outputs and the alternative use case. The parenthetical caveat about CSV separators is useful but slightly tacked-on; overall efficient with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the input provenance, required format, computed statistics, and secondary validation use, which is more than the schema alone. With no output schema, describing the returned metrics and matrix is necessary and done. It could more fully explain edge cases (e.g., fewer than two coders, missing labels), which keeps it below 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents each parameter's meaning and defaults. The description adds the contextual fact that coder columns come from the two coders and that the reference column is the draft label, but does not add new syntax or format detail beyond the schema. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (computes inter-coder agreement) on a specific resource (a coding sheet produced by path_transmission or citation_network). It enumerates the exact outputs (Cohen's kappa with 95% interval, Landis & Koch reading, per-label agreement, confusion matrix, disagreeing edges) and a secondary function (validating draft labels against coders). No sibling tool shares this purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear precondition: 'once two coders have filled their columns' and points to the tools that produce the input (path_transmission, or citation_network with edge_contexts). It also defines the format the sheet must be in (CSV comma/semicolon/tab, as saved from Excel). It doesn't state when NOT to use this tool, but the prerequisites are explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.