Skip to main content
Glama

klax_index_learning_material

Extract text from downloaded PDF/PPTX files and index each page for local search, skipping files already indexed by content hash.

Instructions

로컬에 이미 다운로드된 PDF/PPTX 파일을 읽어 페이지/슬라이드 단위로 텍스트를 추출하고 로컬 검색 인덱스에 반영합니다.

KLAS에는 어떤 것도 쓰지 않으며, klax_download_material로 받은 로컬 경로만 처리합니다. 동일 파일 내용(SHA256)이 이미 색인되어 있으면 재처리하지 않고 기존 결과를 반환합니다 (force_reindex=True로 강제 재처리 가능). PyMuPDF/python-pptx 미설치, 스캔본(텍스트 없음), 미지원 확장자는 각각 unsupported/needs_ocr 상태와 사유(error_reason)로 반환됩니다.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
titleNo
semesterNo
course_idYes
local_pathYes
source_urlNo
force_reindexNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full disclosure burden and does well: it discloses deduplication via SHA256 with existing-result return, the force_reindex=True escape hatch, and the unsupported/needs_ocr states with error_reason for missing libraries, scanned files, and unsupported extensions. This is rich edge-case transparency, though it omits the return payload shape and any permission/rate considerations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with the primary purpose front-loaded in the first line. The following sentences earn their place by covering deduplication, force_reindex, and failure modes without redundancy. Appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Behavioral coverage is solid and an output schema exists, so return values need not be described. However, with 0% parameter coverage in the schema and a required course_id left unexplained, plus no explicit when-not-to-use guidance, the definition is not fully complete for a 6-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it only explains two of six parameters: local_path (processes paths from klax_download_material) and force_reindex. The required course_id and the optional title, semester, and source_url are left entirely unexplained, creating a real gap for a required parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise verb+resource+action: reads locally downloaded PDF/PPTX files, extracts per-page/slide text, and reflects it into the local search index. It clearly distinguishes itself from siblings like klax_download_material (downloads) and klax_search_learning_materials (searches), and explicitly notes it does not write to KLAS.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: it only processes local paths obtained from klax_download_material and never writes to KLAS. However, it does not explicitly name sibling alternatives or state when-not-to-use, leaving the routing mostly implicit rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.