download_document
MyDART MCP의 download_document 도구는 공시 원문을 마크다운으로 읽고 검색합니다(heading·표 구조 보존). 사업보고서 ZIP 안의 감사보고서 선택 포함.
[Purpose]
Full disclosure text; audit analysis (감사의견·KAM): get_audit_report.
비상장 non-filers: the F 감사보고서 rcept IS the route to 재무제표·주석.
HWP·PDF attachments·sibling docs (정관·내부회계 운영실태보고서): get_attachments. XBRL: get_financials(rcept_no).
[Usage]
"이 사업보고서 원문 읽어줘" → rcept_no="20260310002820"
"연결 감사보고서 본문" → documents[].role=consolidated_audit → doc_index
"본문이 잘렸어, 전체로" → truncate_at=2000000
"원문에서 횡령 찾아줘" → find="횡령"
[Response]
role: main_body / separate_audit / consolidated_audit / internal_control
documents[]: metadata only — content ONLY with all_docs=true
[Rules]
A 사업보고서 ZIP holds MULTIPLE docs — 감사의견·KAM·내부회계 live in the 별도(_00760)/연결(_00761) sub-docs, not index 0: check documents[].role first.
find returns 표 단위 발췌 for EVERY doc in the ZIP, not content. 0 matches ≠ absent — 스캔 이미지 표는 get_attachments(mode=images).
If truncated=true, raise truncate_at (compare char_count).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| find | No | Full-text search over EVERY 원문 XML in the ZIP (본문+별도/연결 감사보고서). Returns `find` (표 단위 발췌 bundles + 나머지 매치 위치 locations) INSTEAD of content — 전문을 나르지 않는다. Case-insensitive, 리터럴 매칭(정규식 아님 — `(주)카카오` 를 그대로 넣어도 된다). 0건이면 표기를 자간 띄운 형태(`핵 심 감 사 사 항`)로 한 번 더 찾는다. doc_index 를 함께 주면 그 문서 하나로 좁힌다(범위 밖이면 좁히지 않고 전 문서를 검색하고 notes 로 알린다 — 전문 조회의 index-0 폴백과 다르다: 잘못된 index 로 좁히면 답이 있는데 0건이 나간다). format=markdown 에서만 쓸 수 있다. | |
| format | No | Output format. markdown=DART XML → 마크다운, raw=original XML, text=tags stripped (table structure is lost) | markdown |
| all_docs | No | Convert every 원문 XML in the ZIP and return them as documents[] — the 사업보고서 본문 plus 별도/연결 감사보고서 in one call. | |
| rcept_no | Yes | 14-digit 접수번호 (the rcept_no from search_disclosures). Hyphens and spaces are stripped. | |
| doc_index | No | Selects one 원문 XML inside the ZIP (0-based; 0=본문 when omitted). Use the index from the documents list. The 감사보고서 body is usually index 1~2 (별도/연결). An out-of-range value silently falls back to the main body WITH a note in `notes` — check it. | |
| truncate_at | No | Max text length (the excess is cut). Default 300,000 chars; out-of-range values are clamped to the bound. With all_docs=true the budget is divided across the documents (the per-document share comes back as per_doc_truncate_at). Ignored when find is set (no content is returned). | |
| find_max_bytes | No | 발췌 본문 합계의 UTF-8 **바이트** 상한 (기본 12,000). 문자수가 아니라 바이트인 이유: 한글은 자당 3바이트라 '12,000자' 로 재면 응답이 36,000B 가 된다(실측). 단일 발췌가 이 값을 넘으면 표 행 단위로 잘리고 clipped=true 가 붙는다(제목행·구분행은 보존). | |
| find_max_bundles | No | Max 발췌 bundles (기본 5). 나머지 매치는 locations 로 위치만 나열된다. 바이트 예산(find_max_bytes)이 먼저 차면 이 수에 못 미칠 수 있다. |