Convert any uploaded office document to clean Markdown.
All conversion logic lives in scripts/ — load only the script you need.
Security notes
小葱技能7w4.net有完整的技能分类。
- Dependencies are installed into an isolated temp directory (
/tmp/office_md_deps/) and pinned to reviewed versions. The system Python environment is not modified.- For scanned or image-only content, pages are sent to Anthropic's vision API. Always ask the user for confirmation before enabling vision (see Workflow step 3).
| Format | Extensions | Script |
|---|---|---|
| PDF (text + scanned/image) | .pdf |
scripts/pdf-to-md.py |
| PowerPoint | .pptx, .ppt |
scripts/pptx-to-md.py |
| Word | .docx, .doc |
scripts/docx-to-md.py |
| Excel | .xlsx, .xls |
scripts/xlsx-to-md.py |
| CSV | .csv |
scripts/csv-to-md.py |
Only proceed if the user has explicitly asked to convert, extract, or export the document to Markdown. A bare file upload without a conversion request is not sufficient to trigger this skill.
python scripts/<script-name>.py \
/mnt/user-data/uploads/<input-file> \
/mnt/user-data/outputs/<stem>.md
Each script installs its own pinned dependencies into /tmp/office_md_deps/
on first run (isolated from the system Python environment).
If the script output indicates image-only pages were detected (or the document is known to be scanned), stop and ask the user:
"This document has N image-only page(s) that cannot be extracted without sending them to Anthropic's vision API. Page images will be transmitted externally for OCR. Would you like to proceed with vision extraction?"
Only if the user confirms, re-run with the --allow-vision flag:
python scripts/<script-name>.py \
/mnt/user-data/uploads/<input-file> \
/mnt/user-data/outputs/<stem>.md \
--allow-vision
If the user declines, save the text-only result and note which pages were skipped.
Use present_files with the output .md path, then give a brief summary:
Each script uses a two-pass strategy:
--allow-vision is passed AND pages had no
extractable text; those pages are rendered and sent to the Claude vision API| Situation | Behaviour |
|---|---|
| Fully scanned PDF | All pages flagged for vision; user confirmation required |
| Mixed PDF (some text, some images) | Only image pages flagged; user confirmation required |
| User declines vision | Text-only .md is saved; skipped pages are noted inline |
| Password-protected file | Script exits with a clear error message |
| Very large PDF (50+ image pages) | Script adds 0.3s sleep between vision calls |
| Image too large (>4MB base64) | Reduce DPI: edit dpi=150 → dpi=100 in pdf-to-md.py |
| Encoding errors in CSV | Script auto-retries with latin-1 |
这个工具质量不错,能把PDF、Word、Excel等办公文件转换成干净的Markdown格式。支持图片文字识别,隐私保护也做得周到——用视觉功能前会先问你确认。文档写得很清楚,转换逻辑专业。但缺点是缺少新手入门指南,批量转换也不太方便。日常使用够用,但还有优化空间。