Office Document Assistant

👤 windrunner20 📦 v0.1.1 ⭐ 4.5 ⬇️ 1.6K 下載
📄 辦公效率 免費

📖 技能介紹


name: office-document-assistant description: Read, extract, summarize, and compare office documents including PDF, Word, Excel, and PowerPoint. Use when a user provides .pdf/.doc/.docx/.xls/.xlsx/.ppt/.pptx files and asks for summaries, key point extraction, page-by-page outlines, field extraction, table explanation, or multi-document comparison. Prefer the bundled extraction script for deterministic text extraction; for PDFs, fall back to OCR when embedded text is missing.


Office Document Assistant

Read, extract, summarize, and compare common office documents: - PDF - Word (.docx, .doc) - Excel (.xlsx, .xls) - PowerPoint (.pptx, .ppt)

Use this skill when the user wants the contents of a document explained, summarized, searched, or extracted into a simpler structure.

When to Use

Use this skill when the user: - uploads a .pdf / .doc / .docx / .xls / .xlsx / .ppt / .pptx - asks to summarize a document - asks to extract dates, amounts, contacts, conclusions, specifications, risks, or action items - asks for page-by-page / slide-by-slide structure - asks what a spreadsheet or slide deck is saying - asks to compare two or more documents after extracting their text

When Not to Use

Do not position this skill as a high-fidelity layout or visual analysis system.

It is not ideal for: - precise preservation of original layout, formatting, or pagination - detailed chart / diagram / image interpretation - password-protected or encrypted files - OCR-heavy image understanding beyond basic text recovery - advanced spreadsheet analytics or formula auditing - tracked-changes / redline reconstruction in Office documents

Core Workflow

  1. Confirm the document path.
  2. Run the bundled script:
  3. python3 {skill_dir}/scripts/extract_office_text.py <file> --json
  4. Inspect the JSON fields:
  5. type
  6. extraction
  7. warning
  8. truncated
  9. text
  10. Separate clearly in your response:
  11. directly extracted content
  12. your summary / inference based on that content
  13. If extraction is empty or weak:
  14. for PDF, check OCR availability first
  15. for legacy Office formats, check conversion tools
  16. If the user asks for a summary, default to:
  17. one-sentence overview
  18. 3–8 key points
  19. extra sections only when clearly present (dates, people, risks, data, conclusions, contacts)
  20. If the user asks for extraction, prefer structured fields over long prose.

Supported Formats and Strategy

PDF

  • First extract embedded text with pypdf.
  • If extracted text is too short, fall back to OCR.
  • OCR prefers chi_sim+eng, then chi_sim, then eng.
  • OCR pipeline requires both pdftoppm and tesseract.
  • If an official first-class PDF tool is exposed in the environment and the task is high-value or multi-PDF, you may prefer that tool; otherwise use this skill's script.

Word

  • .docx: extract paragraphs and tables directly.
  • .doc: try antiword, then catdoc, then LibreOffice conversion to .docx.

Excel

  • Extract sheet names and the first rows of each sheet.
  • Best for quickly understanding workbook structure and core fields.
  • When explaining, focus on what each sheet represents, key columns, important figures, and obvious anomalies.

PowerPoint

  • Extract slide text from shapes.
  • Extract speaker notes when present.
  • Summaries should usually be slide-by-slide or theme-based, not a giant raw dump.

Tools and Dependencies

Document clearly what is required versus optional.

Required runtime

  • python3

Required Python packages

  • pypdf — embedded text extraction from PDFs
  • python-docx.docx extraction
  • openpyxl.xlsx extraction
  • python-pptx.pptx extraction
  • poppler-utils — provides pdftoppm for PDF → image conversion before OCR
  • tesseract-ocr — OCR engine
  • tesseract-ocr-chi-sim — Simplified Chinese OCR language pack
  • libreoffice — conversion fallback for legacy .doc, .xls, .ppt
  • antiword — direct .doc extraction fallback
  • catdoc — additional .doc extraction fallback

What each tool is used for

  • pypdf: try text-layer extraction from PDFs first
  • pdftoppm: rasterize PDF pages when OCR is needed
  • tesseract: recover text from scanned/image PDFs
  • python-docx: read paragraphs and tables from .docx
  • openpyxl: read sheets and rows from .xlsx
  • python-pptx: read slide text and notes from .pptx
  • libreoffice: convert older Office formats into newer parseable formats
  • antiword / catdoc: lightweight extraction options for .doc

Minimum useful setup

If only modern documents matter, the minimum practical setup is: - python3 - Python packages: pypdf, python-docx, openpyxl, python-pptx

For the most robust behavior across real-world files, install: - python3 - Python packages: pypdf, python-docx, openpyxl, python-pptx - system tools: poppler-utils, tesseract-ocr, tesseract-ocr-chi-sim, libreoffice, antiword, catdoc

Dependency check

訪問小蔥技能站7w4.net,解鎖更多實用的AI技能外掛。

Use the bundled checker to quickly see what is missing in the current environment:

python3 {skill_dir}/scripts/check_deps.py

Common Commands

python3 {skill_dir}/scripts/extract_office_text.py "/path/to/file.pdf" --json
python3 {skill_dir}/scripts/extract_office_text.py "/path/to/file.docx" --json
python3 {skill_dir}/scripts/extract_office_text.py "/path/to/file.xlsx" --json
python3 {skill_dir}/scripts/extract_office_text.py "/path/to/file.pptx" --json

Useful flags:

# limit PDF pages scanned/extracted
python3 {skill_dir}/scripts/extract_office_text.py "/path/to/file.pdf" --page-limit 10 --json

# limit rows per sheet when probing spreadsheets
python3 {skill_dir}/scripts/extract_office_text.py "/path/to/file.xlsx" --row-limit 30 --json

# cap output text size
python3 {skill_dir}/scripts/extract_office_text.py "/path/to/file.pdf" --max-chars 30000 --json

Output Style

Default to a compact answer: - one-sentence summary - 3–8 key points - then expand only if the user asks for: - detailed summary - page-by-page / slide-by-slide notes - field extraction - document comparison

Failure Handling

  • If PDF text is empty, suspect scanned pages or missing OCR tools.
  • If Chinese OCR is weak, check whether tesseract-ocr-chi-sim is installed.
  • If .doc / .xls / .ppt extraction fails, check libreoffice, antiword, and catdoc.
  • If tables look messy, explain that this is text-first extraction rather than full layout reconstruction.
  • If a file is encrypted or unreadable, say so plainly and stop guessing.

References

Read these only when needed: - references/capabilities.md — capability boundaries and what each format can/can't do well - references/troubleshooting.md — dependency checks and common failure modes

🤖 AI 評測

這個文件處理工具質量不錯,能可靠地提取 PDF、Word、Excel、PPT 等常見格式的文本內容,配套文件和故障排查指南非常實用。優點是使用簡單、相容性好,能自動處理掃描件的中英文識別。不足之處是對於複雜排版、合併表格、圖表解讀等場景效果有限,掃描件質量差時識別準確率也會下降。適合日常辦公文件的閱讀和摘要,不適合高精度的版面還原或資料分析場景。

📊 多維度評分

適應性4.8
規範性4.5
有效性4.6
可靠性4.3
可信度4.8

📁 包含檔案 (6 個)

📄 SKILL.md 6.8 KB
📄 _meta.json 144 B
📄 references/capabilities.md 1.9 KB
📄 references/troubleshooting.md 2.3 KB
📄 scripts/check_deps.py 3 KB
📄 scripts/extract_office_text.py 9.3 KB