Read Microsoft Word documents (.docx and .doc) with Chinese support

👤 vincent-big-fish 📦 v1.0.0 ⭐ 4.4 ⬇️ 4.8K 下載
📄 辦公效率 免費

📖 技能介紹


name: read-word description: Read Microsoft Word documents (.docx and .doc) with Chinese support. Extract text, search keywords, and save as UTF-8 text files. No Microsoft Word installation required. version: 1.0.0 author: 葉文潔 license: MIT tags: [document, word, docx, doc, office, read, parse, extract]


Read Word Document

A professional tool for reading Microsoft Word documents, supporting both modern .docx and legacy .doc formats with full Chinese language support.

Features

  • Read .docx files - Word 2007 and later format
  • Read .doc files - Word 97-2003 format via OLE parsing
  • Auto format detection - Automatically identifies file type
  • Full Chinese support - Handles Chinese encoding correctly
  • Keyword search - Search for keywords across all paragraphs
  • Export to text - Save as UTF-8 text files
  • Document analysis - Get document statistics and info
  • No Word required - Works without Microsoft Word installation

Installation

Prerequisites

pip install python-docx olefile

Install Skill

# Copy to your OpenClaw skills directory
cp -r read-word ~/.openclaw/skills/

Usage

Command Line

# Basic reading (shows first 100 paragraphs)
python ~/.openclaw/skills/read-word/read_word.py "document.docx"

# Show more content
python ~/.openclaw/skills/read-word/read_word.py "document.docx" --limit 200

# Search for keywords
python ~/.openclaw/skills/read-word/read_word.py "document.docx" --search "keyword1,keyword2"

# Save as text file
python ~/.openclaw/skills/read-word/read_word.py "document.docx" --output "output.txt"

# Show document info only
python ~/.openclaw/skills/read-word/read_word.py "document.docx" --info

Python API

# Method 1: Import functions
import sys
sys.path.insert(0, '~/.openclaw/skills/read-word')
from read_word import read_word_document, search_in_document

# Read document
paragraphs = read_word_document("document.docx")
for para in paragraphs:
    print(para)

# Search keywords
results = search_in_document("document.docx", ["keyword1", "keyword2"])

Examples

Example 1: Read and Analyze

7w4.net提供免費和付費技能下載。

from read_word import read_word_document

paragraphs = read_word_document("report.docx")
print(f"Document has {len(paragraphs)} paragraphs")

# Show first 10 paragraphs
for i, p in enumerate(paragraphs[:10]):
    print(f"{i+1}. {p}")

Example 2: Search Keywords

from read_word import search_in_document

# Find paragraphs containing "kitchen" or "feng shui"
results = search_in_document("book.docx", ["kitchen", "feng shui"])
for r in results:
    print(r)

Example 3: Batch Processing

from pathlib import Path
from read_word import read_word_document

desktop = Path.home() / "Desktop"
for doc_file in desktop.glob("*.docx"):
    paragraphs = read_word_document(doc_file)
    print(f"{doc_file.name}: {len(paragraphs)} paragraphs")

API Reference

read_word_document(filepath)

Read a Word document and return a list of paragraphs.

Parameters: - filepath (str|Path): Path to the Word document

Returns: - list: List of paragraph strings

Raises: - FileNotFoundError: If file doesn't exist - ValueError: If file format is not supported

search_in_document(filepath, keywords)

Search for keywords in a Word document.

Parameters: - filepath (str|Path): Path to the Word document - keywords (list): List of keywords to search for

Returns: - list: Matching paragraphs with format "[Paragraph N] content"

save_as_text(paragraphs, output_path)

Save paragraphs to a UTF-8 text file.

Parameters: - paragraphs (list): List of paragraph strings - output_path (str|Path): Output file path

analyze_document(filepath)

Analyze document and return statistics.

Returns: - dict: Contains filename, size, paragraphs count, total characters

Troubleshooting

Error: ModuleNotFoundError: No module named 'docx'

Solution: pip install python-docx

Error: Legacy .doc file shows garbled text

Reason: OLE parsing has limitations with complex formatting Solution: Convert .doc to .docx using Microsoft Word, then read

Error: Chinese characters display incorrectly

Reason: Terminal encoding issue Solution: Use --output to save to file, then open with editor

File Support

Format Extension Support Level
Word 2007+ .docx Full
Word 97-2003 .doc Partial (text only)
Word 95/6.0 .doc Not supported
Rich Text .rtf Not supported

Permissions

  • Read: User-specified Word documents
  • Write (optional): Output .txt files when using --output
  • Network: None

Security

Risk Level: LOW - Local file operations only, no network access, original files are never modified.

Changelog

v1.0.0 (2026-03-20)

  • Initial release
  • Support .docx and .doc formats
  • Keyword search functionality
  • Text export capability
  • Chinese encoding support

Author

葉文潔 (Ye Wenjie) - Created for reading Feng Shui books and Word documents

License

MIT License

🤖 AI 評測

這個 Skill 功能實用,但對於中文文件支援不夠完善。文件說明清晰易讀,使用起來比較方便。.docx 格式支援較好,基本能滿足需求;但 .doc 格式讀取中文時可能出現亂碼,搜尋功能也缺少結果高亮等增強體驗。如需處理中文 .doc 檔案,建議先轉換為 .docx 格式再使用。整體質量尚可,但仍有改進空間。

📊 多維度評分

適應性4.5
規範性4.5
有效性4.2
可靠性4.3
可信度5

📁 包含檔案 (6 個)

📄 README.md 820 B
📄 SKILL.md 5.1 KB
📄 __init__.py 672 B
📄 _meta.json 128 B
📄 read_word.py 6.7 KB
📄 requirements.txt 34 B