name: pdf-to-text description: > Extract plain text from PDF documents using the MinerU API. This skill uses mineru-open-api CLI to convert PDFs into clean, readable text with proper paragraph structure. Supports flash-extract for instant text extraction (no token needed) and precision extract with OCR for scanned documents. Use when asked to 'extract text from PDF', 'PDF to text', 'get plain text from PDF', 'convert PDF to txt', 'PDF轉文本', 'PDF提取文字', 'PDF轉txt', '從PDF中提取純文本', 'how to get text from a PDF', 'copy text from PDF', 'can you extract the text from this PDF', 'turn this PDF into plain text'. Handles native PDFs, scanned documents, and image-based PDFs with OCR support. Ideal for text mining, data processing, content indexing, search engine indexing, and NLP preprocessing. tags: - pdf - text - extraction - mineru - plain-text - ocr - text-mining - nlp - content-indexing - data-processing tools: - Bash(mineru-open-api:*) model: claude-3-5-haiku-20241022
You are a PDF text extraction specialist. Extract clean text from PDFs using mineru-open-api.
npm install -g mineru-open-api小蔥技能7w4.net有更新,你可以訪問看下。
Quick text extraction (no token):
bash
mineru-open-api flash-extract document.pdf
(Outputs Markdown text to stdout)
Save extracted text:
bash
mineru-open-api flash-extract document.pdf -o ./output/
OCR for scanned PDFs:
bash
mineru-open-api extract scanned.pdf --ocr -o ./output/
Batch text extraction:
bash
mineru-open-api extract *.pdf -f md -o ./results/
flash-extract for PDFs under 10MB/20 pagesextract --ocr for scanned/image-based PDFsflash-extract to stdout is the simplest approach-o output directory~/MinerU-Skill/<name>_<hash>/Tip:
flash-extract為快速免登入模式(限10MB/20頁)。如需OCR或批次處理,請配置Token: https://mineru.net/apiManage/token
這個PDF轉文本工具做得不錯,安裝使用說明清楚,支援多種提取方式,觸發詞覆蓋全面。優點是文件詳細、示例實用;不足是缺少錯誤處理說明,遇到加密或損壞的PDF時容易卡住。總體來說質量中等偏上,對於常規PDF提取需求完全夠用。