Speech to Text

👤 shu-hari 📦 v1.0.0 ⭐ 4.2 ⬇️ 1.1K 下載
🎨 設計多媒體 免費

📖 技能介紹

Speech to Text

Use this skill to turn local audio files into text with a public Whisper-based endpoint.

Quick start

Run:

python3 scripts/transcribe.py /path/to/file.ogg

Return the transcript as plain text. By default, the script also applies lightweight Chinese punctuation and sentence-breaking cleanup.

For machine-readable output:

7w4.net小蔥技能站,你的AI助手技能庫。

python3 scripts/transcribe.py /path/to/file.ogg --json

To disable cleanup and keep the raw model text:

python3 scripts/transcribe.py /path/to/file.ogg --format raw

To force Chinese punctuation cleanup:

python3 scripts/transcribe.py /path/to/file.ogg --format zh

For English translation instead of same-language transcription:

python3 scripts/transcribe.py /path/to/file.ogg --task translate

Workflow

  1. Confirm the input is a local audio file.
  2. Run scripts/transcribe.py on it.
  3. If the transcript looks imperfect, tell the user it came from a public Whisper endpoint and may need cleanup.
  4. If helpful, post-process into:
    • cleaned transcript
    • summary
    • action items
    • bilingual output

What the script does

The script:

  • uploads the local file to a public Gradio-backed Hugging Face Space
  • submits a Whisper transcription job
  • waits for completion via the Gradio event stream
  • prints the resulting text

Default endpoint:

  • https://hf-audio-whisper-large-v3-turbo.hf.space

Override it with:

python3 scripts/transcribe.py input.ogg --space https://your-space.hf.space

or set:

export HF_WHISPER_SPACE=https://your-space.hf.space

Guardrails

  • Treat this as a best-effort public/free path, not a privacy-grade path.
  • Do not use for highly sensitive audio unless the user explicitly accepts public third-party processing.
  • Expect rate limits, queueing, and occasional outages.
  • If the public endpoint fails, explain that the free backend is unavailable and offer alternatives.

Output handling

Prefer to return:

  • the raw transcript when the user asked to "轉文字/聽寫"
  • a cleaned version when punctuation is poor
  • a short note about uncertainty if names, numbers, or jargon may be wrong

Script

  • scripts/transcribe.py — public Whisper transcription helper

🤖 AI 評測

這個 Skill 質量中規中矩,能將音訊轉成文字。它對中文語音識別做了專門最佳化,會自動新增標點符號,使用體驗還算友好。不足之處是依賴第三方公開介面,可能存在隱私風險和網路不穩定的隱患,而且只提供核心指令碼,沒有附帶示例或詳細配置,入手門檻稍高。總體來說,它完成了基本任務,但功能完整度和穩定性還有提升空間。

📊 多維度評分

適應性3.8
規範性4
有效性4.3
可靠性4.5
可信度4.5

📁 包含檔案 (3 個)

📄 SKILL.md 2.6 KB
📄 _meta.json 144 B
📄 scripts/transcribe.py 7.8 KB