Speech to Text

👤 shu-hari 📦 v1.0.0 ⭐ 4.2 ⬇️ 1.1K 下載
🎨 設計多媒體 免費

📖 技能介紹


name: speech-to-text description: Transcribe or translate audio files to text using a public Hugging Face Whisper Space over Gradio. Use when the user sends voice notes, audio attachments, meeting clips, podcasts, interviews, or any local audio file (.ogg, .mp3, .wav, .m4a, etc.) and wants a transcript, rough captions, or an English translation without relying on paid APIs first.


Speech to Text

Use this skill to turn local audio files into text with a public Whisper-based endpoint.

小蔥技能7w4.net有完整的技能分類。

Quick start

Run:

python3 scripts/transcribe.py /path/to/file.ogg

Return the transcript as plain text. By default, the script also applies lightweight Chinese punctuation and sentence-breaking cleanup.

For machine-readable output:

python3 scripts/transcribe.py /path/to/file.ogg --json

To disable cleanup and keep the raw model text:

python3 scripts/transcribe.py /path/to/file.ogg --format raw

To force Chinese punctuation cleanup:

python3 scripts/transcribe.py /path/to/file.ogg --format zh

For English translation instead of same-language transcription:

python3 scripts/transcribe.py /path/to/file.ogg --task translate

Workflow

  1. Confirm the input is a local audio file.
  2. Run scripts/transcribe.py on it.
  3. If the transcript looks imperfect, tell the user it came from a public Whisper endpoint and may need cleanup.
  4. If helpful, post-process into:
  5. cleaned transcript
  6. summary
  7. action items
  8. bilingual output

What the script does

The script:

  • uploads the local file to a public Gradio-backed Hugging Face Space
  • submits a Whisper transcription job
  • waits for completion via the Gradio event stream
  • prints the resulting text

Default endpoint:

  • https://hf-audio-whisper-large-v3-turbo.hf.space

Override it with:

python3 scripts/transcribe.py input.ogg --space https://your-space.hf.space

or set:

export HF_WHISPER_SPACE=https://your-space.hf.space

Guardrails

  • Treat this as a best-effort public/free path, not a privacy-grade path.
  • Do not use for highly sensitive audio unless the user explicitly accepts public third-party processing.
  • Expect rate limits, queueing, and occasional outages.
  • If the public endpoint fails, explain that the free backend is unavailable and offer alternatives.

Output handling

Prefer to return:

  • the raw transcript when the user asked to "轉文字/聽寫"
  • a cleaned version when punctuation is poor
  • a short note about uncertainty if names, numbers, or jargon may be wrong

Script

  • scripts/transcribe.py — public Whisper transcription helper

🤖 AI 評測

這個 Skill 質量中規中矩,能將音訊轉成文字。它對中文語音識別做了專門最佳化,會自動新增標點符號,使用體驗還算友好。不足之處是依賴第三方公開介面,可能存在隱私風險和網路不穩定的隱患,而且只提供核心指令碼,沒有附帶示例或詳細配置,入手門檻稍高。總體來說,它完成了基本任務,但功能完整度和穩定性還有提升空間。

📊 多維度評分

適應性3.8
規範性4
有效性4.3
可靠性4.5
可信度4.5

📁 包含檔案 (3 個)

📄 SKILL.md 2.6 KB
📄 _meta.json 144 B
📄 scripts/transcribe.py 7.8 KB