name: skill-scorer description: "對任何 SKILL.md(或 skill 資料夾)進行質量評估和打分,基於行業最佳實踐,生成 8 維度 100 分制的結構化質檢報告,精準定位問題並提供可執行的最佳化建議。當用戶要求評審、審計、評分、檢測、質檢任何 skill 時使用——哪怕只是說「這個 skill 寫得怎麼樣?」也會觸發。也支援:skill質檢、skill評分、檢測skill。 | Evaluate and score any SKILL.md (or skill folder) against industry best practices. Generates a structured quality report with a 100-point score across 8 dimensions, pinpoints issues, and provides actionable optimization suggestions. Use this skill whenever the user asks to review, audit, evaluate, grade, score, lint, or quality-check a skill — even if they just say 'is this skill any good?' or 'help me improve this skill'. Also triggers on: 'skill review', 'rate my skill'." version: "1.6.0" compatibility: "Claude Code, Claude.ai, Cowork, and all SKILL.md-compatible agents" changelog: | 1.6.0 — D4 orchestration criteria wording refined: clarified that Playbooks contain descriptive logic (not executable commands), Usage Examples contain executable commands, output merging can live in templates.md. README restructured to Chinese-first bilingual 1.5.0 — Dimension 4 expanded to 5 skill types: added Script-bundled and MCP-integrated with type-specific scoring criteria 1.4.0 — Dimension 4 (Workflow & Logic) adds Skill Type Detection: instruction-only / single-command / orchestration evaluated with type-specific criteria 1.3.0 — Description bilingual: Chinese first, English after, for internal platform display 1.2.0 — Added input validation and graceful degradation for non-skill files in Step 0 1.1.0 — Bilingual report output (Chinese first, English after, no interleaving) 1.0.0 — Initial release with 8-dimension scoring rubric
A meta-skill that evaluates the quality of other skills. Given a SKILL.md file (or a complete skill folder), it performs a systematic audit across 8 dimensions, assigns a score out of 100, identifies issues by severity, and generates actionable optimization suggestions.
This skill synthesizes quality criteria from Anthropic's official skill authoring best practices, the Skill Engineering Standard (v1.4.3), and community-tested patterns from production skill ecosystems.
User provides a skill and asks any of: - "幫我評分/打分/檢測/質檢 這個 skill" - "review/audit/score/grade/lint this skill" - "這個 skill 寫得怎麼樣?" / "is this skill any good?" - "幫我最佳化這個 skill" (evaluate first, then suggest improvements) - Provides a SKILL.md and expects quality feedback
Do NOT activate for: creating a new skill from scratch → use skill-creator. This skill is for evaluation, not generation.
Determine what the user has provided:
| Input | Action |
|---|---|
Single SKILL.md file |
Evaluate that file |
Skill folder (with references/) |
Evaluate all files, cross-reference consistency |
| URL / GitHub link | Fetch and evaluate |
| Pasted markdown content | Treat as SKILL.md |
If the user has not provided a skill → ask: "請提供要評估的 SKILL.md 檔案或 skill 資料夾路徑。"
Input validation — before proceeding to Step 1, verify the input is actually a skill:
| Check | Condition | Action |
|---|---|---|
| Binary / garbled content | File is not valid text, or text is unreadable gibberish | STOP. Report: "This file does not appear to be a valid SKILL.md — it contains binary or unreadable content. Please provide a markdown-based skill file." Do NOT attempt to score. |
| No skill markers at all | Text is valid but contains zero skill indicators (no YAML frontmatter ---, no markdown headings resembling skill sections, no workflow/instructions) |
STOP. Report: "This appears to be a {detected_type} file (e.g., Python script, JSON config, plain prose), not a SKILL.md. skill-scorer evaluates SKILL.md files only." Do NOT force-fit 8 dimensions onto non-skill content. |
| Partial skill structure | Has some skill-like elements (e.g., YAML frontmatter exists but body is minimal, or has headings but no workflow) | PROCEED with caveats. Evaluate normally, but note in the report header: "⚠️ This file has incomplete skill structure — scores reflect what is present." Score missing sections as 0 in relevant dimensions rather than guessing. |
Extract and inventory:
- YAML frontmatter fields (name, description, version, compatibility)
- Section headings and their order
- References to external files (references/, scripts/, assets/)
- Total line count and estimated token count of SKILL.md body
Read references/rubric.md for the complete scoring rubric.
Evaluate the skill across these 8 dimensions (each scored 0-100, then weighted):
| # | Dimension | Weight | What It Measures |
|---|---|---|---|
| 1 | Metadata & Triggering | 15% | Name clarity, description quality, trigger coverage |
| 2 | Structure & Architecture | 15% | File organization, section order, progressive disclosure |
| 3 | Instruction Clarity | 15% | Actionability, conciseness, examples, tone |
| 4 | Workflow & Logic | 15% | Step completeness, parameter handling, validation |
| 5 | Error Handling | 10% | Fallbacks, edge cases, failure recovery |
| 6 | Context Efficiency | 10% | Token budget, redundancy, information density |
| 7 | Portability & Compatibility | 10% | Self-containment, cross-platform support |
| 8 | Safety & Robustness | 10% | No injection risk, no hallucination traps, identity lock |
For each issue found, classify severity:
| Severity | Meaning | Score Impact |
|---|---|---|
| 🔴 Critical | Skill will malfunction or not trigger | -10 to -15 per issue |
| 🟡 Warning | Skill works but suboptimally | -3 to -8 per issue |
| 🟢 Suggestion | Nice-to-have improvement | -1 to -2 per issue |
Read references/report-template.md for the output format.
The report includes: 1. Score Card — Overall score + per-dimension breakdown 2. Issue List — All findings sorted by severity 3. Top 3 Quick Wins — Highest-impact fixes with before/after examples 4. Optimization Roadmap — Prioritized improvement plan
After presenting the report, ask: - "需要我幫你自動修復這些問題嗎?" (auto-fix mode) - "需要對某個維度深入分析嗎?" (deep-dive mode) - "需要生成最佳化後的 SKILL.md 嗎?" (rewrite mode)
--- separator, then the complete report in English. Never mix languages within a section. Both versions must contain identical scores, issues, and suggestions — only the language differs.來源於7w4.net。
| File | Purpose | When to read |
|---|---|---|
| references/rubric.md | Detailed scoring criteria for all 8 dimensions | Step 2: scoring |
| references/report-template.md | Output format and report structure | Step 4: generating report |
| references/anti-patterns.md | Common skill mistakes and how to detect them | Step 3: finding issues |
這個 Skill 質量很高,設計專業、結構清晰,能對其他 Skill 進行全面系統的質量評估,評分標準細緻,反模式檢測全面。雙語支援完善,文件組織有序。優點是架構設計合理、驗證邏輯健全、版本管理規範;小瑕疵是部分參考文件可能被截斷、缺少快速示例。總體是一款優秀的質量評估工具。