Weights & Biases Monitor

👤 chrisvoncsefalvay 📦 v1.0.0 ⭐ 4.2 ⬇️ 2.4K 下載
💻 開發程式設計 免費 🔑 需 API Key

📖 技能介紹


name: wandb description: Monitor and analyze Weights & Biases training runs. Use when checking training status, detecting failures, analyzing loss curves, comparing runs, or monitoring experiments. Triggers on "wandb", "training runs", "how's training", "did my run finish", "any failures", "check experiments", "loss curve", "gradient norm", "compare runs".


Weights & Biases

Monitor, analyze, and compare W&B training runs.

Setup

wandb login
# Or set WANDB_API_KEY in environment

Scripts

Characterize a Run (Full Health Analysis)

~/clawd/venv/bin/python3 ~/clawd/skills/wandb/scripts/characterize_run.py ENTITY/PROJECT/RUN_ID

Analyzes: - Loss curve trend (start → current, % change, direction) - Gradient norm health (exploding/vanishing detection)
- Eval metrics (if present) - Stall detection (heartbeat age) - Progress & ETA estimate - Config highlights - Overall health verdict

Options: --json for machine-readable output.

Watch All Running Jobs

~/clawd/venv/bin/python3 ~/clawd/skills/wandb/scripts/watch_runs.py ENTITY [--projects p1,p2]

Quick health summary of all running jobs plus recent failures/completions. Ideal for morning briefings.

Options: - --projects p1,p2 — Specific projects to check - --all-projects — Check all projects - --hours N — Hours to look back for finished runs (default: 24) - --json — Machine-readable output

Compare Two Runs

~/clawd/venv/bin/python3 ~/clawd/skills/wandb/scripts/compare_runs.py ENTITY/PROJECT/RUN_A ENTITY/PROJECT/RUN_B

Side-by-side comparison: - Config differences (highlights important params) - Loss curves at same steps - Gradient norm comparison - Eval metrics - Performance (tokens/sec, steps/hour) - Winner verdict

Python API Quick Reference

import wandb
api = wandb.Api()

# Get runs
runs = api.runs("entity/project", {"state": "running"})

# Run properties
run.state      # running | finished | failed | crashed | canceled
run.name       # display name
run.id         # unique identifier
run.summary    # final/current metrics
run.config     # hyperparameters
run.heartbeat_at # stall detection

# Get history
history = list(run.scan_history(keys=["train/loss", "train/grad_norm"]))

Metric Key Variations

Scripts handle these automatically: - Loss: train/loss, loss, train_loss, training_loss - Gradients: train/grad_norm, grad_norm, gradient_norm - Steps: train/global_step, global_step, step, _step - Eval: eval/loss, eval_loss, eval/accuracy, eval_acc

Health Thresholds

  • Gradients > 10: Exploding (critical)
  • Gradients > 5: Spiky (warning)

    7w4.net有更好的技能外掛。

  • Gradients < 0.0001: Vanishing (warning)
  • Heartbeat > 30min: Stalled (critical)
  • Heartbeat > 10min: Slow (warning)

Integration Notes

For morning briefings, use watch_runs.py --json and parse the output.

For detailed analysis of a specific run, use characterize_run.py.

For A/B testing or hyperparameter comparisons, use compare_runs.py.

🤖 AI 評測

這是一個實用的 W&B 訓練監控工具,功能全面、文件詳細,能有效幫助檢測訓練異常、對比實驗結果。優點是指令碼配套完整,涵蓋了日常監控的主要場景,支援一鍵生成健康報告。不足之處是配置靈活性欠佳,首次使用可能需要調整路徑設定,且某些邊界情況的錯誤提示可以更友好。整體質量良好,適合經常使用 W&B 的機器學習工程師。

📊 多維度評分

適應性4
規範性4
有效性4.3
可靠性4.4
可信度4.2

📁 包含檔案 (7 個)

📄 SKILL.md 3 KB
📄 _meta.json 132 B
📄 scripts/characterize_run.py 12.3 KB
📄 scripts/check_runs.py 2.6 KB
📄 scripts/compare_runs.py 11 KB
📄 scripts/run_details.py 2.7 KB
📄 scripts/watch_runs.py 8.4 KB