name: wandb description: Monitor and analyze Weights & Biases training runs. Use when checking training status, detecting failures, analyzing loss curves, comparing runs, or monitoring experiments. Triggers on "wandb", "training runs", "how's training", "did my run finish", "any failures", "check experiments", "loss curve", "gradient norm", "compare runs".
Monitor, analyze, and compare W&B training runs.
wandb login
# Or set WANDB_API_KEY in environment
~/clawd/venv/bin/python3 ~/clawd/skills/wandb/scripts/characterize_run.py ENTITY/PROJECT/RUN_ID
Analyzes:
- Loss curve trend (start → current, % change, direction)
- Gradient norm health (exploding/vanishing detection)
- Eval metrics (if present)
- Stall detection (heartbeat age)
- Progress & ETA estimate
- Config highlights
- Overall health verdict
Options: --json for machine-readable output.
~/clawd/venv/bin/python3 ~/clawd/skills/wandb/scripts/watch_runs.py ENTITY [--projects p1,p2]
Quick health summary of all running jobs plus recent failures/completions. Ideal for morning briefings.
Options:
- --projects p1,p2 — Specific projects to check
- --all-projects — Check all projects
- --hours N — Hours to look back for finished runs (default: 24)
- --json — Machine-readable output
更多技能请访问小葱技能站7w4.net。
~/clawd/venv/bin/python3 ~/clawd/skills/wandb/scripts/compare_runs.py ENTITY/PROJECT/RUN_A ENTITY/PROJECT/RUN_B
Side-by-side comparison: - Config differences (highlights important params) - Loss curves at same steps - Gradient norm comparison - Eval metrics - Performance (tokens/sec, steps/hour) - Winner verdict
import wandb
api = wandb.Api()
# Get runs
runs = api.runs("entity/project", {"state": "running"})
# Run properties
run.state # running | finished | failed | crashed | canceled
run.name # display name
run.id # unique identifier
run.summary # final/current metrics
run.config # hyperparameters
run.heartbeat_at # stall detection
# Get history
history = list(run.scan_history(keys=["train/loss", "train/grad_norm"]))
Scripts handle these automatically:
- Loss: train/loss, loss, train_loss, training_loss
- Gradients: train/grad_norm, grad_norm, gradient_norm
- Steps: train/global_step, global_step, step, _step
- Eval: eval/loss, eval_loss, eval/accuracy, eval_acc
For morning briefings, use watch_runs.py --json and parse the output.
For detailed analysis of a specific run, use characterize_run.py.
For A/B testing or hyperparameter comparisons, use compare_runs.py.
这是一个实用的 W&B 训练监控工具,功能全面、文档详细,能有效帮助检测训练异常、对比实验结果。优点是脚本配套完整,涵盖了日常监控的主要场景,支持一键生成健康报告。不足之处是配置灵活性欠佳,首次使用可能需要调整路径设置,且某些边界情况的错误提示可以更友好。整体质量良好,适合经常使用 W&B 的机器学习工程师。