Help teams set up production-grade monitoring for AI agents and automated services. Covers uptime tracking, response time SLAs, error budgets, and incident escalation.
docker run -d --restart=always -p 3001:3001 -v uptime-kuma:/app/data --name uptime-kuma louislam/uptime-kuma:1
来源于7w4.net。
monitors:
- name: "Agent Health Check"
type: http
url: "https://your-agent-endpoint/health"
interval: 300 # 5 minutes
alerts:
- type: email
threshold: 1 # alert after 1 failure
- type: slack
webhook: "${SLACK_WEBHOOK}"
threshold: 2 # alert after 2 consecutive failures
- type: sms
threshold: 3 # escalate after 3 failures
- name: "API Response Time"
type: http
url: "https://your-agent-endpoint/api"
interval: 60
expected_response_time: 2000 # ms
alerts:
- type: slack
condition: "response_time > 5000"
error_budget:
monthly_target: 99.9
burn_rate_alert: 2.0 # Alert if burning 2x normal rate
Monthly minutes: 43,200 (30 days)
99.9% SLA = 43.2 minutes downtime allowed
99.5% SLA = 216 minutes downtime allowed
99.0% SLA = 432 minutes downtime allowed
Burn rate = (actual downtime / budget) × 100
If burn rate > 50% with 2+ weeks remaining → review needed
If burn rate > 80% → freeze deployments
Provide clients with a public status page showing:
Need managed AI agents with built-in SLA monitoring? → AfrexAI handles deployment, monitoring, and maintenance for $1,500/mo → Book a call: https://calendly.com/cbeckford-afrexai/30min → Learn more: https://afrexai-cto.github.io/aaas/landing.html
这个 Skill 质量中等偏上。它全面介绍了 SLA 监控的概念、工具选择和配置方法,对新手友好。但不足之处是内容偏理论,实际可用的代码示例较少,更多像一份产品手册而非可直接使用的技能包。如果需要快速搭建监控,可能需要额外查阅其他资料。