Operations Guide
Day-2 operations for AI agent platforms. This guide covers health monitoring, incident response, scaling, backup, secret rotation, agent lifecycle management, cost management, compliance, disaster recovery, and runbook automation — aligned to WAF Operational Excellence (OE:07, OE:10) and NIST AI RMF (MANAGE 2.2).
Table of Contents
- Day-2 Operations Overview
- Health Monitoring
- Incident Response
- Scaling Procedures
- Backup & Restore
- Secret & Certificate Rotation
- Agent Lifecycle Management
- Cost Management
- Compliance & Audit
- Disaster Recovery
- Troubleshooting Playbooks
- Runbook Automation
1. Platform Components
| Component | Azure Service | Health Signal | SLA |
|---|---|---|---|
| AI Models | Azure Foundry | Model deployment status | 99.9% |
| Agent Memory | Cosmos DB | Availability metric | 99.99% |
| Decision Ledger | Cosmos DB | Availability metric | 99.99% |
| Search Index | AI Search | Service status | 99.9% |
| Secrets | Key Vault | Service health | 99.99% |
| Telemetry | App Insights | Ingestion latency | 99.9% |
| AI Gateway | API Management | Gateway health | 99.95% |
| Networking | VNet / PEs | Resource health | 99.99% |
Operational Responsibility Model
Platform Team Agent Developers Security Team
├── Infrastructure ├── Agent manifests ├── Policy packs
├── Networking ├── Prompt engineering ├── Cert pipeline
├── Monitoring ├── Toolbox configs ├── Audit reviews
├── Scaling ├── Eval datasets ├── Incident triage
├── DR procedures ├── SDK integration ├── Compliance
└── Cost governance └── Agent certification └── Trust scores2. Health Monitoring
Critical Metrics
| Metric | Source | Threshold | Alert Severity |
|---|---|---|---|
| Agent response latency (P95) | App Insights | > 30s | Sev 2 |
| Agent error rate | App Insights | > 5% | Sev 1 |
| Trust score average | Cosmos DB | < 0.7 | Sev 2 |
| Token consumption rate | APIM | > 80% budget | Sev 3 |
| Cosmos DB RU consumption | Cosmos DB | > 80% provisioned | Sev 2 |
| Key Vault throttling | Key Vault | > 0 events/5min | Sev 3 |
| Decision ledger write failures | App Insights | > 0 | Sev 1 |
| Policy violation count | App Insights | > 0 (block mode) | Sev 1 |
Alert Escalation
Sev 3 (Warning) → Slack/Teams notification → 4h response
Sev 2 (Error) → PagerDuty/on-call page → 1h response
Sev 1 (Critical) → PagerDuty P1 + bridge → 15min response
Sev 0 (Outage) → Incident Commander → Immediate3. Incident Response
Severity Classification
| Severity | Impact | Example | Response Time |
|---|---|---|---|
| Sev 0 | Full platform outage | Foundry endpoint unreachable | Immediate |
| Sev 1 | Agent degradation | All agents returning errors | 15 min |
| Sev 2 | Partial degradation | Single agent class failing | 1 hour |
| Sev 3 | Non-critical issue | Elevated latency, warnings | 4 hours |
| Sev 4 | Informational | Minor config drift detected | Next business day |
Response Procedure
1. DETECT → Automated alert or user report
2. TRIAGE → Classify severity, assign owner
3. CONTAIN → Isolate affected agent/service
4. DIAGNOSE → Use troubleshooting playbooks (Section 11)
5. RESOLVE → Apply fix, verify in staging first
6. VERIFY → Run eval pipeline against affected agent
7. POSTMORTEM → Document root cause, update runbooksEscalation Matrix
| Condition | Escalation Target | Action |
|---|---|---|
| Agent trust score < 0.5 | Security Team | Disable agent, review ledger |
| Policy violation (block) | Compliance Officer | Review, approve/reject |
| Model endpoint down | Platform Team | Failover to secondary model |
| Data boundary violation | Security + Legal | Immediate containment |
| Cost spike > 200% | FinOps + Platform | Budget alert, throttle |
4. Scaling Procedures
Cosmos DB (Serverless): Auto-scales to 5000 RU/s. No action needed below that threshold. If exceeding limits, convert to provisioned throughput (requires data export/import).
AI Search: Scale replicas for throughput, partitions for index size. Use az search service update to adjust.
API Management: Scale units for gateway throughput. Premium tier supports multi-region deployment for HA.
5. Proactive Health Checks
Run weekly:
# Validate all configurations
python scripts/validate_all.py
# Check Key Vault expiring secrets (< 30 days)
az keyvault secret list --vault-name kv-foundry-prod \
--query "[?attributes.expires < '<30d-from-now>']" \
-o table
# Verify Azure Policy compliance
az policy state summarize \
--resource-group rg-foundry-prod \
--query "policyAssignments[?results.nonCompliantResources > `0`]" \
-o tableRelated Resources
- Dashboards — Azure Monitor Workbook setup and KPIs
- Telemetry Schema — OTel spans and custom attributes
- Troubleshooting — Common issues and fixes
- RBAC Matrix — Role assignments for operations personas
- Deployment Guide — Infrastructure provisioning