Runbooks and SLOs
Operational objectives cover service reliability and privacy workflow integrity. A fast API that silently drops a withdrawal task is unavailable for its purpose.
Service indicators
Section titled “Service indicators”| Service | Indicator | Standard target |
|---|---|---|
| public consent API | valid requests successfully durably recorded | 99.95% monthly |
| Principal portal | successful authenticated page/actions | 99.9% monthly |
| admin console | successful non-destructive operations | 99.9% monthly |
| job delivery | due instructions durably attempted | 99.9% within 5 min |
| webhook | accepted events delivered or visibly dead-lettered | 99.9% within configured window |
| audit | successful domain mutations with audit event | 100%; fail mutation if audit cannot commit |
| evidence export | approved pack completed | 99% within 4 h |
Customer connector/source availability is measured separately from control plane. Planned maintenance and exclusions are explicit and cannot erase statutory workflow impact.
Alert priorities
Section titled “Alert priorities”- P0: cross-tenant exposure, active mass exfiltration, audit integrity failure, total public outage during active statutory clock.
- P1: consent/withdrawal/rights/incident writes failing, connector fleet compromise, key service failure, data loss.
- P2: partial connector/provider failure, evidence/export delay, capacity risk.
- P3: stale non-critical integration, documentation or cosmetic issue.
P0/P1 pages 24×7 for supported tiers. Alerts include runbook, safe first action, tenant impact, correlation and current clocks but no raw PII.
Common runbook shape
Section titled “Common runbook shape”- declare incident and commander; preserve first awareness;
- confirm alert and affected control/service;
- contain safely without deleting evidence;
- identify tenants/entities/data/workflows;
- start DPDP/CERT-In/sector/contract assessments independently;
- communicate factual status;
- recover and reconcile queued/ambiguous work;
- verify audit, data and tombstones;
- close with two-person review and post-incident actions.
Degraded modes
Section titled “Degraded modes”| Dependency failure | Safe behavior |
|---|---|
| IdP | break glass for authorised staff; public existing sessions policy; no bypass |
| queue | accept only if command/outbox commits; workers recover in order |
| customer agent | queue bounded tasks, show stale, manual fallback |
| messaging provider | alternate approved channel; no sensitive email/SMS body |
| KMS | fail protected read/write; use tested recovery, never plaintext |
| regulator portal | record attempts and use approved manual route |
| legal config service | pin last approved version; show staleness |
Maintenance
Section titled “Maintenance”Published window, customer notice by tier, backup/preflight, migration compatibility, rollback decision, smoke tests including tenant isolation/receipt/audit, and post-change evidence. Emergency changes require incident link and retrospective review.
Capacity and error budgets
Section titled “Capacity and error budgets”Capacity alerts at sustained 60/75/85% for database connections/storage, queue age, worker concurrency, object requests and KMS limits. Error-budget burn gates feature releases. Security or audit-integrity risk overrides availability optimisation.
Operational reviews
Section titled “Operational reviews”Daily active clocks/dead letters; weekly reliability/support access; monthly SLO, capacity, restore sample and dependency risk; quarterly full restore, region/dependency exercise, incident tabletop and runbook contact validation.