Skip to content

Runbooks and SLOs

Operational objectives cover service reliability and privacy workflow integrity. A fast API that silently drops a withdrawal task is unavailable for its purpose.

ServiceIndicatorStandard target
public consent APIvalid requests successfully durably recorded99.95% monthly
Principal portalsuccessful authenticated page/actions99.9% monthly
admin consolesuccessful non-destructive operations99.9% monthly
job deliverydue instructions durably attempted99.9% within 5 min
webhookaccepted events delivered or visibly dead-lettered99.9% within configured window
auditsuccessful domain mutations with audit event100%; fail mutation if audit cannot commit
evidence exportapproved pack completed99% within 4 h

Customer connector/source availability is measured separately from control plane. Planned maintenance and exclusions are explicit and cannot erase statutory workflow impact.

  • P0: cross-tenant exposure, active mass exfiltration, audit integrity failure, total public outage during active statutory clock.
  • P1: consent/withdrawal/rights/incident writes failing, connector fleet compromise, key service failure, data loss.
  • P2: partial connector/provider failure, evidence/export delay, capacity risk.
  • P3: stale non-critical integration, documentation or cosmetic issue.

P0/P1 pages 24×7 for supported tiers. Alerts include runbook, safe first action, tenant impact, correlation and current clocks but no raw PII.

  1. declare incident and commander; preserve first awareness;
  2. confirm alert and affected control/service;
  3. contain safely without deleting evidence;
  4. identify tenants/entities/data/workflows;
  5. start DPDP/CERT-In/sector/contract assessments independently;
  6. communicate factual status;
  7. recover and reconcile queued/ambiguous work;
  8. verify audit, data and tombstones;
  9. close with two-person review and post-incident actions.
Dependency failureSafe behavior
IdPbreak glass for authorised staff; public existing sessions policy; no bypass
queueaccept only if command/outbox commits; workers recover in order
customer agentqueue bounded tasks, show stale, manual fallback
messaging provideralternate approved channel; no sensitive email/SMS body
KMSfail protected read/write; use tested recovery, never plaintext
regulator portalrecord attempts and use approved manual route
legal config servicepin last approved version; show staleness

Published window, customer notice by tier, backup/preflight, migration compatibility, rollback decision, smoke tests including tenant isolation/receipt/audit, and post-change evidence. Emergency changes require incident link and retrospective review.

Capacity alerts at sustained 60/75/85% for database connections/storage, queue age, worker concurrency, object requests and KMS limits. Error-budget burn gates feature releases. Security or audit-integrity risk overrides availability optimisation.

Daily active clocks/dead letters; weekly reliability/support access; monthly SLO, capacity, restore sample and dependency risk; quarterly full restore, region/dependency exercise, incident tabletop and runbook contact validation.