Skip to content

[Enterprise] Add residency/deployment profiles, backups, disaster recovery, and SLO evidence #122

Description

@rahuliitk

Parent epic: #83

Priority: P2

Problem and current baseline

QuickVoice is deployable with a server/control plane, database and object data, the LiveKit voice runtime, and external Twilio/Telnyx/STT/LLM/TTS/tool providers. Existing security/privacy primitives include organization scoping, secrets, audit, redaction, recording controls, zero-PII behavior, and retention. Enterprise operators still need a truthful deployment/residency inventory, encrypted and tested backups, documented recovery objectives, capacity/SLO evidence, and explicit visibility into every third-party path that may cross a requested boundary. Configuration alone cannot prove recoverability or residency.

User-visible outcome

An administrator selects or records a supported deployment profile, sees where each data/media/control dependency is expected to operate, configures protected backups, runs restore and disaster-recovery drills, tracks RTO/RPO and service indicators, and exports evidence of configuration and exercises—without QuickVoice claiming guarantees that depend on customer or provider infrastructure.

Functional requirements

Deployment and residency profiles

  • Add versioned deployment profiles for supported hosted, customer-cloud/VPC, private, and self-hosted topologies as applicable, with control-plane, media, database, object storage, cache/queue, analytics/eval, log/audit, backup, and external-provider locations/capabilities.
  • Add data-flow/dependency inventory mapping data classes (audio, recording, transcript, metadata, contacts, secrets, analytics, evaluations, tool payloads, backups) to storage/processing/transit components, region, encryption/key profile, retention, and responsible operator.
  • Validate an organization/project/environment's requested residency/deployment policy against LiveKit, Twilio/Telnyx/SIP carrier, STT/LLM/TTS, integration/MCP/tool, storage, telemetry, support, and backup capabilities before publication.
  • Surface supported, unsupported, unknown, and customer-attested separately. Do not silently treat a configured region label as proof that network transit, support access, or a third party remains in-region.
  • Track configuration/version drift and dependency inventory changes, require review for material boundary changes, and retain evidence of the profile active for each agent/call/export.

Backup and restore

  • Define backup coverage for relational data, object/audio artifacts allowed by policy, agent/workflow versions, configuration, audit/evidence, encryption metadata, and external-secret references. Document excluded/ephemeral data and reconstruction dependencies.
  • Support encrypted scheduled backups, retention/expiry, immutability/object-lock where available, per-tenant/environment scope, completion verification, inventory/manifests, checksums, and alerting.
  • Provide supported full and scoped restore workflows into an isolated validation target before promotion. Prevent accidental overwrite/cross-tenant restore and preserve identity, key, retention, and legal-hold rules.
  • Track restore points, dependency/software/schema versions, key availability, validation results, data loss window, duration, actor/approval, and post-restore reconciliation.

Disaster recovery, capacity, and SLO evidence

  • Define versioned recovery tiers with documented RTO/RPO targets, eligible topology, failover/failback sequence, dependency assumptions, communication/approval, and explicit non-covered provider failures.
  • Add regional/service drain and recovery hooks that respect residency, provider compatibility, call state, idempotency, and data consistency; never imply an incompatible in-flight PSTN/media call can move without interruption.
  • Add automated and scheduled recovery drills using isolated/synthetic data where possible, plus controlled documented production exercises. Capture each step, observed RTO/RPO, data validation, gaps, owner, and remediation.
  • Define SLIs for API/realtime availability, call setup/success, media/runtime health, STT/LLM/TTS latency, tool reliability, queues/jobs, webhook/SIEM delivery, backup freshness/success, restore readiness, and data pipeline lag.
  • Provide SLO definitions, measurement window, exclusions/burn rate, error budget, alert status, dependency attribution, and evidence export. Clearly distinguish internal target, observed result, and contractual SLA.
  • Add capacity profiles and tests for concurrent calls, campaign pacing, worker/provider limits, API/realtime connections, queues, database/storage, backup/restore throughput, and regional failover headroom.

Data, API, and events

  • Add/normalize DeploymentProfile/version, ComponentInventory, DataFlow, ResidencyDecision, BackupPolicy, BackupRun/manifest, RestoreRun, RecoveryPlan, RecoveryExercise, SliDefinition, Slo/measurement/burn event, and CapacityProfile/test.
  • Expose scoped APIs to validate/publish/compare profiles, inspect inventory/drift, configure/run/verify backup/restore, start/record recovery exercises, query SLO/burn/capacity evidence, and export a redacted evidence bundle.
  • Emit correlated events for profile validation/publication/drift, backup start/success/failure/expiry, restore start/validation/promotion/failure, recovery exercise/failover/failback, SLO burn/violation/recovery, capacity threshold, and dependency capability change.
  • Evidence must reference immutable configuration and test versions and exclude secret values, raw PII/audio/transcripts, and unrelated tenant data.

UX and accessibility

  • Add an infrastructure/residency map with text/table equivalent, unsupported/unknown dependency warnings, profile diff/drift, backup freshness, restore-point explorer, recovery runbook/exercise timeline, SLO/error-budget dashboard, and capacity forecast.
  • Require production scope, data-loss/interruption risk, region, approval, and rollback to be explicit before restore/failover actions. Do not rely on color alone; support keyboard navigation and screen readers for maps, tables, timelines, and charts.
  • Evidence exports show generation time, source coverage, unknowns, observed versus target values, and provider/customer attestations.

Security, privacy, compliance, provider, and cost boundaries

  • Reuse existing RBAC, secrets, audit, redaction, zero-PII, retention, recording, and legal-hold controls. Encrypt backups in transit/at rest, restrict restore/evidence access, require strong reauthentication/approval where configured, and audit every operation.
  • Backup/restore must preserve tenant isolation and deletion/retention/hold precedence. External key loss can make data unrecoverable; surface key dependencies and validate them in drills without storing plaintext keys.
  • Third-party carrier/media/model/tool routing, network transit, support access, and customer deployment can limit residency and recovery. Report verified facts and unknowns; do not claim compliance, certification, absolute residency, uptime, or disaster immunity.
  • Storage replication, immutable backups, cross-region standby, provider redundancy, exercises, and capacity headroom have cost. Estimate and meter them, require explicit configuration, and avoid surprise replication across forbidden regions.

Failure modes and backward compatibility

  • Existing deployments receive an unclassified/current profile populated from discoverable configuration; no workload is moved and no region or SLO is inferred as guaranteed.
  • Handle missed/corrupt backup, incomplete manifest, schema/software mismatch, unavailable KMS/secret/provider, cross-region policy conflict, partial restore, stale replica, split brain, failed failover/failback, call interruption, SLO data gap, and insufficient standby capacity.
  • A backup is not verified until integrity checks pass, and a recovery objective is not exercised until restoration and validation complete. Failed/partial operations remain visible with last known-good state and documented manual recovery.

Dependencies and out of scope

  • Depends on current deployment tooling, database/object storage, LiveKit runtime, Twilio/Telnyx/SIP and model/tool provider inventories, projects/environments/RBAC/approvals, encryption/secrets/BYOK, canonical events/analytics, retention/holds, and operational alerting.
  • Coordinates with runtime/telephony failover, SIEM/evidence, billing budgets, data export, incident operations, and provider capability metadata.
  • Out of scope: promising a contractual SLA in product UI, guaranteeing third-party or customer infrastructure, automatically certifying residency/compliance, building every cloud's disaster-recovery service, or moving all active calls without disruption.

Acceptance criteria

  • Every environment has a versioned profile listing all known control/media/data/backup/external-provider paths as supported, unsupported, unknown, or customer-attested with source and timestamp.
  • Profile validation blocks or requires the configured exception/approval when a dependency conflicts with requested region/encryption/retention policy, and material drift is detected and audited.
  • Scheduled backups produce encrypted manifests/checksums and freshness alerts; no run is marked successful or verified when required coverage/integrity is missing.
  • A restore can run in an isolated target, validates tenant/schema/version/key/retention/hold invariants, records observed data-loss window and duration, and cannot overwrite production without explicit authorization/approval.
  • Each recovery tier has documented assumptions and a drill records observed RTO/RPO, consistency checks, failed steps, owner, and remediation; untested targets are labeled unexercised.
  • SLO dashboards distinguish target, observed SLI, exclusions/data gaps, dependency attribution, error budget, and contractual SLA status.
  • Capacity tests cover concurrent calls and critical control/data paths plus declared failover headroom and provider limits.
  • Evidence exports are immutable-reference-based, redacted, scoped, and explicit about unknowns; documentation makes no compliance/certification/residency/uptime guarantee.

Test plan

  • Profile/data-flow validation tests for hosted/self-hosted variants, provider capability changes, unknown/customer-attested paths, drift, cross-region conflicts, and immutable call/config references.
  • Backup/restore tests for full/scoped data, encryption/key rotation, checksums, legal hold/retention, deleted resources, schema/version upgrade/downgrade, cross-tenant prevention, corrupt/missing artifacts, and idempotent retry.
  • Disaster-recovery exercises/failure injection for region loss, database/object store/queue/LiveKit/provider outage, KMS/secret loss, stale replica, split brain, failover/failback, active calls, and reconciliation.
  • SLI/SLO tests for accurate windows/exclusions/burn, data gaps, delayed events, dependency attribution, alerts, and evidence export reproducibility.
  • Authorization/reauthentication/approval, audit/redaction/zero-PII, accessibility, migration/rollback, storage/cost metering, and concurrent-call/campaign/API/worker/capacity load tests.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area: ai-runtimePython AI API and LiveKit workerarea: consoleCustomer consolearea: securitySecurity-sensitive project workarea: serverExpress API and server control planearea: toolingLocal development and repository toolingenhancementNew feature or requeststatus: needs-designNeeds maintainer design agreement before implementation

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions