Case Study: Multi-Cloud Resilience, Out-of-Band Observability & 3-2-1 DR
| Engineering Dimension | Production Specification |
|---|---|
| Architecture Pattern | Multi-Cloud Hybrid Support, External Heartbeat Probing & 3-2-1 Disaster Recovery |
| Core Technologies | Oracle Cloud Infrastructure (OCI Mumbai), Google Cloud (GCP), Uptime Kuma, Restic, Terraform S3 |
| Primary Code Paths | infrastructure/oci/, infrastructure/on-prem/backend.tf, configuration/playbooks/deploy_uptime_kuma.yml |
| Relevant Decisions | ADR-011, ADR-015 |
| Operational Status | Production Verified (Out-of-Band Alerts <60s, Remote State Locked, Nightly Encrypted Sync) |
1. Executive Summary
A monitoring system running inside the very cluster it monitors is fundamentally flawed: if the local network, electricity, or physical server suffers a catastrophic failure, internal monitoring goes dark and cannot alert the engineering team. Furthermore, co-locating backup archives on the same physical machine leaves the platform defenseless against hardware death or physical disaster.
This project engineered an enterprise Multi-Cloud Hybrid Support Architecture leveraging Oracle Cloud Infrastructure (OCI) and Google Cloud Platform (GCP). It decouples critical operational dependencies—external availability probing, Terraform remote state locking, and encrypted 3-2-1 backup repositories—into geographically independent cloud regions while maintaining an ultra-lean, cost-optimized operating model.
2. The Problem: The "In-Band Monitoring" & "Co-Located Backup" Trap
- The Silent Outage Trap: If the homelab's residential ISP drops connection or the local circuit breaker trips, internal Prometheus and Grafana instances become unreachable and cannot dispatch webhook alerts. The engineer remains completely blind to outages until an external user reports it.
- State Locking Vulnerability: Storing Terraform state on local disks risks permanent infrastructure blindness if the local drive fails or state becomes corrupted during concurrent executions.
- Disaster Recovery Failure (3-2-1 Violation): Storing VM snapshots strictly on a local hard drive violates the industry 3-2-1 Backup Standard, leaving the platform vulnerable to drive mortality, physical board failure, fire, or theft.
3. Multi-Cloud Hybrid Support Topology
Resilience Architecture:
[ Oracle Cloud Infrastructure (OCI Mumbai ap-mumbai-1) ]
├── Uptime Kuma Compute Instance
│ └── Continuous HTTPS probes to docs, hooks, dash (<60s interval)
├── S3 Object Storage: homelab-terraform-state
│ └── Offsite state locking and historical versioning
└── S3 Object Storage: restic-backups
└── AES-256 encrypted nightly VM and database dumps
▲
│ HTTPS / S3 API (Port 443)
▼
[ On-Premise Sovereign Host (Proxmox VE + k3s-prod) ]
├── Tier 1 (NVMe): Active workloads & CloudNativePG HA
├── Tier 2 (SATA): Local vzdump snapshots (/mnt/hdd/dump/)
└── Tailscale P2P WireGuard Mesh (Management Plane)
▲
│ Multi-Cloud Support Connectivity
▼
[ Google Cloud Platform (GCP us-east1) ]
└── Secondary Support Compute Instance (Operational Automation & Registry)

4. Key Architectural Implementations
1. Independent Out-of-Band Health Probing (Uptime Kuma on OCI)
An independent cloud compute instance in OCI Mumbai (ap-mumbai-1) runs Uptime Kuma in a dedicated container runtime:
- Executes active HTTP status probes against public endpoints (
docs.vijaysingh.cloud,hooks.vijaysingh.cloud,dash.vijaysingh.cloud) over the public internet every 60 seconds. - Pings hypervisor heartbeat endpoints over the private Tailscale WireGuard mesh.
- If response status codes fail or latency exceeds SLA boundaries (TTFB > 2000ms), Uptime Kuma immediately dispatches high-priority incident notifications to Slack (
#homelab-alerts) completely out-of-band.
2. Off-Site Terraform Remote State Backend with S3 State Locking
Terraform state is decoupled from local disks and stored in an OCI Object Storage bucket using standard S3 compatibility API:
- State Locking: Eliminates race conditions and prevents catastrophic concurrent state mutations.
- Object Versioning: Retains an immutable audit trail of every applied state change.
- Disaster Durability: Even if the physical Mini PC is completely destroyed, the exact infrastructure state remains intact in cloud storage.
# Remote State Configuration (infrastructure/on-prem/backend.tf)
terraform {
backend "s3" {
bucket = "homelab-terraform-state"
key = "on-prem/terraform.tfstate"
region = "ap-mumbai-1"
endpoint = "https://<OCI_NAMESPACE>.compat.objectstorage.ap-mumbai-1.oraclecloud.com"
skip_region_validation = true
skip_credentials_validation = true
skip_requesting_account_id = true
skip_s3_checksum = true
skip_metadata_api_check = true
use_path_style = true
}
}
3. True 3-2-1 Disaster Recovery with Encrypted Restic Replication
The platform adheres strictly to the 3-2-1 Backup Rule: 1. 3 Copies of Data: Production data, local snapshot dump, and offsite cloud repository. 2. 2 Different Media Types: Solid-state NVMe flash (hot transactional data) and mechanical SATA magnetic disk (local cold backups). 3. 1 Offsite Geographically Independent Copy: Automated systemd timers execute Restic, encrypting snapshot archives with client-side AES-256-GCM before uploading to OCI Object Storage in Mumbai.
5. Verification & Operational Health
1. Out-of-Band Incident Alerting Simulation
Simulated sudden platform failure by scaling down the edge connector:
kubectl -n platform scale deployment cloudflared --replicas=0
#homelab-alerts:
[FIRING] Service Down: docs.vijaysingh.cloud
Status: HTTP 530 (Edge Ingress Unreachable)
Timestamp: 2026-09-26 13:10:04 IST
kubectl -n platform scale deployment cloudflared --replicas=1
[RESOLVED] Service Restored: docs.vijaysingh.cloud (Ping: 12ms).
2. Restic Offsite Snapshot Integrity Audit
restic -r s3:https://<OCI_NAMESPACE>.compat.objectstorage.ap-mumbai-1.oraclecloud.com/homelab-backups snapshots
ID Time Host Tags Paths
--------------------------------------------------------------------------------
a4b1c2d3 2026-09-25 03:00:01 pve-mumbai vzdump /mnt/hdd/dump/vzdump-500.vma.zst
e5f6a7b8 2026-09-26 03:00:01 pve-mumbai vzdump /mnt/hdd/dump/vzdump-500.vma.zst
--------------------------------------------------------------------------------
2 snapshots, 0 errors
6. Quantified Engineering Impact
| Resilience Dimension | Legacy Single-Host Baseline (v2) | Multi-Cloud Hybrid Architecture (Current) | Engineering Yield |
|---|---|---|---|
| Outage Notification Latency | Silent failure until manually discovered | Instant Webhook Alert (< 60s) | Immediate Incident Visibility |
| Terraform State Durability | Trapped on local operator disk | Offsite S3 Object Storage with Locking | Zero Risk of State Loss |
| Disaster Recovery Posture | Co-located local backups only (0-copy offsite) | Full 3-2-1 Compliance (AES-256 Cloud Sync) | Resilient to Total Bare-Metal Loss |
| Recovery Time Objective (RTO) | 12+ Hours (Manual rebuilding) | < 45 Minutes (Automated Cloud Retrieval) | ~93% Faster Recovery Time |
| Recovery Point Objective (RPO) | 24+ Hours (Unscheduled dumps) | < 15 Minutes (WAL Streaming + Nightly Dumps) | Minimal Data Loss Window |
| FinOps Cloud Footprint | Fragile, costly ingress gateways | Optimized Multi-Cloud Architecture | Enterprise Resilience with Zero Bloat |