Context
A MAS-licensed insurance carrier, SG-domiciled, ran a mature production estate: Microsoft 365 for office, ~140 virtualized servers across two datacenters, two database clusters, file shares, and a Veeam backup environment that the SRE team trusted.
What they didn’t have — and what their internal audit kept flagging — was a credible answer to one question: at 2 AM on a Sunday, when the primary datacenter showed degraded performance, who pulled the trigger on failover? The answer in practice was “wait for the head of infrastructure to wake up, then have a meeting.” MAS examiners weren’t satisfied; neither was the CRO.
A second question was hovering: their annual DR drill had become an exhausting five-day stop-the-world exercise that nobody trusted as representative of real failure. The auditor had asked to see live cutover evidence at the next review. The team was anxious.
TWO TWO proposed a 90-day Jumborca POC — SupInsight as a decision brain over the existing Veeam estate, SupDRC layered for active-active failover between the two datacenters. Zero fee for the POC, terminate any day, no explanation required.
What we did
Weeks 1-2 — Discovery without disruption
We inventoried the production estate — every VM, database, file share, M365 tenant, criticality tier — without touching production. The existing Veeam backup chain stayed as-is; SupAI.ONE attached non-intrusively as a backup-of-backup target. Nothing changed for the SRE team’s day-to-day.
Weeks 3-6 — SupInsight observation phase
SupInsight watched the production telemetry, every backup event, every alert. It made predictions in shadow mode — “at this moment, I would have recommended initiating failover for cluster A; here’s the probability and the cost analysis” — without acting. The CRO and SRE lead reviewed each shadow recommendation weekly. After three weeks, the shadow accuracy was running at 94%; the team began trusting it.
Weeks 7-10 — SupDRC active-active enablement
Configured SupDRC for active-active failover between the two SG datacenters. RTO target ≤ 10 minutes; RPO target ≤ 10 seconds. First synthetic cutover at the end of week 8: 7-minute RTO, 3-second RPO. Repeated weekly through week 10; numbers stayed consistent.
Weeks 11-12 — Live cutover drill
The annual MAS-evidenced DR drill — historically five days of pageantry — was executed as a two-hour live cutover with SupDRC
- SupInsight running. Auditor witnessed. Result: 8-minute RTO, 4-second RPO, automated audit report generated in real time and signed off at the end of the drill.
The 90-day POC converted to a multi-year SupAI.ONE subscription plus annual SupDrill — the drill-automation product — to keep the discipline current.
Results
- 8 minutes RTO at live cutover (target ≤ 10)
- 4 seconds RPO at live cutover (target ≤ 10)
- 5 days → 2 hours annual DR drill duration
- Automated audit report generated and accepted by MAS examiner at first review post-POC
- 94% SupInsight recommendation agreement rate in shadow phase (rising to 97% in production)
- Veeam stays — SupAI.ONE sits on top, doesn’t replace; team’s trusted backup discipline preserved
What we learned
The 2 AM decision problem is not a technology problem; it is a decision-rights problem. SupInsight worked because it stopped being a black-box recommender and became a documented, replayable judgment — the SRE team could explain to the CRO why a recommendation was issued, and the CRO could explain it to the board.
The Jumborca POC’s “terminate any day, no fee, no explanation” posture mattered more than the technology in the early weeks. The team felt safe enough to engage seriously; they were not selling their freedom upfront.
The SupVault long-term retention product was added after engagement month 9 — the customer’s compliance team wanted a WORM-protected audit-evidence archive separate from Veeam. The SupDrill product keeps the drill discipline alive without re-burdening the team each year.