Flagship System
Oracle Customer Notification Service OCI Migration Case Study
Architected and led a production notification platform migration across 32 global data centers while preserving delivery stability and enabling long-term OCI ownership.
Transition-State to Steady-State Architecture View
Impact Signals
What This Migration Required
Summary: This migration required defining both the migration architecture and the final OCI target architecture, then executing in phases with explicit governance and risk controls.
This case study describes a distributed systems migration to OCI across global data centers, with emphasis on architecture clarity, rollout safety, and operational stability.
Core problem
Summary: The legacy delivery path reduced control during a customer-critical migration window.
Legacy pathways constrained supportability and rollout control during a customer-critical timeline.
My role
I architected the migration strategy end-to-end: first defining the transition architecture for safe staged cutover, then defining the final-state OCI architecture for steady-state operations after migration.
I presented the architecture and rollout approach to the engineering review board, addressed design questions and operational concerns, and drove approval to proceed with a phased implementation plan.
I mentored engineers who were new to OCI so they could onboard quickly, operate confidently in the target environment, and contribute to migration execution without creating delivery risk.
I partnered with product, operations, and leadership stakeholders to identify migration risks, define mitigation and rollback criteria, and maintain alignment during high-scrutiny milestones.
After establishing a stable process and target-state operating model, I managed the migration through critical phases and handed the program off to the broader team with clear ownership, runbooks, and continuation plans.
Migration architecture (transition state)
Summary: The transition architecture used Oracle GoldenGate for bidirectional replication so tenant traffic could be moved incrementally without forcing a high-risk one-time cutover.
Oracle GoldenGate replicated data bidirectionally between legacy data centers and OCI so writes remained synchronized during staged rollout windows. This enabled incremental tenant migrations in controlled waves, with explicit pause points between cohorts to validate health before expanding blast radius.
I added custom Prometheus migration metrics to track replication lag, tenant wave completion, queue depth, worker throughput, and error-rate drift in near real time. Those signals became hard rollout gates for continue, hold, or rollback decisions.
Target architecture (steady state)
Summary: The final-state platform was OCI-native and Kubernetes-based, with service-mesh controls, autoscaling, and Terraform-driven infrastructure orchestration.
The steady-state architecture ran on Kubernetes with Istio service mesh for traffic policy, resilience controls, and service-to-service observability. Data and artifact paths used OCI Object Storage and Block Volume where appropriate, with Autonomous Oracle Database services for managed persistence and operational simplification.
Ingress and edge controls were implemented through API gateway patterns, while dynamic scaling policies handled demand variability for API and worker tiers. Infrastructure provisioning and environment consistency were orchestrated as code with Terraform so platform changes remained reviewable, repeatable, and auditable.
Key engineering decisions
- Controlled traffic shifts instead of all-at-once cutover
- Separate architecture definition for transition-state and final-state OCI operation
- Bidirectional Oracle GoldenGate replication to support safe incremental tenant migration waves
- Formal engineering review board checkpoints before irreversible rollout gates
- Queueing and worker isolation to reduce overload amplification
- Custom Prometheus migration metrics for real-time go/no-go and rollback decisions
- Hands-on OCI onboarding and pairing to accelerate team readiness
- Kubernetes + Istio service mesh as the core target-state runtime and policy layer
- OCI storage and Autonomous Oracle Database adoption for managed, scalable persistence
- API gateway and dynamic autoscaling policies to handle variable production load
- Terraform-managed infrastructure as code for reproducible environment operations
- Risk register with mitigation owners, rollback criteria, and stakeholder communication cadence
- Fallacy-aware architecture constraints for reliability, latency, topology change, and multi-team ownership (see The Fallacies of Distributed Computing Still Break Modern Systems)
- Observability-first rollout with explicit metrics checkpoints
Tradeoffs
Balanced migration speed against production safety, support continuity, and team enablement so the system could be sustained by a wider engineering group after handoff.
Outcome
Summary: The rollout held delivery stability, secured governance approval, reduced migration risk through phased controls, and transitioned ownership cleanly to the long-term team.
The migration improved supportability and delivery stability while supporting a $2M enterprise deal and leaving behind an architecture and operating model that the OCI-enabled team could continue with confidence.