OpenStack HA Recovery Automation
Rebuilt recovery behavior in Java and adapted it for HA, priority, security, and closed-network requirements.
This was not a thin API integration. In an environment managing 50+ VMs, I changed recovery behavior itself so a single failed node's 5-10 VMs could be detected, verified, and sequentially restarted in roughly 3-5 minutes, under stricter security, access control, and closed-network requirements.
- The Python-oriented flow did not fit the existing service structure well.
- Government access control and closed-network constraints were not directly supported.
- Flat recovery behavior could not reflect workload importance.
- Rebuilt the recovery path in a Java-centered structure that fit the product better.
- Added account-level access control, closed-network operability, and safer HA handling.
- Introduced priority-based recovery so more important workloads moved first.
This was not a thin wrapper over default OpenStack behavior. I restructured the recovery workflow itself to fit the existing service environment.
The system no longer treated every VM equally and could recover higher-priority workloads first.
I extended recovery behavior to support account-level access control, closed-network delivery, and safer HA recovery at larger scale.
In a 6-month team project, I owned the backend of the HA recovery logic end to end - OpenStack API integration, failed-VM detection, target-node selection, priority handling, and recovery-result logging - and validated it in an environment managing 50+ VMs.
How it works
Analyze upstream OpenStack recovery behavior and identify customer gaps.
Rebuild the recovery workflow in Java around the existing service environment.
Apply account-level access control and customer-specific operational constraints.
Set recovery priorities so critical workloads move first.
What I owned
Integrated OpenStack recovery APIs and built the failed-VM detection logic that kicks off the recovery path.
Designed the criteria for selecting a healthy target node and rebuilt that selection flow in a Java-centered structure, replacing the original Python-oriented one.
Implemented priority-based recovery settings so higher-priority workloads on a failed node recover first among its 5-10 VMs.
Questions you may have first
I started my engineering career in Korea on cloud infrastructure products, then continued in Canada delivering production web experiences and payment flows for business clients.