- The Python-oriented flow did not fit the existing service structure well.
- Government access control and closed-network constraints were not directly supported.
- Flat recovery behavior could not reflect workload importance.
OpenStack HA Recovery Automation
Rebuilt recovery behavior in Java and adapted it for HA, priority, security, and closed-network requirements.
In a 6-month team project, I owned the backend of the HA recovery logic end to end - OpenStack API integration, failed-VM detection, target-node selection, priority handling, and recovery-result logging - and validated it in an environment managing 50+ VMs.
- The recovery path was rebuilt in a Java-centered structure that fit the product better.
- Account-level access control, closed-network operability, and safer HA handling were added.
- Priority-based recovery made it possible to move more important workloads first.
This was not a thin wrapper over default OpenStack behavior. The recovery workflow itself was restructured to fit the service environment.
The system no longer treated every VM equally and could recover higher-priority workloads first.
The default recovery behavior was adapted for government security expectations and safer HA recovery at larger scale.
A 15-second HA flow from failure detection to successful VM recovery
This keeps the key labels and flow of the original architecture while simplifying the motion so a third party can understand it more quickly.
This was not a thin API integration. In an environment managing 50+ VMs, I changed recovery behavior itself so a single failed node's 5-10 VMs could be detected, verified, and sequentially restarted in roughly 3-5 minutes, under stricter security, access control, and closed-network requirements.
Private Cloud Recovery Platform
Selected tools and technologies used in the work.
A closer look at the questions behind the work
The summary above is for fast scanning. The Q&A below keeps the deeper story available without making the page feel noisy.
01Why is this the project to lead with?
Because of the three projects, this one required the most judgment. GPU monitoring was mostly about wiring a pipeline together well, and the payment work was about shipping fast and safely alone - but here I had to design answers to questions with no textbook solution on top of open-source code: which VMs on a failed node should recover first, and by what criteria a healthy node gets picked as the target. In an environment managing 50+ VMs, the result was a sequential recovery path - detect, verify, select a target node, restart, confirm state - that completed in roughly 3-5 minutes for the 5-10 VMs on a single failed node. It is the project where I can walk through that reasoning in the most detail, which is why it leads.
02What was my direct scope of ownership in this project?
- Integrated OpenStack recovery APIs and built the failed-VM detection logic that kicks off the recovery path.
- Designed the criteria for selecting a healthy target node and rebuilt that selection flow in a Java-centered structure, replacing the original Python-oriented one.
- Implemented priority-based recovery settings so higher-priority workloads on a failed node recover first among its 5-10 VMs.
- Logged recovery results per VM and tested the extended workflow in an environment managing 50+ VMs before customer delivery.
03What evidence best supports the strength of this work?
In an environment managing 50+ VMs, a single failed node's 5-10 VMs are detected, verified, and sequentially recovered in roughly 3-5 minutes.
Added priority handling so higher-importance workloads on a failed node recover before the rest.
Extended recovery behavior for closed-network operability and stronger, account-level access control required by government and enterprise customers.
04Which technical decisions mattered most here?
The customer requirements made it clear that simply exposing OpenStack APIs would not be enough. The recovery workflow had to be extended to behave like an internal feature set tailored to the environment.
In a larger environment, not every VM should recover with the same urgency. Priority settings let the system reflect operational importance instead of treating every failure the same way.
I used a Raft-based approach for more stable coordination and state handling so recovery logic could behave more safely across multiple nodes and a broader VM footprint.
05How did the system work end to end in practice?
06What trade-offs did I make, and why?
| Option | Pros | Cons | Decision |
|---|---|---|---|
| Use upstream behavior as-is | Fastest initial integration | Could not satisfy customer-specific security and network requirements | Extended and adapted the workflow for the real environment |
| Simple sequential recovery | Easy to implement | Ignored workload importance and scale pressure | Introduced priority-based recovery behavior |
| Focus only on basic API success | Less engineering effort up front | Would not hold up in closed-network and HA-heavy customer conditions | Designed for the actual delivery environment instead |
07What changed as a result?
The result was a Java-based HA recovery workflow, validated in an environment managing 50+ VMs, where a single failed node's 5-10 VMs are detected and sequentially recovered in roughly 3-5 minutes - going beyond generic OpenStack API usage to match the security, network, and priority requirements of government and enterprise customers.