The value of the project was not data collection by itself. Across a 50-100 GPU cluster, usage, temperature, and memory metrics were collected every minute, and a threshold breach reached the operator dashboard and alert within 1-2 minutes - turning anomaly detection that used to take hours of manual checking into a minutes-level workflow for the 5-10 people who operate it.
GPU Monitoring & Alerting
Built a monitoring path from raw telemetry to operator-facing alerts and readable dashboard visibility.
I worked across ingestion, time-series storage, threshold evaluation, alert generation, and dashboard presentation instead of staying in one layer, for a cluster of 50-100 GPUs used by 5-10 research-institution operators.
Cloud Monitoring System
Selected tools and technologies used in the work.
A closer look at the questions behind the work
The summary above is for fast scanning. The Q&A below keeps the deeper story available without making the page feel noisy.
01Why does this project belong in the portfolio?
Monitoring only matters when people can act on it, not when the most data is collected. Across a 50-100 GPU research cluster, I collected usage, temperature, and memory metrics every minute and got threshold breaches onto the operator dashboard and alert within 1-2 minutes, cutting anomaly-detection time from hours of manual checking to minutes for the 5-10 people who operate it. I never measured an exact false-positive rate, but combining thresholds with a minimum duration was the specific decision that kept short-lived spikes from turning into alert fatigue - and that judgment call is what this project is here to show.
02What was my direct scope of ownership in this project?
- Implemented GPU metric collection and ingestion paths.
- Stored time-series metrics in InfluxDB so the system could support recent trends and operational queries.
- Processed threshold rules in Spring Boot and generated alerts from those evaluations.
- Helped shape the React dashboard so users and operators could read system state quickly instead of relying on raw logs.
03What evidence best supports the strength of this work?
Spring Boot + InfluxDB pipeline collects usage, temperature, and memory metrics every minute across a 50-100 GPU cluster.
Threshold breaches reach the operator dashboard and alert within 1-2 minutes, not just passive metric collection.
Combined thresholds with a minimum duration so short-lived spikes do not trigger repeat alerts, cutting anomaly-detection time from hours to minutes for 5-10 operators.
04Which technical decisions mattered most here?
The goal was operational usefulness, not maximum sensitivity. A noisier system would have made alerts easier to ignore.
This matched the need for recent trends, threshold checks, and time-series queries without overcomplicating the first version.
The first release was about helping users and operators understand GPU state quickly, not exposing every possible monitoring control.
05How did the system work end to end in practice?
06What trade-offs did I make, and why?
| Option | Pros | Cons | Decision |
|---|---|---|---|
| Polling-based collection | Simpler to build and operate | Less immediate than streaming | Used polling for the first implementation |
| Alert on every notable event | Lower chance of missing anomalies | High alert fatigue | Used threshold-based alerting to balance signal quality |
07What changed as a result?
The system delivered real-time monitoring and alerting for a 50-100 GPU research cluster, cutting anomaly-detection time for its 5-10 operators from hours of manual checking to minutes.