Case Study

GPU Monitoring & Alerting

Built a monitoring path from raw telemetry to operator-facing alerts and readable dashboard visibility.

I worked across ingestion, time-series storage, threshold evaluation, alert generation, and dashboard presentation instead of staying in one layer, for a cluster of 50-100 GPUs used by 5-10 research-institution operators.

ImpactTurned passive metrics into operational signals

The value of the project was not data collection by itself. Across a 50-100 GPU cluster, usage, temperature, and memory metrics were collected every minute, and a threshold breach reached the operator dashboard and alert within 1-2 minutes - turning anomaly detection that used to take hours of manual checking into a minutes-level workflow for the 5-10 people who operate it.

RoleBackend and Frontend Contributor

Cloud Monitoring System

StackSpring Boot · InfluxDB · React · Time-Series Monitoring

Selected tools and technologies used in the work.

Q&A Depth View

A closer look at the questions behind the work

The summary above is for fast scanning. The Q&A below keeps the deeper story available without making the page feel noisy.

01Why does this project belong in the portfolio?

Monitoring only matters when people can act on it, not when the most data is collected. Across a 50-100 GPU research cluster, I collected usage, temperature, and memory metrics every minute and got threshold breaches onto the operator dashboard and alert within 1-2 minutes, cutting anomaly-detection time from hours of manual checking to minutes for the 5-10 people who operate it. I never measured an exact false-positive rate, but combining thresholds with a minimum duration was the specific decision that kept short-lived spikes from turning into alert fatigue - and that judgment call is what this project is here to show.

02What was my direct scope of ownership in this project?
  • Implemented GPU metric collection and ingestion paths.
  • Stored time-series metrics in InfluxDB so the system could support recent trends and operational queries.
  • Processed threshold rules in Spring Boot and generated alerts from those evaluations.
  • Helped shape the React dashboard so users and operators could read system state quickly instead of relying on raw logs.
03What evidence best supports the strength of this work?

Spring Boot + InfluxDB pipeline collects usage, temperature, and memory metrics every minute across a 50-100 GPU cluster.

Threshold breaches reach the operator dashboard and alert within 1-2 minutes, not just passive metric collection.

Combined thresholds with a minimum duration so short-lived spikes do not trigger repeat alerts, cutting anomaly-detection time from hours to minutes for 5-10 operators.

04Which technical decisions mattered most here?
Start with threshold-based alerting instead of noisier event-level signaling.

The goal was operational usefulness, not maximum sensitivity. A noisier system would have made alerts easier to ignore.

Use InfluxDB as the first storage layer for monitoring-style access patterns.

This matched the need for recent trends, threshold checks, and time-series queries without overcomplicating the first version.

Make the dashboard readable before making it feature-heavy.

The first release was about helping users and operators understand GPU state quickly, not exposing every possible monitoring control.

05How did the system work end to end in practice?
01
GPU telemetry is collected from live infrastructure.
02
Metrics are written into InfluxDB for time-series storage and recent trend access.
03
Spring Boot evaluates thresholds and detects alert conditions.
04
Alerts and processed status are exposed to the application layer.
05
React shows current status and trends for users and operators.
06What trade-offs did I make, and why?
OptionProsConsDecision
Polling-based collectionSimpler to build and operateLess immediate than streamingUsed polling for the first implementation
Alert on every notable eventLower chance of missing anomaliesHigh alert fatigueUsed threshold-based alerting to balance signal quality
07What changed as a result?

The system delivered real-time monitoring and alerting for a 50-100 GPU research cluster, cutting anomaly-detection time for its 5-10 operators from hours of manual checking to minutes.

GPU Monitoring & Alerting | Somin Moon | Somin Moon