Reliability for a distributed device network
Managed connectivity provider
99.99% uptime · 1000% fulfillment throughput · 85% smaller firmware, 12x faster startup
The problem
Hundreds of thousands of remote devices under contractual uptime obligations. Failures meant expensive field dispatches, and the firmware delivery process couldn't keep up with sales growth, with concurrent updates causing instability.
The approach
Established a Site Reliability Engineering practice with a clear charter and reliability metrics. Re-architected the firmware flashing system for parallel, reliable updates over wired and wireless connections, and guided Kubernetes adoption with zero-downtime deployments and shared observability.
The outcome
Core service uptime reached 99.99%, sharply reducing field visits. Firmware update success improved dramatically and fulfillment throughput scaled with revenue.