Nadia Belanger — Site Reliability Engineer
Chicago, IL
Site reliability engineer with nine years keeping high-traffic systems up, in freight logistics and healthcare claims. Took mean time to recovery across 60 services from 41 minutes to 7 by replacing host alerts with service-level objectives, and cut
.4M a year from the cloud bill in the same period.
Experience
Senior Site Reliability Engineer, Harborlight Freight Systems · May 2021 – present
Reliability for a logistics platform moving 90,000 shipments a day across 60 services.
- Replaced 400 host-level alerts with 38 service-level objectives, taking pages from 61 a week to 9 and mean time to recovery from 41 minutes to 7.
- Moved 60 services onto Terraform and Argo CD, turning a three-hour change window into an 11-minute rollout with automatic rollback.
- Cut .4M a year of cloud spend by right-sizing 220 workloads and running batch reconciliation on spot capacity with checkpointing.
- Runs the incident review program; repeat incidents fell from 31% of all incidents to 6% over 18 months.
Site Reliability Engineer, Cindermill Health · February 2018 – April 2021
- Built multi-region failover for a claims API handling 12,000 requests a second, proven quarterly by shifting real traffic in a game day.
- Cut p99 latency from 3.1s to 420ms with a read-through cache and by removing an N+1 query in the eligibility service.
- Automated certificate rotation across 180 hosts after an expiry took the member portal down for four hours; no expiry incident since.
Systems Engineer, Alderway Systems · July 2016 – January 2018
- Containerized 24 legacy Java services, taking environment-specific bugs from 18 a quarter to 2.
- Wrote the restore verification job that found 3 of 11 database backups had been failing silently for months.
Education
- B.S. Computer Engineering · University of Illinois Urbana-Champaign · August 2012 – May 2016
Skills
- Platform: Kubernetes, Terraform, Argo CD, AWS, Linux, Envoy
- Observability: Prometheus, Grafana, OpenTelemetry, Loki, PagerDuty
- Languages: Go, Python, Bash, SQL
- Practice: SLO design, Incident command, Game days, Capacity planning, Cost engineering
Certifications
- AWS Certified Solutions Architect – Professional — Amazon Web Services
- Terraform Associate — HashiCorp
Loading CVAurum…
