Uploaded June 2026 | Updated September 2026, 1 week ago
Abstract
We've all been there: it's 3AM, the pager fires, and the only person who knows what to do is asleep in another timezone. If your incident response looks different depending on who's on call or what time it is, that's not a staffing problem, it's a systems problem. This talk argues that "boring" is the highest compliment you can pay your operations. Drawing from hands-on experience automating a Tier-1 IP backbone, I'll walk through a practical framework built on three pillars: architecture that isolates failure domains so outages stay small, automation that actually heals infrastructure instead of just screaming about it, and documentation that lives in runbooks, not in someone's head. The goal is simple: the same incident should produce the same response whether it's a Tuesday afternoon or Christmas morning.
I'll also talk about tabletop exercises, chaos engineering, and how to get your org to treat reliability as a first-class investment rather than an afterthought. The benchmark I keep coming back to: if a junior engineer three months into the job can't handle the same incident as your senior DBA, you have homework to do. Boring isn't lazy, it's disciplined, and it's how your team sleeps through the night.
Ryan Hamel:
Ryan is a Senior Site Reliability Engineer at Zayo, where his work spans across optical transport, IP networks, and cloud platforms, with a focus on making complex systems more predictable, observable, and easier to operate.
Within the NANOG community, Ryan serves as Co-Chair of the Moderation Committee and helps run the organization’s online platforms. He is particularly interested in reducing operational toil, applying SRE principles to networking, and helping engineers ship practical solutions.
nanog.org/events/nanog-97/content/5737
Abstract
We've all been there: it's 3AM, the pager fires, and the only person who knows what to do is asleep in another timezone. If your incident response looks different depending on who's on call or what time it is, that's not a staffing problem, it's a systems problem. This talk argues that "boring" is the highest compliment you can pay your operations. Drawing from hands-on experience automating a Tier-1 IP backbone, I'll walk through a practical framework built on three pillars: architecture that isolates failure domains so outages stay small, automation that actually heals infrastructure instead of just screaming about it, and documentation that lives in runbooks, not in someone's head. The goal is simple: the same incident should produce the same response whether it's a Tuesday afternoon or Christmas morning.
I'll also talk about tabletop exercises, chaos engineering, and how to get your org to treat reliability as a first-class investment rather than an afterthought. The benchmark I keep coming back to: if a junior engineer three months into the job can't handle the same incident as your senior DBA, you have homework to do. Boring isn't lazy, it's disciplined, and it's how your team sleeps through the night.
Ryan Hamel:
Ryan is a Senior Site Reliability Engineer at Zayo, where his work spans across optical transport, IP networks, and cloud platforms, with a focus on making complex systems more predictable, observable, and easier to operate.
Within the NANOG community, Ryan serves as Co-Chair of the Moderation Committee and helps run the organization’s online platforms. He is particularly interested in reducing operational toil, applying SRE principles to networking, and helping engineers ship practical solutions.
nanog.org/events/nanog-97/content/5737










