https://thenewstack.io/diary-of-a-first-time-on-call-engineer/ SEARCH (ENTER TO SEE ALL RESULTS) Cancel Search [ ] POPULAR TOPICS Contributed News Analysis The New Stack Makers Tutorial Podcast Research Feature Profile Science The New Stack Logo Skip to content * Podcasts * Events * Ebooks + DevOps + DevSecOps + Docker Ecosystem + Kubernetes Ecosystem + Microservices + Observability + Serverless + Storage + All Ebooks * Newsletter * Sponsorship * * * * + Podcasts o TNS @Scale Series o TNS Analysts Round Table o TNS Context Weekly News o TNS Makers Interviews o All Podcasts + Events + Ebooks o Machine Learning o DevOps o Serverless o Microservices o Observability o Kubernetes Ecosystem o Docker Ecosystem o All Ebooks + Newsletter + Sponsorship Skip to content * Architecture + Cloud Native + Containers + Edge/IoT + Microservices + Networking + Serverless + Storage * Development + Development + Cloud Services + Data + Machine Learning + Security * Operations + CI/CD + Culture + DevOps + Kubernetes + Monitoring + Service Mesh + Tools Search The New Stack 2022-03-09 08:16:39 Diary of a First-Time On-Call Engineer contributed,sponsored, Culture / What Is DevOps? / Sponsored / Contributed Diary of a First-Time On-Call Engineer 9 Mar 2022 8:16am, by Anna Baker [040bc11a-women-in-tech-1280x855-1-1024x683] LaunchDarkly sponsored this post. [f218b0cc-a] Anna Baker Anna Baker is a software engineer at LaunchDarkly, where she has spent the past few years advocating for inclusive processes and bridging the gap between product engineering and site reliability. She is also a part-time graduate student pursuing an M.S. in computer science at Georgia Tech. There is a clear business need for having engineers on call. One of the best ways to retain and gain customers is to deliver excellent, world-class service and a product they can rely on. Developing skills that drive business value not only will help engineers in their current role, but can prepare them for any role or level they may want in the future. Volunteering for on-call rotation can help the individual and the company. While taking on a new opportunity is exciting, it also comes with nerves. Maybe you have commitments outside of work like grad school classes or family obligations. Maybe you enjoy your evenings and weekends. What do you do if an alert comes in about a service that you're not familiar with? Change can be scary, which is often why processes don't change within organizations, but as companies grow, they may need to rethink their on-call rotation. Last year, LaunchDarkly modified our on-call process for engineers, as we had outgrown our previous model. A few years ago, engineers worked together on one of two large teams: the Application team and the Backend Services team. Small groups of people would come together for projects, but then disperse when the project was over. The Backend Services team handled all the on-call responsibilities. As we grew, we adopted the squad model. Each squad has an engineering manager, a product manager, a designer and five to seven engineers that work on a subset of our features. After some time, the squad model evolved and adopted service ownership. Each squad became responsible for a subset of our backend services. However, we didn't change the engineers who were on call in a substantial way. They were almost exclusively engineers who were at one point, or would have been, on the defunct Backend Services team. We decided a new process was needed. Change Happens There were multiple discussions internally about what the new on-call process should look like. We needed to make sure the rotation was equitable and that we had appropriate coverage. In the end, we decided on the following: * Squad members are on call during regular business hours for the services their squad owns. * For off-hours, responsibilities are distributed between the U.K. team during their normal business hours and the Virtual On-Call squad, a volunteer-based group of engineers split across two rotations that cover different subsets of our services, who take on primary responsibilities for the evening and weekend shifts. * On-call engineers with off-hours responsibilities are paid for their contribution. If you're considering changing your on-call rotation, have open conversations to get various perspectives on what works and what challenges might be encountered. How to Onboard New Engineers to the On-Call Rotation One of the most important aspects is to have an engineering culture that fosters learning and psychological safety. On-call engineers need an onboarding process that sets them up for success, knowing that they will have help, that if something goes wrong, they won't be blamed. Feeling safe to learn and explore means knowing it's OK to make mistakes. Be clear with people who are thinking to join an on-call rotation about what the expectations are. I received the following message from my manager when I was thinking about joining the rotation. "In general the expectation is to try your best with what you know, and if you don't know how to address the issue, escalate. Over time the people that are being escalated to will think 'hmm, next time if I don't want to get a page, I should arm the virtual squad with whatever it needs to handle this.'" In the weeks leading up to an inaugural shift, consider the following: * + Provide online or in-person training to give engineers confidence in the process and in their ability to succeed at being on call. + Host meetings and conduct question-and-answer sessions for the on-call rotation. + Have managers or leads check in on how the new engineers are feeling and send a test page. Normalize that it's OK to feel a rush of adrenaline when you get paged: Manager: Was this your first time being paged? Me: Yup Manager: Did your heart skip a beat? At least 10 years into on-call, I still sort of jump when I get an alarm :-) * Establish a co-pilot system with experienced on-call engineers who will pair up with onboarding on-call engineers as their on-call backup for their first few shifts. * Assign new engineers to an experienced co-pilot for the rotation and devise a plan for how to communicate if needed. Diary of a First-Time On-Call Engineer While the above advice may sound good in theory, you may be wondering, in practice, how do things go for new on-call engineers? I signed on to join the inaugural Virtual On-Call rotation, and below is my log of my first week on call. Day 1, Monday At around 8:30 p.m., I was at home in my jammies playing "The Sims 4" when I got my first legitimate page. It was thrilling! I hopped on my computer. Within minutes, I got another notification that a colleague on the other rotation had been paged for something related. We both hopped online. I suggested that we start a public thread in the virtual squad Slack channel instead of direct messaging so people could learn from our mistakes and help us improve the onboarding process. We spent about 45 minutes looking at the alert catalog, the runbook for the services and trying to fix the underlying problem. After getting more information, we realized it was not affecting customers and could wait until the team that owned the service came online. I spent another 30 minutes updating the Captain's Log, the log we use to communicate about events that might affect backend services, and notifying the squad that owned the service. Day 2, Tuesday Silence. Day 3, Wednesday At around 5:15 p.m., I got paged while still working. Ironically, the cause for this alert was the remediation for Monday's alert We realized it was an alert for a piece of system architecture due for retirement and no longer serving customer traffic. We throttled the offending service and silenced the alert until the morning. Badda bing, badda boom. Day 4, Thursday At 9 a.m., the alert from the night before unsnoozed itself and let me know that I need to figure out what to do about it. Note that typically I would not have been responsible for on call during business hours, but I had set up the alert to unsnooze then since I knew I'd be at my computer. I created a thread in my squad's on-call channel since the alert was for a service my squad owned and within minutes, it was clear that we could delete the alert and permanently wind down the service. Day 5, Friday Silence. Day 6, Saturday This was by far the most eventful day. Morning At 5 a.m., I got paged. I jumped out of bed and ran over to my computer. The page self-resolved at 5:01 a.m. Now wide awake, I spent some time on Google reading documentation about the service that had paged me before eventually going back to sleep. Afternoon I dared to venture to a park nearby. I brought my laptop and my cell phone with tethering capabilities and was prepared to run home if needed. Of course, Murphy's law, I got paged. I pulled out my laptop, tethered my phone and popped online. Sitting at a playground picnic table, I re-ran the test that had failed and alerted me. It passed. I was on standby but able to enjoy the rest of my day. Day 7, Sunday I woke up and thanked PagerDuty for no 5 a.m. alert. The rest of Sunday was quiet as well. Sponsor Note sponsor logo Unleash developer productivity for the software-powered world by fundamentally changing how you deliver software to your customers. With LaunchDarkly's feature management platform, empowered developers can empower the business to release new features faster and more efficiently than ever. LaunchDarkly and TNS are under common control. What I Learned At the beginning of the week, I was prepared to have to declare multiple incidents and getting paged constantly. Instead, most days were quiet. I was surprised by how few alerts there were, especially low-priority alerts, which I thought were going to be incessant. I attribute this to LaunchDarkly's commitment as an organization to scalability and investment in making our services production-ready. Perhaps I got exceptionally lucky this week, but overall, it was enthralling, and I'm glad I volunteered. It gave me an incentive to Google things I wouldn't typically research, read the alert catalog and service runbooks, and I learned a bunch. * Was there an increased cognitive load from having LaunchDarkly in the back of my mind 24/7? Yes. * Did I check my phone too many times, anxious about missing an alert? Also, yes. * Will those things improve over time? Probably. * Was it thrilling? Did I learn anything? Yes and yes! If you're interested in learning more about incidents, sign up to attend IRConf, a virtual event dedicated to all things incident response. The New Stack is a wholly owned subsidiary of Insight Partners, an investor in the following companies mentioned in this article: LaunchDarkly. LaunchDarkly is a sponsor of The New Stack. Feature image via Nappy ContributedSponsored The New Stack Newsletter Sign-Up A newsletter digest of the week's most important stories & analyses. [ ] Do you also want to be notified of the following? [ ] Send me everything :-D [*] TNS Weekly Update [ ] Upcoming ebook notifications [ ] Research surveys [ ] Upcoming event notifications [ ] New product & service notifications Subscribe [tns-button] We don't sell or share your email. By continuing, you agree to our Terms of Use and Privacy Policy. Related Stories [] [f3851b14-businessman-g3909e689a_1280-290] API Management / Microservices / Sponsored / Contributed Why GraphQL for Microservices? 6:22am, by Roy Derks [] Five grain silos standing next to each other DevOps Tools / Kubernetes / Security / Software Development / Technology / Sponsored Software Supply Chain Security: Tearing Down the Silos 3:00am, by B. Cameron Gain Sponsored Feed [35c9e830-o] Brain to the Cloud - Part III - Examining the Relationship Between Brain Activity and Video Game Performance March 16, 2022 [20f61f0a-r] As an OpenShift Administrator, what questions should you ask your developers March 14, 2022 [NewRelic] A faster and cheaper way to send log insights to New Relic One with the Cloudflare Logpush integration March 14, 2022 [e69eebf6-o] Deploy and Configure Kustomize and MongoDB for Kubernetes Stateful Applications March 14, 2022 [7160e7db-i] Start with Python and InfluxDB March 14, 2022 [SAP] A Fiori Launchpad Sandbox for all your CAP-based projects - Sample project setup March 14, 2022 [0b8a7166-g] Installing GitLab on Raspberry Pi 64-bit OS March 13, 2022 [5e8bc8be-w] Accelerating Banking Innovation with Speed and Agility March 13, 2022 [59ec0897-q] Deploying Docker Containers on AWS: Elastic Beanstalk vs ECS vs EKS March 12, 2022 [383b2ee5-s] Major Government Attack Highlights How Log4j is Still Unresolved March 11, 2022 [b208f6dc-l] Tecton and Redis partner to improve real-time ML services March 11, 2022 [44399732-l] Welcome, Tim Silva! March 11, 2022 [3feae3d9-l] Lacework named only security company in top-10 of Forbes' Best Startup Employers March 11, 2022 [3ef8550d-h] Vault, Boundary, and Zero Trust Videos from HashiTalks 2022 March 11, 2022 [d79020ae-t] March Monthly Tech Talk: Transform Data on the Fly | TriggerMesh Blog March 11, 2022 [CNCF] Flux Security: More confidence through fuzzing March 11, 2022 [69778945-p] What You Can Do With Different AWS Database Services and Kubernetes March 11, 2022 [78f349c9-n] Using DNS to Minimize Cyber Threat Exposure March 11, 2022 [d2f23bca-c] Predict the cost of IP ranges with new enhancements to the Resources tab March 11, 2022 [d011076d-s] Animating API Results (On a Budget) March 11, 2022 [7db0ad57-b] Evolving Our Business by Putting Our People First March 10, 2022 [4cab8290-p] Celebrating the Changing Face of Leadership March 10, 2022 [86b6391d-s] How to Implement Global View and High Availability for Prometheus | Squadcast March 10, 2022 [3eabf683-t] Why Single Sign On Sucks. March 10, 2022 [e97c353e-t] Zero-trust for cloud-native workloads March 10, 2022 [9fad41ed-n] Exposing APIs in Kubernetes March 10, 2022 [d47c41d2-m] Announcing Mirantis Training Badges March 10, 2022 [b22fecc6-l] Block Joins the Linux Foundation March 10, 2022 [25c13654-a] New - Amazon EC2 X2idn and X2iedn Instances for Memory-Intensive Workloads with Higher Network Bandwidth March 10, 2022 [dd39254e-m] Understanding the MongoDB Stable API and Rapid Release Cadence March 10, 2022 [2bac37aa-c] An Introduction to Data Mesh March 10, 2022 [599a31ec-s] Password Policy Best Practices | strongDM March 10, 2022 [6996f2a0-a] Armory Agent: Managing Kubernetes at Scale March 10, 2022 [d3745942-t] Russia and the Global Internet March 10, 2022 [3976e109-r] Why Mutability Is Essential for Real-Time Data Analytics March 10, 2022 [c1030d28-s] Synopsys contributes to the Linux Foundation Census II of the most widely used open source application libraries March 10, 2022 [09390306-s] Moving Your Healthcare Organization to the Cloud? Here's What You Need to Know First March 10, 2022 [7fb6d54e-l] Getting Started with Svelte and LaunchDarkly March 10, 2022 [67e6d14a-e] Why We're Raising Prices and Not Blaming the Supply Chain March 09, 2022 [f96d5427-t] Getting more from codeless test automation with strategic test management March 09, 2022 [d5d3aa3b-s] Liveblog: Second Day at SoloCon 2022 Wrap Up March 09, 2022 [7942c19b-s] Introducing SansShell: A Non-Interactive Local Host Agent March 09, 2022 [eaae9108-t] 5 Challenges to Security Operations Strategies March 09, 2022 [9c7dadea-l] Cloud Native Journey Part 2: Technical Adventure March 09, 2022 [d2d0180b-a] New Linux Kernel Vulnerability: Escaping Containers by Abusing Cgroups March 09, 2022 [4a53d638-d] Optimizing Java XPath CPU and memory overhead by 98% March 09, 2022 [06874d86-k] Dogfooding Testkube - Part1 - How to Test a Testing Framework - Kubeshop March 08, 2022 [e73fdfe4-c] Partner Opportunities with Curity March 08, 2022 [669d57ea-r] Run Containers and VMs Together with KubeVirt March 08, 2022 [e690512a-t] Microsoft's March 2022 Patch Tuesday Addresses 71 CVEs (CVE-2022-23277, CVE-2022-24508) March 08, 2022 [c2c15fb4-h] On the brittleness of dashboards March 08, 2022 [042245f3-p] Inspiring Women to Join the DevOps Movement by Natalie Fair March 07, 2022 [ad93ed2e-m] Multi-Cloud Monitoring and Alerting with Prometheus and Grafana March 03, 2022 [348c14df-i] Upgrades to Our Internal Monitoring Pipeline: Using Redis(tm) as a Cassandra(r) Cache March 02, 2022 [c2b4a6fb-k] SUSE Rancher Ropes in Kasten K10 by Veeam March 02, 2022 [b22fecc6-l] Participate in the 2022 Open Source Jobs Report! March 02, 2022 [d533f22f-a] Trusting SBOMs in the Software Supply Chain: Syft Now Creates Attestations Using Sigstore March 02, 2022 [0467a55b-s] Scylla University LIVE - Spring 2022 February 28, 2022 [831e201b-n] Orchestrate Spark pipelines with Airflow on Ocean for Apache Spark February 23, 2022 [d120d17f-i] itopia Secures $5 Million in Funding as the Global Workforce Embraces the Cloud at Record Pace March 02, 2021 [ff5fbb9d-c] Citrix Deployment Builder: Simplifying Citrix cloud-native deployments March 02, 2021 [dd39254e-m] We Replaced an SSD with Storage Class Memory. Here is What We Learned. August 27, 2020 [951da3dd-v] Join us at our new blog home July 02, 2020 [116b3300-i] What Is Quantum-Safe Cryptography, and Why Do We Need It? December 31, 1969 Architecture * Cloud Native * Containers * Edge/IoT * Microservices * Networking * Serverless * Storage Development * Cloud Services * Data * Development * Machine Learning * Security Operations * CI/CD * Culture * DevOps * Kubernetes * Monitoring * Service Mesh * Tools The New Stack * Ebooks * Podcasts * Events * Newsletter * About / Contact * Sponsors * Sponsorship * Disclosures * Contributions * Twitter * Facebook * YouTube * Soundcloud * LinkedIn * Slideshare * RSS (c) 2022 The New Stack. All rights reserved. Privacy Policy. Terms of Use.