Tag: SLO

  • Driving DevOps Excellence: Implementing SLOs for Your DevOps Team

    Driving DevOps Excellence: Implementing SLOs for Your DevOps Team

    In today’s fast-paced world of software development, when continuous deployment and frequent releases are the norm, DevOps teams are essential to facilitating smooth communication between development and operations. In a blog post on LinkedIn, I talked about the importance of service level objectives (SLOs) for the testing team to improve the overall product quality. On top of that, let’s look at how adding SLOs to your DevOps team can improve productivity, dependability and client happiness.

    The Role of SLOs for DevOps Teams

    • Aligning DevOps Goals with Business Objectives:

    SLOs help DevOps teams coordinate their efforts with business objectives. The DevOps team can concentrate on delivering real business value by setting precise performance indicators, such as deployment success rates or infrastructure provisioning timeframes.

    • Promoting Collaboration and Accountability:

    By implementing SLOs, several stakeholders—including development, operations, quality assurance and business teams—are encouraged to work together. Over each stage of the software development life cycle, this shared responsibility promotes a sense of ownership and accountability.

    • Driving Reliability and Stability:

    SLOs play a key role in ensuring the stability and dependability of the system. You can be sure that your services consistently fulfill client expectations when your DevOps pipeline adheres to defined SLOs.

    • Proactive Issue Mitigation:

    Monitoring and alerting systems are used together with SLOs. The DevOps team could avoid service outages and downtime by aggressively identifying possible problems and resolving them before they become more serious by regularly monitoring important indicators.

    • Data-Driven Decision Making:

    Decision-making is driven by quantitative data. With the help of such indicators, the team is able to identify bottlenecks, prioritize improvements and streamline processes based on quick feedback.

    Now, let’s explore some key areas where DevOps teams can set SLOs to improve their performance:

    1. Continuous Integration (CI):
    SLO: “xx% of builds complete within Y minutes.”
    Measurement: Monitor build times and queue times regularly.
    Action: Optimize CI infrastructure and configurations to meet the SLO.

    2. Continuous Deployment (CD):
    SLO: “xx% of deployments are successful.”
    Measurement: Track deployment success rate.
    Action: Improve the deployment process to meet the SLO and reduce deployment failures.

    3. Infrastructure Management:
    SLO: “xx% of infrastructure is provisioned within Y minutes.”
    Measurement: Monitor infrastructure provisioning time.
    Action: Optimize infrastructure provisioning scripts to meet the SLO.

    4. Monitoring and Logging:
    SLO: “DevOps tools and system uptime should be at least xx%.”
    Measurement: Monitor the availability of the DevOps pipeline, deployment system and other tooling, including monitoring and logging systems.
    Action: Ensure high availability for DevOps tools and components.

    5. Artifact Management:
    SLO: “Artifact retrieval time should be less than x seconds on average.”
    Measurement: Monitor artifact retrieval time and availability.
    Action: Optimize artifact storage and distribution mechanisms.

    6. Testing and Quality Assurance:
    SLO: “Code must have at least xx% unit test coverage.”
    Measurement: Track test coverage regularly.
    Action: Encourage developers to write more tests.

    7. Security and Compliance:
    SLO: “xx% of compliance checks must pass.”
    Measurement: Monitor the results of compliance checks.
    Action: Ensure necessary security measures to meet the compliance SLO.

    8. Standardized Tools Selection:
    SLO: “xx% of teams must use the approved CI/CD tools stack.”
    Measurement: Track the percentage of teams using the approved tools stack.
    Action: Encourage teams to adopt standardized tools and provide necessary training and support.

    9. Training and Skill Development:
    SLO: “xx% of team members should undergo relevant DevOps training annually.”
    Measurement: Monitor training completion rates.
    Action: Provide training opportunities and resources to help team members enhance their skills.

    Initially, teams should baseline which percentage values to track by observing the current state. In the event they don’t have the bandwidth to baseline, start with any reasonable arbitrary number over a period of time, and it will be automatically refined.

    Implementing service-level objectives empowers DevOps teams to focus on delivering reliable and high-performing services that meet user expectations. By setting clear performance and reliability targets, teams can proactively identify and resolve issues, leading to improved collaboration, efficiency and overall user satisfaction.

    SLOs are not rigid constraints but are a means to drive continuous improvement and foster a culture of excellence within the DevOps ecosystem. As organizations strive to keep up with the ever-changing demands of the digital world, embracing SLOs is a crucial step toward achieving DevOps excellence and ensuring a competitive edge in the market.

  • Black Box SLIs

    Black Box SLIs

    This article is a preview of a talk by Stephan Lips for SLOconf 2023, on May 15 – 18. To watch this talk and many more like it, register for free at sloconf.com.

    SLOs are fast becoming the industry standard to measure reliability and help teams decide when to prioritize it. The first step in adopting a service level objective (SLO) culture is to identify the metrics that matter without drowning in noise and alert fatigue. This article explores how to apply the black box concept to aggregate granular metrics into service level indicators (SLIs) that focus on the user experience as an indicator of system reliability.

    To SLI or Not to SLI

    In general terms, SLOs define targets for the proper level of reliability of a given product, such as a service or a website. SLOs are applied to or informed by SLIs. An SLI is a measurement determined over a metric, or a piece of data, representing some property of a service. And this is where we, as engineers, can get lost in the details, since the perpetual proximity to the systems we build and support often leads us to think of system reliability in technical terms or metrics (e.g., response time, error rate, throughput). While these are certainly valuable metrics, the user experience may be compromised even if the error rate is zero and the duration is well within SLOs. Consider, for example, the response data. Even if well-formed, it may not be current, or flat-out wrong. An error-free and quick response is of no value to a user that expects current and correct data. Error rate and response time remain valuable metrics and SLIs, but focusing exclusively on them would leave higher-level issues undetected.

    We could add freshness and correctness SLIs, but by doing so, we increase the number of signals we monitor. And with each signal—or SLI and associated SLO and error budget—we increase alert frequency and make reliability reports unnecessarily complex. In other words, adding SLIs may address a particular aspect of system reliability, but it also introduces additional complexities.

    Tales of Black and White Boxes

    So, let’s take a step back and borrow a concept from a related discipline: Quality engineering—in particular, software testing. Tests commonly fall into one of two categories: Black box tests or white box tests.

    In systems theory, the black box is an abstraction representing a class of concrete open systems that can be viewed solely in terms of its stimuli inputs and output reactions, without any knowledge of its internal workings. A given input is expected to result in a particular output, without any consideration for the processing steps. Common examples include end-to-end tests.

    White box tests, on the contrary, are designed with knowledge of, and to test, internal structures and workings of an application. Common examples include unit and integration tests.

    User Journey as Black Box

    Now that we understand the concept of black box versus white box tests, let’s apply it to our SLIs. As mentioned above, a good SLI considers the entire user journey. Conceptually, a user journey aligns with the black box paradigm: For a given input, a particular output is expected. For example, requests to our API (the “input”) result in responses that provide fresh data to clients within a given time frame (the “output,” including success criteria). There are several aspects worth mentioning with this SLI:

    ● The SLI is applied at a system level
    ● The SLI aggregates lower-level metrics implicitly and explicitly
    ● The SLI is binary; it is either true or false.

    These aspects combine to inform an SLI that represents the user experience (system level), via measuring many indicators by measuring only a few and supporting pass/fail attribution to an SLO target (by being binary). In other words, the user journey is measured as a black box SLI.

    White Box to Black Box: An Example

    Let’s consider a concrete example. A user requests a new account for a website. After the request is processed successfully, the user receives a confirmation email with an activation link. The user follows the link to activate the new account and log in. This workflow is visualized in the following sequence diagram.

    Of particular interest are the account creation and user notification via email steps. Both steps occur asynchronously. In particular, the event processing engine where the request for account creation is queued offers several opportunities for insightful SLIs: Queue length, average processing time, etc. Those SLIs, however, fall into the white box category: They contribute to the user experience, yet are opaque to the user (black box). The user journey begins with the initial request for account creation (input) and ends with the email containing the activation link (output). Rephrasing the example from earlier—a user-focused (black box) SLI could be a request for a new account (the ‘input’) that results in sending an email with a valid activation link within 1 minute (the ‘output’, incl. success criteria). This single high-level SLI aggregates several lower-level metrics; it measures many things by measuring only a few.

    Let’s switch to the engineering mindset mentioned in the introduction and assume the processing queue is stuck. The high-level black box SLI does not capture queue-specific metrics, suggesting a more granular SLI specific to queue size may be needed. However, white box metrics like this will affect the error budget burn of the aggregate SLO associated with the high-level black box SLI. Monitoring and observability tools will allow engineers to diagnose and troubleshoot particular issues, such as a stuck queue, while understanding the impact on system reliability (via the higher level, black box SLO’s error budget and burn rate). The solution to the stuck processing queue used in this example is not an SLI dedicated to the queue, but reliability-focused work to diagnose and correct the root cause of the queue getting stuck.

    Summary

    This article introduces an SLI thought model that uses a common paradigm from quality engineering. This thought model offers a different way to think about SLIs. It supports implementing the fundamental objective of SLIs and the associated reliability stack they inform: Ensure a positive and reliable user experience by measuring reliability and providing quantitative support for decisions on prioritizing development efforts. Only a happy user is a continuous user.

  • How IT Ops Can Exceed Service Level Objectives in Digital Transformations

    How IT Ops Can Exceed Service Level Objectives in Digital Transformations

    The pace of change can be managed successfully by defining service level objectives and more in dev environments

    Mobile applications, data lakes, microservices, data visualizations, SaaS integrations, automations, IoT data streams, machine learning models—in proof of concepts, pilots and scaling production environments, for customer-facing capabilities and employee workflows—all of these technical capabilities are developed, deployed and enhanced faster today more than ever before.

    I spoke to Jason Walker, field CTO at AIOps platform BigPanda, about how the speed of deployment, the breadth of technology services transforming businesses are developing, the greater security threats, and the increase in reliability and performance requirements impact IT Ops.

    Walker believes that of all the things we’re trying to do in IT—more, faster, smarter, safer, innovative, secure, reliable—it’s the speed that’s the driving force. “The most significant impact is velocity; the dev-test-deploy cycle time is drastically reduced,” he said. “Without the right guardrails, that breeds unnecessary complexity and a gradual loss of operational awareness.”

    He explained that much of the barriers that once slowed down development teams are addressable today when developing service-based architectures on the cloud. “Developers realize that the traditional constraints, either technical dependencies or organizational capability, are much reduced,” Walker noted. “Developing in the cloud for a cloud-based service, leveraging an ecosystem of microservices for inputs, an agile team can move very quickly.”

    Why IT Ops Can’t Slow Down Transformations

    IT Ops can’t easily say “no” or “slow down” to business stakeholders investing in digital transformation to improve customer experiences and gain competitive advantages with data, analytics and machine learning. Some IT leaders attempted that command-and-control approach during the early days of public clouds, but today, DevOps practices, SRE responsibilities and AIOps capabilities are integral to mainstream IT Ops teams in keeping up with transformational velocities.

    So instead of saying “no,” progressive IT Ops teams say “yes, but” by defining service level objectives, (SLOs), capturing service level indicators (SLIs) and managing to error budgets.

    Walker agreed. “SLIs, SLOs and error budgets are a very useful way to manage the critical inputs and outputs at the interfaces between microservices and at a high level across the business service, allowing developers to keep changing the ‘interior’ pieces.“

    These tools change the operating model and mindset across the entire IT organization by exposing trade-offs to business stakeholders. For example, if a web application has a 99.9% SLO, the whole IT team has a 0.1% error budget. If the SLO is missed, a service level policy identifies areas of investment to improve performance, reliability, security and automation or to address technical debt.

    What IT Teams Should Do to Implement Service Level Objectives

    Defining service level objectives helps bring business stakeholders, development teams, SREs and IT Ops together and align on reliability objectives and trade-offs. It’s an important step, but not sufficient for teams that want to exceed service levels during digital transformation.

    Walker offered several technical recommendations for IT Ops groups transitioning to service level objectives:

    1. Where there are dependencies between microservices and interfaces, frequent check-ins between adjacent teams are required.
    2. APIs, inputs and outputs, need to remain consistent over time and changes communicated and receipt confirmed.
    3. Knowledge management processes delivering accurate, up-to-date documentation have to be baked into the SDLC to prevent the generally small, modular teams from sprawling away from each other and developing incompatibilities.
    4. Consolidated change awareness for operations is also a must-have. Whether changes are human or automated, they have to be tracked and relatable to service events and alerts.
    5. The later phases of the SDLC, when cloud-native, containerized microservices are running and supporting customers, have to be monitored. A monitoring strategy is necessary to ensure effectiveness and minimize the work and noise involved.
    6. Synthetics and client telemetry can be very useful macro-indicators of overall service performance. As with all monitoring efforts, actionability is key. Signal-to-noise in monitoring has to be measured.

    These are balanced recommendations, with the first two focused on how development teams engineer microservices and the last two on how IT Ops teams use monitoring and AIOps to implement actionable SLIs. The middle two recommendations on knowledge and change management processes help the entire organization stay in sync through a fast-paced operating environment.

    “The velocity, flexibility and variability of this type of development means leaders at all levels need to understand and align the strategic and tactical goals, and to prevent their teams from drifting away from business goals,” Walker said.

    Why AIOps is a Force Multiplier for IT Transformation

    So, the business leaders align on service level objectives, development teams engineer observable containerized microservices, IT Ops executes a monitoring strategy and the CIO ensures communication, collaboration and knowledge sharing. Is that all that’s required?

    The issue is that most IT Ops teams are understaffed and get overwhelmed supporting the new cloud-native microservices, legacy systems and everything in between. To address the gap, IT leaders are investing in AIOps, and even hundred-year-old enterprises are successful at adopting machine learning and automation to accelerate IT Ops.

    Automation and machine learning event correlation applied to monitors, alerts and observable artifacts are the force multipliers. Open-box machine learning enables IT Ops to triage incidents and improve their mean time to resolution, while automation reduces manual efforts and keeps teams in sync. Organizations that are modernizing applications and supporting hybrid clouds require these capabilities to manage the complexities, run at business speed and manage databases, microservices and applications to higher service level objectives.

    As Walker noted, increasing velocity is important for responding to customer opportunities and changing conditions. Driving faster digital transformations is more achievable today when IT Ops leverages automation and AIOps to stabilize speed with reliability and performance.

  • Achieving Reliable Observability Part 1 – Making Cloud-Native Observability More Robust

    Achieving Reliable Observability Part 1 – Making Cloud-Native Observability More Robust

    I was having a conversation with a CxO level customer as part of an AIOps/Observability workshop, and from what I could tell, most are confused about how to properly operationalize cloud-native production environments – especially the monitoring/observability portion. Here is how the conversation went.

    “Andy, we are thinking about getting [vendor] to use for our observability solution based on your recent research. What do you think?”

    “Well, I don’t want to endorse any specific vendor, as they are all good at what they do. But let’s talk about what you want to do, and what they can do for you, so you can figure out whether or not they are the right fit for you.” The conversation continued for a while, but the last piece is worthy of being called out specifically.

    “So, we will be running our production microservices in AWS in the ____ region. And we are planning to use this particular observability provider to monitor our Kubernetes clusters.”

    “Couple of items to discuss. First, you realize that this particular provider you are speaking of also runs in the same region of the same cloud provider as yours, right?”

    “We didn’t know that. Is that going to be a problem?”

    “Not particularly. However, you may get into a ‘circular dependency’ situation.”

    “What is that?”

    “Well, as an enterprise architect, I always call for separation of duties as a best practice. For example, having your developer testing the code is a bad idea, having your developer figuring out how to deploy is a bad idea. In much the same way as when your production services run in the same region as your monitoring software – how would you know about a production outage if the cloud region takes a hit, and your observability solution goes down at the same time your production services do?”

    “Should we dump them and go get this other solution instead?”

    “No, I am not saying that. Figure out what you are trying to achieve and have a plan for it. Selection of an observability tool should fit your overall strategy.”

    For those who don’t understand the above conversation, here is the reason why this scenario could be a problem.

    Coming from an enterprise architecture background, we were taught, as a best practice, to operationalize production systems to avoid circular dependencies. This includes not having two services depend on each other, or not to colocate monitoring, governance and compliance systems as part of the production systems themselves. If you were to monitor your production system, you would do it from a separate and isolated sub-system (server, data center rack, sub-net, etc.) to make sure that if your production system goes down, the monitoring system doesn’t go down, too.  The same goes for public cloud regions – although it’s unlikely, individual regions and services do experience outages. If your production infrastructure is running on the same services in the same region as your SaaS monitoring provider, not only won’t you be aware that your production systems are down, but you also won’t have the data to analyze what went wrong. The whole idea behind having a good observability system is to quickly know when things went bad, what went wrong, where the problem is and why it happened so you can quickly fix it. You can check out this blog where I explain this in detail.

    The best practice would be to either:

    1. Consider an observability solution that runs in a different region than your production workloads. Better yet, consider something that runs on a different cloud provider altogether. Although it is exceptionally rare, there have been instances of cloud service outages that cross regions. The chances of both cloud providers going down at the same time would be slim.For example, if the cloud region goes down (a region-wide outage in the cloud is quite possible, and seems to be more frequent of late), then your observability systems will also be down. You wouldn’t even know your production servers are down to switch to your backup systems unless you have a “hot” production backup. Not only will your customers find out about your outage before you do, but you won’t even be able to initiate your playbook, as you are not even aware that your production servers are down.
    2. Consider having your observability solution in a different location/region, yet still close enough to your production services so latency is very low. Most cloud providers operate in close proximity, so it is easy to find one.
    3. Another option is to get a solution that gives you deployment flexibility. For example, there are a couple of observability solutions that allow you to deploy it in any cloud and observe your production systems from anywhere – both in the cloud and on-premises.
    4. You can also consider sending the monitoring data from your instrumentation to two observability solution locations, but that will cost you slightly more. Or, ask what the vendor’s business continuity/disaster recovery plans are. While some think the costs might be much higher, I disagree, for a couple of reasons. First, because monitoring is mainly time-series metric data, so the volume and the cost to transport is not as high as logs or traces. Second, unless your observability provider is VPC peered, the chances are your data will be routed through the internet even though they are hosted in the same cloud provider. Hence, there will not be much more additional cost. Having observability data, all the time, about your production system is very critical during outages.
    5. A very commonly overlooked consideration is monitoring your full-stack observability system. While it is preferable to have the monitoring instance in every region where your production services run, it may not be feasible either because of cost or manageability. On such occasions, monitor the monitoring system. You could do synthetic monitoring by checking the monitoring API endpoints (or do random data inputs and check to ensure it worked) to make sure that your monitoring system is properly watching your production system. Better yet, find a monitoring vendor that will do this on your behalf.

    When/if you get the dreaded 2 a.m. call, what is your plan of action? Just think it through thoroughly before it happens and have a playbook ready, so you won’t have to panic in a crisis.

  • Nobl9 Ties Business Goals to Observability Data

    Nobl9 Ties Business Goals to Observability Data

    Nobl9 today announced it has made available via an open public beta program a software-as-a-service (SaaS) platform for correlating business goals against data collected by observability tools.

    Unlike observability platforms that only aggregate metrics, the Nobl9 Service Level Objective (SLO) Platform applies that data to specific reliability targets defined by the business, company CEO Marcin Kurc said.

    Compatible with monitoring platforms such as Datadog, New Relic and Prometheus, Nobl9 SLO Platform calculates uses monitoring data to calculate acceptable rates of error per service threshold and can be configured to trigger alerts and even workflows in anticipation of outages, Kurc said. The critical difference is knowing which metrics are truly impacting customer experiences.

    The platform also enables DevOps teams to create business rules and define “facets” of users based on application experiences required. DevOps teams can identify users who are being serviced poorly in addition to tracking groups of users. They also codify contractual service level agreements (SLAs) down to critical business periods, said Kurc. The Nobl9 SLO Platform is accessible via a CLI/GUI/API for everything, Kubernetes-like YAML SLO-as-code or a sloctl command line.

    This approach also has a significant impact on the cost of monitoring because the Nobl9 SLO Platform makes it easier to identify what data needs to be stored versus deleted, noted Kurc.

    Armed with those insights, Kurc said it then becomes easier for DevOps teams to either determine how to make systems more reliable or lower reliability goals to reduce costs when possible. IT organizations will also spend much less time finger-pointing because the root cause of any issue will be much more readily apparent, he noted, adding the Nobl9 SLO Platform provides access to both real-time and historical reports to enable DevOps teams to achieve both those goals.

    In the wake of the COVID-19 pandemic, there’s a lot more focus on observability than ever. Organizations are relying more on digital business process to survive. The difference between surviving and thriving in this new era will come down to observability of not just the IT environment but also the entire digital business process.

    It may be a while before most organizations fully equate profit and revenue to the performance of their IT environments. However, it’s becoming increasingly clear that it’s not enough to simply make sure IT environments are available. Customers are judging organizations by their ability to drive a quality experience within the context of a digital business process spanning multiple applications. Organizations that are unable to meet those expectations will soon find themselves being cast aside in favor of those that can.

    Naturally, those requirements will put a lot of pressure on IT teams that often don’t have a lot of visibility into a business process. There is no end of alerts being generated by the system that are closely monitored, but the bulk of those alerts lack business context. Undoubtedly, there soon will be plenty of ways to gain that context. The challenge now is explaining to business leaders now why it’s needed before it’s too late.

  • How To Build a Culture of Resilience Through Good Habits

    How To Build a Culture of Resilience Through Good Habits

    Good habits are hard to form. I’ve been listening to the audiobook “Atomic Habits“ by James Clear on my morning runs, and something struck me. At Gremlin, along with our software, what we’re trying to promote are positive new habits for our customers. According to the author, one of the primary reasons new habits don’t stick is because there’s often sacrifice without immediate gratification.

    Psychologically, we’re wired to want instant gratification. But not all habits give immediate rewards; in fact, many delay our gratification for some time. So how do we help ourselves pursue good habits? Put simply: We need to make them obvious, attractive and easy.

    One of the ways to make them easy is to focus on what the author calls “gateway habits”—the smallest piece of the habit that can reasonably be achieved in two minutes. So if your goal is to eventually run a marathon, the gateway habit is putting on your running shoes every day.

    To build a culture of resilience at your company, start small and create getaway habits. If a team runs a GameDay once a month (time dedicated to experimenting on your systems) or even simply runs their first single chaos engineering experiment, then award that person or team with immediate recognition.

    Here are some other ways to build a culture of resilience at your organization:

    • Recognize the change to new habits.
    • Create DNR (do not repeat) items.
    • Adopt “You build it, you own it.”
    • Track the four golden signals.

    Recognize the Change to New Habits

    Incentives are a great way of kick-starting a new habit, but they don’t necessarily sustain the good behavior. It’s identifying the improvements that result from the new habit that really makes it stick. We’ll get to some specific metrics you can track later in the article, but notice for now that identifying the improvements is when gratification starts to drive enthusiasm. Ideally, that enthusiasm grows until the new habit becomes part of your identity.

    In our example of running a marathon, the moment of most significant change is when the person starts to self-identify as a runner. Then, the habit is no longer a chore, but rather part of who they are. Ideally, we want all computer engineers to adopt a specific set of habits until they consider themselves site reliability engineers (SREs) as part of their identity. That’s when the habit is solidified. 

    Create DNR (Do Not Repeat) Items

    Engineers and product managers want to ship new products and features. There’s nothing quite as satisfying in software development as deploying new code and seeing what you built running out in the world. But, if what you built is consistently breaking or providing a bad user experience, then you are hurting your customers and ultimately your business.

    To make sure we are always learning and getting better, at Gremlin we have what we call “DNR,” or Do Not Repeat work. This work consists of action items from outages and incidents that must not ever be repeated, lest we fail to learn our lesson from these failures. Practically, what this means is all feature work is halted until the issues highlighted as DNR work are remedied and the fixes are verified. In other words, you don’t get to write new code until your old code is fixed. We all know that many teams struggle with the trade-offs of moving fast, but ultimately, if you don’t have strict guidelines in place, then more often than not engineers will prioritize shipping something over making sure it’s reliable.

    Creating a DNR item is an easy way to incentivize the behavior you want to see internally by appealing to the engineer’s desire to produce new features. We convince them to write better code because better code means they get to spend less time fixing things.

    Adopt “You Build It, You Own It”

    This is the driving principle of DevOps. It is the reason behind shifting left. When the team that develops the software is different from the team that operates it, then there’s a misalignment of incentives. If I am a developer being tracked (and promoted) solely on the amount of code I ship, my focus will be on getting more bits out the door and not on ensuring the features I release will withstand the burdens of operation. That’s another team’s concern.

    That’s the motivation behind the proliferation of the “you build it, you own it” mindset at top-performing organizations. Hell, my first day at Amazon, they tossed me a pager and said “Good luck.” And while that may sound daunting (and it was), I can tell you that it not only motivated me to make sure I built systems to last, but it also fundamentally changed the way I thought about architecting systems. Spoiler: I’m not a big fan of my pager going off at 3 in the morning.

    In other words, the person or team building the system needs to be the same person or team that feels the pain if that system is failing. But it’s not just about punishment and pain, that team also needs to be recognized and rewarded when their system is running reliably. This creates an alignment of incentives that promotes the kind of habits seen across top-performing teams.

    Track the Four Golden Signals

    In monitoring distributed systems, Google’s SRE book outlines the four golden signals of monitoring as latency, traffic, errors and saturation. If I could wave a magic wand and immediately improve the culture of an organization, then I would have service-level objectives (SLOs) tied to these four metrics. But going back to habit formation: If you make acquiring a habit too difficult upfront, then an engineering team will ultimately reject it. So if your organization isn’t mature enough to create SLOs, simply beginning to track these metrics will up-level your game tremendously. They will give you an understanding of what is not working well and help guide your priorities.

    To Summarize

    Imagine a world where, after a major incident happens, the points of failure involved become DNR work. The same failure is not allowed to happen again and no new feature work will be completed until the fixes are implemented. And more importantly, no new feature work will be completed until those fixes are verified via a chaos engineering experiment, which is then cataloged and run continuously against your system. Then you take that knowledge and share it with other teams so they can run the same experiments and make sure they are immune to those failures as well.

    This is how you build a culture of resilience.

  • The SRE Pressure Cooker: Balancing Velocity Against Risk

    The SRE Pressure Cooker: Balancing Velocity Against Risk

    Delivering fast, reliable digital services today is a lot like Olympian alpine skiing. These services must deftly maneuver a series of perilous passages en route to end users, all while maintaining the astounding speed we now take for granted.

    In an SRE’s world, those passages are today’s increasingly complex and interconnected internet infrastructure through which they must deliver and support cutting-edge application functionality–all while rolling out updates at a relentless pace. It may appear easy to end users, but SREs are working incredibly hard behind the scenes, and as the internet grows more intertwined and treacherous, the risk for catastrophe increases.

    Judging from the results of our second annual SRE survey, striking this critical balance between supporting development velocity while maintaining site reliability is an onerous and stressful task. Here’s a look at some of the key findings, and takeaways for organizations increasingly reliant on this emerging role.

    The SRE Role Is Relatively New

    Even though Google coined the term in 2003, 64% of respondents noted their SRE disciplines have only been in existence for three years or less. This means many SREs and SRE teams—those professionals dually knowledgeable in software development and adapting IT systems to meet the needs of particular software—are still adjusting to their roles and responsibility. Many SRE teams are also small relative to the size and scope of the infrastructure to be managed, and they remain a scarce resource, with many organizations highlighting skills shortages in this area.

    IT Incidents Are the “New Normal”

    Almost half (49%) of survey respondents reported they had worked on a service incident over the course of the past week. In the month of June alone, Google—which is known for having a very mature SRE practice—had two significant outages: One that hit YouTube, G Suite, and several popular third-party apps relying on Google Cloud for their back-ends (including Discord and Snapchat), and the Google Calendar outage just a few weeks later.

    Catchpoint has detected a significant uptick in outages in recent months, likely the result of growing internet complexity which naturally increases the potential for problems impacting digital service reliability. This is a challenge for all companies and their SRE teams, including the biggest and the strongest—the last few weeks have shown us no one is immune. Organizations can help their SREs maintain a healthy perspective by bearing this in mind.

    Many SREs Don’t Have Clearly Defined Service Level Objectives (SLOs)

    This is problematic, because without SLOs, it is nearly impossible to identify what is an incident in the first place. Twenty-seven percent of SREs reported they don’t have any SLOs, and when flags are raised on everything, this leads to more alerts, a greater number of false positives, and naturally, greater SRE fatigue and stress. Our own conversations with SREs have revealed that even among those with SLOs in place, these are often unrealistic or not correlated to end-user satisfaction and happiness—for example, targeted at 100% availability (impossible perfection) when 98% availability may suit end users just fine. If you’re striving for perfection, every little thing that goes wrong constitutes an emergency, even if customers don’t notice. Setting unrealistic targets and not letting customer satisfaction define SLAs is a recipe for SRE burnout.

    Many SREs Feel Burnt Out and Stressed

    According to our survey, 21% of respondents indicated they never experience post-incident stress. However, it was interesting to note the answers to the following question, where we asked what types of changes SREs notice after an incident—in concentration, ability to sleep, mood and more. One would suspect if an SRE never experienced post-incident stress, he/she would select “none” for the second question—but that wasn’t the case. Rather, one third of SREs said they have noticed such symptoms. Sixty-seven percent of SREs who reported feeling stress after each incident, indicated they don’t believe their company cares about their well-being.

    We are aware of other surveys showing SREs are quite happy, particularly given their impressive compensation. While we don’t believe being an SRE automatically makes someone stressed and unhappy, organizations should never assume these professionals know and understand how appreciated they are. There is often a reluctance to discuss topics such as job-related stress, as evidenced by the fact that 50% of the SREs we targeted opted not to complete our survey—perhaps a result of SREs’ predominant hero culture. In addition to compensation, organizations should consider other ways to show SREs—whether they vocalize their stress levels or not—they care about their well-being, such as offering an extra vacation day or other perks after a particularly trying incident.

    SREs Could Benefit from Automation, but Don’t Have Enough

    Another thing we consistently heard in our conversations with SREs is they feel so busy putting out fires today they don’t get a chance to create a better tomorrow—in other words, they lack the opportunity to be proactive and preemptive in their approaches. A key factor in this is the amount of toil SREs face daily—meaning manual, repetitive, automatable, tactical work. Our survey found 59% of SREs believe there is too much of it in their jobs, and too little automation. Nobody strongly agreed with the statement, “We have used automation to reduce toil,” while almost half disagreed or strongly disagreed.

    Considering SRE teams are often small and the talent pool tight, it becomes critical for organizations to automate as much work as possible, including areas such as application performance monitoring. The SRE team at Alaska Airlines recently automated performance monitoring and deep-dive diagnostics, and as a result has been able to reduce their alert noise by 92%. Ultimately, this has freed up the team to focus on real issues versus wasting time investigating false positives. This has helped them bring down the average mean time to detection for real issues from hours to less than 10 minutes.

    Conclusion

    The stage has been set for SREs to experience significant stress, at a time when organizations need their exceedingly rare skillsets the most. As internet complexity increases, it is critical for organizations to continually nurture, encourage and motivate these valuable team members. This means understanding and empathizing with the intense pressures they’re under; having SLOs well-defined, realistic and meaningful; leveraging soft skills to communicate effectively and non-monetary incentives to show appreciation; and providing them with needed automation so they can focus on proactive improvements that advance the business.

    — Mehdi Daoudi