Tag: application availability

  • Defining Availability, Maintainability and Reliability in SRE

    Defining Availability, Maintainability and Reliability in SRE

    In the world of reliability engineering, you’ll frequently encounter the three “-ability” words: Availability, maintainability and reliability. They sound similar and have similar meanings. In fact, these words may seem so similar that it can be tempting to use them interchangeably.

    That would be a mistake. Availability, maintainability and reliability all have distinct—if related—meanings, and they each play different roles in reliability operations.

    Definitions of Availability, Maintainability and Reliability

    Let’s start by succinctly defining each of the terms.

    Availability, Defined

    Availability is the extent to which an IT resource is ready to perform a task when requested.

    Thus, an application or server that is responding to requests is available (even if it takes longer to respond than desired or otherwise operates suboptimally, in which case it may have a reliability issue but not an availability issue).

    Availability is usually measured as a percentage. A resource that has 99% availability, for example, is one that is up and responding 99% of the time.

    Maintainability, Defined

    Maintainability is a measure of how quickly and easily a resource can be fixed when something goes wrong.

    If a buggy application release can be quickly fixed by rolling back to a stable version, the application would have a high degree of maintainability. On the other hand, if you have a server that needs to be rebuilt manually after it fails, it’s not very maintainable.

    Reliability, Defined

    Reliability is the extent to which a resource functions as required upon request (as opposed to simply being available).

    For example, if your users require an application to process each transaction within one second, and it does this 99%of the time while also maintaining a 99% availability level, that would be a relatively reliable application. In contrast, an application that responds to almost all requests but that suffers from high latency or error rates would not be very reliable, although it might be highly available.

    Differences Between Availability, Maintainability and Reliability

    The simplest way to spell out the differences between availability, maintainability and reliability is to highlight what’s unique about each concept.

    Availability

    Unlike maintainability and reliability, availability is an essentially binary metric, in the sense that a system is either available or it’s not. Although availability status can change over time, there is no such thing as varying degrees of availability.

    Availability also stands out because it is often the most important metric in defining SLAs and SLOs. Contracts typically specify that resources will achieve a certain level of availability, defined in percentage terms. They may sometimes also specify metrics like response times, which are a reflection of reliability, but availability is more likely to be the most important metric within service agreements.

    Maintainability

    Maintainability is unique in that it’s a pretty subjective concept. A system that one SRE considers easy to maintain could seem difficult to maintain to another SRE. The methodologies that engineers leverage to optimize maintainability could vary, too; for instance, someone from a DevOps background may be more likely to focus on optimizing maintainability within the software delivery chain than someone trained in classic site reliability engineering, which focuses on automating IT operations through code more than on the software delivery process.

    That said, automation tools help virtually every team maximize maintainability, no matter their preferences or background.

    Reliability

    Reliability stands out because it reflects how well a system performs. Here again, this is a somewhat subjective metric, because performance requirements may vary from one user to another. Nonetheless, reliability is the most effective means of measuring whether a system meets the performance levels it needs to, even if those requirements change over time.

    Why do availability, maintainability and reliability matter?

    Because availability, maintainability and reliability each measure different aspects of a system’s status, putting them together is a useful means of gaining insight into the overall reliability of a system.

    In order to be reliable, a system requires both availability and maintainability. A system can’t be reliable if it’s not available. It’s also unlikely to be highly reliable if it takes a long time to fix issues due to low maintainability.

    At the same time, by focusing on availability, maintainability and reliability individually, you can drill down into specific issues within the IT resources you manage. For instance, you might have systems that are high-performing when they’re available, but that have low reliability rates because of availability issues. In that case, you’ll know that an investment in increased availability is likely to yield the greatest reward for increasing overall reliability.

    The Bottom Line

    The bottom line: While availability, maintainability and reliability are all reflections of the quality of an IT resource, they measure quality in different ways. Tracking each category separately is important for ensuring that you know where your weakest links exist within overall system performance and health.

    At the same time, however, it’s important to compare and correlate availability, maintainability and reliability data so that you can achieve continuous insight into a resource’s status. When you can track each item separately but also combine them together to gain a complete picture of your system, you are in the best position to optimize reliability operations.

  • But What About When You Fail?

    But What About When You Fail?

    One thing Agile and DevOps definitely brought IT was a more accepting view of the whole “mistakes happen” mantra. Crashing systems is no longer a guarantee of a free ticket out the door. Indeed, off the top of my head, I can think of at least two cases where a CIO themselves checked in sloppy code and it trashed the system—CIOs at orgs big enough that they probably should not have been coding in the first place.

    And that brings us to today’s blog. The standard response to “But what if there is a bad update?” is “We’ll just roll out a new version!” That works in some instances. In a lot of instances—most that involve GitOps as part of DevOps—it doesn’t. Implosions can be huge and take out chunks of infrastructure. So you need a better plan than just assuming you can fix it with a forward update. Do you have rollback capability? Are you making use of it, assuring all is set to roll back if needed? Testing to make sure rollback works with the other changes made across the system in this update?

    That is part of the problem. Fixing with a new update is the best option, simply because in the age of massively distributed, microservices-based solutions, rolling back comes with a ton of baggage—enough that it may not be viable for you. Okay. So you can’t quickly roll forward. You can’t quickly roll back. Quick! What do you do?!

    And I don’t have the answer to that question, because I’m not in your organization working on your systems. At some point, all IT is personal. And this is one of those points. You need to know what your best options are if an update spirals everything, and you need to have a plan to implement it. But what your best options are is not going to be the same as the next org. This is that point.

    One option is to keep the ability to build the entire system. That’s a large setup, but every minute systems are down is hurting the company. And that’s what you need to plan for. Think of it as disaster planning for DevOps. A man-made disaster that destroys systems and infrastructure.

    So, like any other disaster planning scenario, walk through the chain that makes the system work, identify weaknesses and list ways to address them. Then test those to make certain they do what you hope they will. Then, set up an automated system to keep this whole plan up-to-date.

    In short, we’re in a new automated landscape, and we need a new automated tool to pull our rears out of the fire when the inevitable happens. Be it from dev error or malicious attacker, it is a safe bet that, sooner or later, you will have a massive systems outage that your DevOps toolchain can’t adequately address. Know what you are going to do. Or at least have thought about it, so you’re not just starting to think it through while coworkers and customers can’t access systems.

    And keep rocking it. This is just another layer of protection for all the hard work you’re doing. Take the extra step. Like insurance, if you ever need it, you will absolutely be glad you did.

  • Achieving Application Health Through Integrated APM

    Achieving Application Health Through Integrated APM

    Performance and availability mean everything to application users at both the enterprise and consumer level. Think about how much we rely on technology for our everyday communication—when Slack suffered from an outage this summer, it affected thousands of business professionals across the globe who use Slack to communicate, organize tasks and share information. In a sense, we are only as efficient and effective as the technology we use.

    With this in mind, it’s clear achieving and maintaining application health is paramount to overall business success. In part one of this series, we explored how creating and maintaining a healthy app involves not only ensuring it’s performing well via lifecycle APM, but also making sure it’s available when needed—making a log management solution that monitors availability, enables proactivity and reacts quickly to overcome problems is crucial.

    In order to help ensure a healthy application, we must first consider a few defining characteristics.

    Understanding a Healthy Application

    Even for humans, health is more than just being alive—if we’re not performing optimally, we’re not having the best experience possible. Optimum health for an app or website doesn’t mean it’s just up, but that it’s also performing in the way the market demands. For this, organizations need to implement an integrated APM solution to ensure top performance and continuous availability.

    However, achieving application health is easier said than done. Issues that arise for teams working in silos can lead to digital war rooms and blame storming. Organizational structure can get in the way of disparate teams working together to achieve performance and availability. How can organizations overcome this issue?

    Optimizing Performance and Availability

    As we know, each piece of the APM puzzle must be incorporated to achieve full-stack visibility and ultimately top performance: user experience, metrics, traces and logs. Following a few key best practices can help achieve overall application health.

    • Early Adoption: Implement application performance and logging tooling as early in the application development lifecycle as possible. Without integrating these tools from the beginning, organizations run the risk of having errors in availability and performance when they deploy the app. This is because one of the key benefits of a full-stack APM solution is the ability to analyze the performance characteristics of the application code and logic itself. If an application isn’t properly written, it can be so inefficient that no amount of resource in production can assure appropriate service delivery.
      Another critical benefit of early lifecycle APM adoption is the knowledge it delivers about app performance can be reviewed with business stakeholders to ensure complete alignment against shared goals—before going live in full production.
    • Integrated Solution: As mentioned above, the biggest challenge to achieving availability and performance is typically organizational structure. For most organizations of a reasonable size, separate teams monitor performance and availability; in some cases there may even be multiple teams monitoring performance and availability, specifically aligned with the individual technologies (e.g. compute, storage, network, web, cloud, etc.) underpinning each application. To get around the obstacles this creates, it’s key to have an integrated solution where every team is looking at the same set of information—the so called single pane of glass. The ability to share information proactively with management across teams is paramount to overall success and avoiding digital war rooms, or at least minimizing time spent in them.
    • Implement High-Value and Appropriately Priced Solutions: Powerful, easy-to-use solutions that will achieve integrated application performance and availability management for any type of organization operating either on-premises and/or cloud-based applications are key. Traditional APM solutions were either so complex to deploy or so expensive that ubiquitous adoption was not operationally or fiscally possible—depriving much of the organization and its applications the benefits of APM by restricting its application to only the most important subset. As a result, quick-to-value and affordable solutions are needed to enable tech pros who may be working in the organizational performance silo to get running quickly and show success to the organization.

    Conclusion

    Although apps serve their own unique purpose—whether that’s for consumer use enabling us to watch TV and movies on-the-go or at the business level to streamline work and communication—they’re all connected by a universal need: performance and availability. Companies must therefore prioritize application health, and in turn creating and delivering an exceptional user experience.

    — David Wagner

  • Understanding and Implementing the APM Puzzle – Performance and Availability

    Understanding and Implementing the APM Puzzle – Performance and Availability

    While every organization is unique, there’s one universal goal everyone can get behind—creating and delivering an exceptional end-user experience. For both classic IT organizations and companies built in the age of cloud, delivering continuous availability and appropriate application performance is key.

    Although digital transformation has become a buzzword some would say is meaningless, the idea of digitizing a traditional company and building new businesses whose foundations are established upon the latest web technologies is critical in today’s IT and non-IT organizations. Think about the prevalence of cloud migration efforts during the past several years. Traditional organizations typically prioritize on-premises applications with increasing quantities of hybrid applications. But, because they aren’t yet fully transitioned to cloud—whether IaaS, PaaS or SaaS—their IT organizations need to live in both universes.

    Monitoring these environments at both the infrastructure and application tiers and across each stage in this transition requires support of a full stack of tools to cover web uptime, user experience, transaction monitoring, application tracing, log analysis and appropriate infrastructure monitoring. That said, modern organizations that have fully embraced cloud and leverage modern microservice architectures have similar challenges, but the exact same level of monitoring visibility requirement.

    Whether public, private, hybrid or multi-cloud, nearly every organization has moved to the cloud in some capacity. Regardless of what the company produces or the services IT provides, it must maintain uptime, availability and seamless performance to create and continuously deliver a user experience rivaling competitors and exceeding user expectations.

    Because let’s face it, even though every business is different and virtually every market is saturated, consumers and B2B companies have more options than ever. Without delivering a solid end-user experience, there’s an ever-increasing plethora of alternative vendors and providers from which customers can choose.

    Public Cloud Platforms for Classic IT Organizations and the Cloud Generation

    Let’s explore a few differences between how traditional IT organizations and cloud-era companies interact with the cloud. At a high level, Azure is a natural choice for IT companies with significant investments in apps traditionally acquired (not built) by the company and managed by their IT team, that are now tasked with digitizing their business by leveraging public cloud platforms. These types of companies often have Microsoft Office, Microsoft Exchange and other Microsoft products already implemented on-premises, and Infrastructure-as-a-Service (IaaS) or Platform-as-a-Service (PaaS) solutions as well. If they’re now adding a public cloud platform to their mix, staff expertise and solution familiarity make adopting Microsoft Azure a natural next step. Microsoft has long recognized the need to make these types of transitions and mixtures easier, demonstrated by its Office365 suite (as an example).

    On the flip side, companies that were born fully in the cloud, (i.e., relatively newer organizations without a legacy anything on-premises, with respect to IT infrastructure) will likely have a higher propensity to choose a provider such as AWS or Google Cloud even though Microsoft, Oracle and others are certainly competing there as well. So, it’s not a matter of one being intrinsically better or more fully-featured, but about which is the best fit for the organization and its priorities.

    At the end of the day, however—regardless of organization type or public cloud platform implemented—every company needs monitoring tools and solutions that keep applications up, running and performing seamlessly. With an integrated set of simple, easy-to-use, quick to value and powerful solutions, organizations can achieve truly proactive lifecycle assurance of application availability and performance.

    Uptime and Seamless Performance: Putting Together the APM Puzzle Using All the Pieces

    Each and every piece of the APM puzzle must be incorporated to achieve full-stack visibility and ultimately top performance: user experience, metrics, traces and logs. Having a log management solution monitoring availability, enables proactivity and enables IT to react quickly to help identify and remediate problems is paramount—it’s table stakes that everything be available. After all, if a website or page isn’t accessible, business transactions cannot be processed and that business is effectively down.

    Creating and maintaining a healthy application involves not only ensuring it’s functionally available to users; it must also be performing appropriately—continuously delivering both a reasonably responsive user experience and supporting the required volume of transactional work. This means a complete view into all aspects of performance—from the website and webpages all the way into and across all relevant application components and their associated infrastructure—be proactively visible and monitored.

    There is thus a natural and increasingly urgent workflow affinity between logs and APM—across the development lifecycle. When errors occur, it’s critical to understand, precisely and immediately, where are the impacts on service performance, or if there even are any impacts. Conversely, when application performance issues arise, it’s critical to immediately know exactly which errors—across the complex and constantly changing infrastructure and application stack—are directly related to the problem at hand.

    Without this bi-directional workflow integration, not only can application performance and availability not be managed proactively, but any issues with application service availability and performance will take far longer to diagnose and repair—seriously impacting the business bottom line.

    Just because everything is available doesn’t necessarily mean it’s appropriately performing for the end user trying to use the application.

    With all this in mind, it’s paramount to implement powerful, easy-to-use solutions that’ll achieve integrated application performance and availability management for any type of organization operating on-premises and/or cloud-based applications across all cloud environments.

    Conclusion

    While traditional IT organizations and cloud-only companies are different in many ways, providing exceptional end-user experience and achieving seamless availability and performance (no matter what) is a common and increasingly critical business goal. To achieve business success, organizations should choose a complete suite of simple, powerful and integrated solutions to optimize and manage their cloud-hosted applications and platforms. Anything less leaves the APM puzzle unsolved.

    — David Wagner