Tag: performance management

  • Upgrade SRE Performance Management With AI

    Upgrade SRE Performance Management With AI

    In my recent blog, Revolutionizing the Nine Pillars of SRE with AI-Engineered Tools, I indicated that AI could analyze application and infrastructure performance data to identify optimization opportunities and predict capacity requirements. In this blog, I explain in more detail how AI-engineered tools can be used to improve the pillar that I call performance management of apps and infrastructure.

    AI Use Cases for Performance Management

    Performance management in the SRE context generally refers to monitoring, analyzing and optimizing the performance of IT systems to ensure they are functioning efficiently and effectively to meet the needs of users and the business.

    Performance Monitoring and Analysis: With the complexity and scale of modern IT systems, manually sifting through performance data to find issues can be like finding a needle in a haystack. AI can help identify patterns and anomalies that might indicate a performance problem. Tools like Dynatrace and Datadog use AI to analyze performance data and automatically alert you to potential issues.

    Capacity Planning: Predicting future capacity needs is often more of an art than a science, but AI can help make these predictions more accurate. Machine learning algorithms can analyze historical usage data and predict future capacity needs based on patterns and trends in this data. A tool like Amazon Forecast can generate accurate capacity forecasts based on time-series data.

    Load Testing and Stress Testing: AI can enhance these testing techniques by dynamically generating realistic load and stress scenarios based on real-world usage patterns. AI can also analyze the results of these tests to identify potential bottlenecks. Tools like BlazeMeter with Taurus provide capabilities for AI-enhanced load and stress testing.

    Performance Optimization: AI can identify optimization opportunities that humans might miss. Machine learning algorithms can analyze performance data to identify inefficient operations or configurations that are impacting performance. Tools like Akamas leverage AI to automate the performance optimization process.

    Performance Modeling: Traditional performance modeling techniques can struggle to accurately model the behavior of complex, distributed systems. AI can analyze historical performance data to create more accurate models of system behavior. Tools like BMC’s TrueSight Capacity Optimization employ AI to enhance performance modeling capabilities.

    Overcoming Challenges

    Implementing AI in any area brings its own set of challenges, and performance management is no exception. Here are some potential challenges and ways to overcome them:

    Data Quality and Management: AI relies heavily on data to provide accurate predictions and optimizations. If the data is inaccurate, incomplete or not representative of the system’s behavior, the AI’s effectiveness will be limited. To overcome this, it’s important to implement robust data management and governance practices. Regularly review and clean your data to ensure its accurate and relevant.

    Skills Gap: Implementing AI often requires a specific set of skills in data science, machine learning and AI, which may not be present in existing SRE teams. Investing in training and upskilling for existing team members, as well as considering hiring or partnering with AI experts, can help fill this gap.

    Resistance to Change: As with any major change, there can be resistance from team members who are comfortable with existing practices. Open communication about the benefits of AI and how it can make their jobs easier and providing the necessary training can help overcome this resistance.

    Integration with Existing Systems: AI tools need to be able to integrate with existing systems to access necessary data and take actions based on its analysis. Choosing AI tools with robust integration capabilities and planning for potential integration challenges can help overcome this issue.

    Cost: Implementing AI can require significant investment in tools, infrastructure and skills. Careful planning and budgeting as well as calculating potential ROI from the improved performance management can help justify these costs.

    Security and Privacy: AI tools need access to a lot of data, which can raise security and privacy concerns. It’s important to choose AI tools that follow best practices for data security and privacy and to conduct regular security audits.

    Summary

    Performance management is a critical part of site reliability engineering, ensuring that systems function at their optimal efficiency and efficacy. With the integration of AI, this facet of SRE is getting an unprecedented boost. From using AI in performance monitoring to identify patterns and anomalies to employing machine learning for accurate capacity predictions, AI is redefining how SRE practices are carried out. Tools like Dynatrace, Amazon Forecast and Akamas, among others, are leading the charge in AI-fueled performance management, streamlining processes and delivering superior results.

    Yet, as promising as the future of AI in performance management appears, organizations must navigate challenges to fully harness its potential. From managing data quality to bridging the skills gap, countering resistance to change, integrating with existing systems, managing costs and ensuring security and privacy, the transformation involves overcoming significant hurdles.

    However, with robust data governance, investment in upskilling, open communication, careful planning and strict adherence to security best practices, organizations can fully leverage AI to revolutionize their performance management practices and elevate their SRE capabilities. This not only ensures optimal system performance but also positions the organization well for future growth and innovation.

  • Of Max and Min: When Performance Engineering Plans Go Awry

    Of Max and Min: When Performance Engineering Plans Go Awry

    Modern software developers have access to powerful tools and services which allow them to quickly develop, demo and deploy fully functional applications. But what happens when the single-user prototype satisfies all required functionality but the user says the system seems slow? Or if the initial implementation turns out fine for one user, but multiple users start to experience delays? Or if users are happy, but the cost of auto-scaling grows prohibitive?

    The Problem with Performance Engineering

    Concern over non-functional aspects of computing systems—response times, resource utilization and costs—falls under the domain of computer systems performance engineering (CSPE). Unfortunately, the practice of CSPE has evolved into something more akin to art than engineering—with few standardized principles and practices. Because software is constantly evolving its languages and abstractions, CSPE practitioners appear obligated to evolve and change their abstractions and languages, too.1 As a result, part of our jobs as performance engineers is to understand what the heck other performance engineers really intend or mean.

    There are many reasons for this state of affairs, but ultimately this lack of a common language and engineering standards for CSPE is a deficiency in the education and training of software engineers. There is very little in the typical university software engineering curriculum that prepares software engineers for doing CSPE. This sometimes leads to practices that are not backed by scientific or engineering principles. As a result, practitioners are forced to “wing it”—creating their own art or hitching their wagons to someone else’s star and learning (or not) through trial and error.


    The PMWG – A POSIX for Performance?

    The Performance Management Working Group (PMWG) was started in the late 1980s during the “UNIX wars,” when computer vendors were feuding over claims to the UNIX leadership mantle and had formed two factions: UNIX International and the Open Software Foundation. The group was unique in that its members included computer systems vendors on both sides of the “war.” Its members put aside those differences to try to come up with performance management standards and practices that would put UNIX on par (at least) with the performance management tools found on mainframes.

    In retrospect, the PMWG was probably the best chance UNIX had for establishing, promoting and enshrining some basic performance engineering principles—a kind of POSIX for performance. Unfortunately, the group failed—in large part, in my opinion, because it put too much priority on the latter stages of the performance management data pipeline and did not spend enough effort toward establishing key primitives. Today, in the absence of established rules and practices for performance engineering, history is repeating itself. Current efforts in observability continue to focus mainly on data presentation and data transport and storage, without much focus on the right metrics and data quality based on basic performance engineering principles. Presentation, transport and storage are important, but the efforts around metrics, logging and tracing coverage and data quality are sporadic at best.


    Inertia, Expediency and Superstition

    Why are more formal, engineering-based practices around CSPE needed? Consider the use of CPU utilization as a key performance indicator (KPI) of system performance. I’m a member of Bloomberg’s Trading Solutions SRE team, which manages hundreds of machines on which the company’s Trading Solutions software runs. We often see or hear statements like “This host’s CPU is too busy,” or “That host is out of CPU resources.” I’ve found that these observations are not actually very helpful and sometimes divert our attention from the real problem(s). When misused as a system KPI, focusing solely on CPU utilization can lead us down the wrong path and draw incorrect conclusions. To understand why, let’s first look at how we assess the performance of real-world, non-computing systems.

    Take fast-food restaurants. One of the most important attributes of fast-food restaurants (besides their health inspection rating) is that they are fast. But we’ve all been to fast food restaurants that are anything but fast. Without actually standing in line and waiting, we can often tell that we will have a slow experience at a fast food joint if there is a long line. The line length gives us a lot of information. Our assessment of the situation is different whether there are 100 customers versus 10 customers versus a single customer waiting in line.

    In the real world, this line length (or queue length) assessment is almost universal. Supermarket checkouts, ATMs, airport security screening, toll booths, elevators and COVID-19 testing are all areas where we apply this semi-conscious assessment. In contrast, we almost never use “cashier utilization” or “self-serve soda fountain utilization” as a way of assessing how slow a fast-food restaurant is. We always say “the lines are long” or, if we actually choose to wait in line, we’ll make the more direct statement that “service is slow.”

    Which brings us back to our computing systems. Why do we reflexively look to CPU utilization when computing systems are slow while, in the physical world, we rely on queue lengths? The question is even more relevant when we consider that the concept of CPU is less well-defined today. Are we measuring the CPU, its cores or hardware threads?2 I believe the contributing factors to this misuse include inertia, expediency and superstition.

    In the early days of computing, CPUs were perhaps the most expensive component of the large monolithic computer systems of the time. Measuring how much these expensive components were being used was an important part of maximizing the financial investment in these large machines. Fortunately, measures of CPU utilization were relatively easy to estimate. Ease of implementation meant that CPU utilization measures were almost universally available.

    The universal availability of this simple metric, combined with some correlation (sometimes weak) between high CPU utilization and system slowness made it a default indicator of systems performance. When a belief becomes ingrained to the point of superstition, even weak correlations are enough to validate and perpetuate it. Better metrics—for example, run queue lengths at the CPU(s)3—can help us shed our superstitions to better understand our systems and lead us in the right direction in identifying and correcting problems.

    Dispelling Superstition and Other Irrational Practices

    Which brings us to the title and purpose of this series. As a title, “Of Max and Min” is meant to evoke John Steinbeck’s classic novella “Of Mice and Men,” which in turn got its title from the famous line in Robert Burns’ poem “To a Mouse”: The best-laid plans of mice and men oft go awry. The two main characters in “Of Mice and Men” fail in their efforts to better their lives during the Great Depression because they are tragically ill-equipped to do so. Similarly, we should not be surprised if plans go awry when software engineers are tasked with doing performance engineering without the proper training in performance engineering principles.

    “Max and Min” also highlights the importance of data analysis and statistics in understanding systems behavior. Effective CSPE requires us to be familiar with concepts in statistics, numerical analysis, and even operations research—in addition to the more standard computer science areas of algorithmic complexity and computer architecture.

    With “Of Max and Min,” in addition to calling out superstitions, I will be pointing out some common mistakes we make and blind spots that we have.4 Through case studies, I’ll illustrate how our blind spots can lead to mistakes that can persist for years. I will also try to cover some basic CSPE principles that I believe are skipped in most software engineering training. Like everyone else, I’ve been winging it for 40+ years. And I’m still learning. If you disagree with any of my points or have further insights into the topics that I am presenting, please let me know.


    1But, as in the application of the principles behind algorithmic complexity, there are some performance engineering principles that can and should be adhered to regardless of the system and language du jour.

    2There are in fact some bizarre implementations of “CPU utilization”. Here’s an example of how AIX “broke” the meaning of CPU utilization on their multiprocessor systems running in hyper-threading mode.

    3UNIX and Linux do in fact have metrics that could be used as CPU queue length indicators.  In a future installment, I will go into the quirks of some of the CPU queue length implementations and further argue for more consistent mathematical bases for queue length metric implementations.

    4For example, have you ever wondered why, after so many outages caused by logging, malloc, DNS and other low-level services, we still don’t have good, out-of-the-box visibility into these and other low-level components? We are accepting key blind spots as a fact of life when we really shouldn’t.

    About Me

    Like many software engineers, I started off studying a different field (Physics). Unlike most software engineers, I’ve always wanted to be a performance engineer. When I started taking computer classes as an undergraduate in Columbia College’s Physics program, I did not wish for a career writing programs in Fortran (yes, it was that long ago). But I did envision myself making systems run better and faster.

    To that end, when I added a Computer Science major to my undergraduate studies, I also added a minor in Industrial Engineering and Operations Research. For those unfamiliar with IEOR, it is a multidisciplinary engineering field that leverages mathematical and analytical methods to address complex system problems like resource allocation, supply chains, wait times, fairness, etc. Tools and methodologies from IEOR are invaluable to CSPE.

    At the start of my career, I joined the Bell Labs Computer Center as a computer systems performance engineer, where I made the somewhat contrarian decision (for 1980) to work in the group supporting their in-house, low-key UNIX systems (as opposed to their commercially popular MVS mainframes). I didn’t realize it at the time, but having access to the source code of the system you support/study can be a boon to understanding. At Bell Labs, I had the opportunity to contribute to the UNIX System V kernel, where I co-developed the first general-purpose kernel and user-land tracing system for UNIX and made significant performance enhancements to the virtual memory subsystem, and participated in the PMWG.

    Since then, I’ve spent the bulk of my professional career working on financial software systems—ranging from market data to messaging to transaction management. I’ve worked on monitoring systems on an enterprise level. I’ve made and witnessed many CSPE mistakes. In my current role as an SRE on the Trading Solutions team at Bloomberg, I’m seeing and helping remedy lots of operating system issues that stem from our runtime scale.


    Thanks

    Thanks to everyone who reviewed this document and provided constructive feedback. Special thanks to Nate McNamara and Peter Wainwright for suggesting substantive organizational and structural improvements. Peter also reined in my default stream of consciousness, run-on style. None of them had any hand in the writing of this paragraph.

     

     

  • Riverbed Aims to Unify App, Network Performance Management

    Riverbed Aims to Unify App, Network Performance Management

    Much of the initial focus on DevOps has been on accelerating the rate at which applications are built and deployed. But once those applications are deployed, it’s not too long before the focus shifts to the quality of the end-user experience. Aiming to make it easier to determine what the quality of that experience really is, Riverbed Technology has launched a Digital Experience Management initiative based on the SteelCentral management platform.

    Erik Hille, director of product marketing for Riverbed, says the latest release of SteelCentral provides tighter integration between SteelCentral Portal and the SteelCentral Aternity end-user monitoring tool and SteelCentral AppInternals, an application performance management (APM) platform. The goal, says Hille, is to make it easier for an organization to holistically manage individual end-user experience by correlating information at the device, network and application levels. As part of that effort, Riverbed also has integrated the workflow between SteelCentral Aternity and AppInternals to create an integrated monitoring system.

    Hille says IT operations teams can use a comprehensive approach to managing digital experiences that can be set up in fewer than 10 days. Once in place, IT operations teams can more quickly identify the root cause of any issue by correlating application, device and networking data through a single pane of glass. In addition, Riverbed is employing REST APIs to make it possible to consume alerts generated by collaboration tools such as Slack and HipChat. IT operations teams can not only automatically open incident management tickets, but also collect additional metrics via those APIs.

    Riverbed is also providing integration between NetProfiler, a network performance monitoring tool, and NetIM, a network device management tool, to provide more visibility into how network infrastructure impacts network performance.

    Hille notes that most IT operations teams are swiveling between different sets of tools to correlate the same information. The problem is that as more organizations strive to become a digital business, any performance issue has a direct impact on customer satisfaction and the amount of revenue being generated. Faced with that increased business pressure, Hille says IT organizations need to re-evaluate legacy approaches to IT operations based on silos. The goal should be to implement a more proactive approach to IT focused on proactively preventing issues from occurring or escalating, versus merely reacting to isolated events as they occur.

    One of the few things in DevOps that developers and IT professionals can agree on is that whenever there’s an issue, it’s the fault of the network guy. That’s not always true, but it does take networking professionals a long time to prove their innocence. Until then, there’s a tendency to assume their guilty until proven otherwise. Obviously, that kind of bias isn’t in keeping with the whole spirit of integrated DevOps. A more comprehensive approach to application and networking performance management should go a long way toward reducing the amount of finger-pointing inside any organization.

    There will never be a perfect IT environment. The goal is to discover issues before the end-user notices—and, failing that, resolve that issue in as little time as possible. Any argument about what exactly happened to cause the problem can occur at a later date and time.

    — Mike Vizard