Category: Application Performance Management/Monitoring

  • Time-Series Database Basics

    Time-Series Database Basics

    There’s been a lot of buzz about time-series data for the past few years. No matter what you’re monitoring—financial data, server performance, social media activity or something else entirely—time-series data can give you valuable insights into trends and patterns.

    But what exactly is a time-series database, and why do you need one? This article will look closely at time-series databases and how they can help you make sense of your data.

    What is a Time-Series Database?

    A time-series database (TSDB) is a database management system optimized for storing and querying data that changes over time. Time-series data is often used in monitoring applications where it’s important to quickly retrieve information about a system’s current state and trends and patterns over time.

    In most definitions, time-series data, often referred to as time-stamped data, is a sequence of values indexed in time order. Time stamping refers to data collected at various times, where each value is time stamped. These data points are usually gathered from the same source and are used to measure progress over time.

    Important Notes:

    You can use relational or NoSQL databases to crunch time-series data, but purpose-built time-series databases are tailored to exploit the unique features of time-series data. This implies that time-series databases ingest at a faster rate, query more quickly and compress data more efficiently. Furthermore, time-series databases include special analytical capabilities and management features that are not found in most relational or NoSQL databases.

    Some Advantages

    This list is not exhaustive, but here are some of the advantages that you might get from using a time-series database:

    1.) Time-series databases can handle high-velocity data very well.

    2.) Time-series databases are purpose-built for storing time-series data, making them more efficient in storage and querying.

    3.) Time-series databases often have built-in analytics and management features designed specifically for time-series data.

    4.) Many time-series databases are open source, which means they’re free to use.

    Some Disadvantages

    There are also some disadvantages that you should be aware of before using a time-series database:

    1.) Time-series databases are often more complex to set up and manage than relational or NoSQL databases.

    2.) There are relatively few time-series databases to choose from, so you might not have as much flexibility in choosing a platform that meets your specific needs.

    3.) Time-series data can be very large, so you’ll need enough storage capacity to accommodate your data.

    What Are the Characteristics of Time-Series Data?

    There are several characteristics of time-series data that make it unique and require special handling. From a database standpoint, the most essential ones are as follows:

    1. Timestamp: Every data point has its timestamp. The timestamp is crucial for calculating or analyzing the information.
    2. Structure: Metrics from devices or monitoring are almost always structured, unlike those produced by Internet applications. They have predetermined data types and lengths, and the structure will not alter until the device’s firmware is updated.
    3. Stream-like: Data sources, such as audio or video programs generate data at a set or constant rate. These data streams are completely unrelated to one another.
    4. Stable flow: Time-series data traffic is constant over time, and it may be calculated and predicted if enough data sources are used within a certain sampling period.
    5. Immutability: The source of this data is a time-series data store. Each data point is created only once and never updated or corrected. Time-series data, like log data, is typically append-only.

    How Is Time-Series Data Used?

    Many of you may be wondering why time-series data is so important. The answer is that it allows us to track changes over time, which can be incredibly valuable for monitoring and troubleshooting purposes.

    For example, let’s say you’re a system administrator tasked with keeping an eye on server performance. By monitoring various metrics—such as CPU usage, memory usage and network traffic—you can quickly identify when a server is starting to experience issues.

    Time-series data can also be used for predictive purposes. For example, if you notice that CPU usage spikes at certain times of day, you can use that information to plan for future capacity needs.

    Of course, time-series data isn’t just limited to server performance. It can be used for any application where it’s important to track changes over time. This includes everything from weather and financial data to social media and website analytics.

    Let’s Get Specific

    Here’s what you should know about time-series data: It’s typically used to seek insights into operations, create alerts based on real-time analysis and forecast future trends.

    The following characteristics are found in time-series data applications:

    1. High write-read ratio: Twitter and LinkedIn are internet applications with single articles that millions of people read, but raw time-series data is primarily scanned and evaluated by apps and algorithms.
    2. Retention policy: In general, time-series data is not kept permanently. Organizations have a retention policy that specifies when and how their data is destroyed.
    3. Real-time analytics and computing: Time-series data must be calculated in real-time to identify out-of-the-ordinary activity and sound alarms based on the acquired information or aggregate findings.
    4. Query scope: Time-series data is generally requested over a period or a set of data sources, and filters are used to prevent all historical information from being requested. Furthermore, all or a portion of the data sources with a filter condition are always aggregated.
    5. Trends: Single data points are typically not significant in time-series data. The emphasis is on how data evolves, such as fluctuations in the previous hour or day.

    Popular time-series solutions like TDengine or InfluxDB, for example, create more efficient processing of time-series data and greater performance than general databases by taking advantage of these qualities.

    Do Time-Series Databases Need Specialized Databases?

    We live in a modern world where data is constantly being generated at an unprecedented rate. To keep up with the demand, businesses need to be able to store and process large amounts of data quickly and efficiently. This is where specialized time-series databases come in.

    Everything is online, including meters, automobiles, lifts, assembly lines and even bicycles. And these items are sending out a never-ending stream of numbers and events that IoT and the cloud have unleashed an explosion in time-series data generation.

    With this in mind, it’s not surprising that several specialized time-series databases have been created to deal with this type of data influx.

    Time-series data sets are enormous and pose a significant problem for general database management systems, such as relational and NoSQL databases. Non-specialized databases have trouble with the following elements of time-series data:

    1. Data ingestion rate: In many time-series data scenarios, millions of data points are generated every second and must be ingested in real-time. Relational databases are not built to handle this quantity of data, and while NoSQL databases may be scaled to do so, the amount of resources required soon becomes prohibitive.
    2. Query latency: It is often the case that a time-series system must scan a massive number of data points to get an aggregate result, which might lead to sluggish performance. For example, it would take days for a basic database to compute the average response time of all Amazon.com clicks, by which time the overall conclusion may be incorrect.
    3. Storage cost: Data generated by internet-connected devices and apps is nonstop 24/7, with a single day producing as much as a terabyte of data. Because relational and NoSQL databases cannot compress this data effectively, storage costs can quickly skyrocket.

    These problems are generally linked to efficiency in processing big data sets, although there are some places where general databases frequently do not meet even the fundamental demands of time-series applications:

    1. Data life cycle management: Time-series data is generally removed in bulk, not one data point at a time, as it ages out.
    2. Roll-up: In most cases, time-series data is collected and rolled up over a set period before being stored in the new table. Raw data and rolled-up data can have distinct life cycles and retention policies.
    3. Special analytic functions: Time-series applications need more specialized features than general databases, such as time-weighted average, moving average, cumulative sum, rate of change, elapsed time for a specific state and the delta between two consecutive data points.
    4. Interpolation: The database management system must be able to interpolate data based on the adjacent data points and rules to regularize data sets when applications or algorithms require.
    5. Continuous query: Therefore, if you hear a user or customer say they want to know when their data is updated, it’s reasonable to assume they want this functionality.
    6. Session and state windows: Aggregations and analytical procedures may be performed on a session or state window, not just time—for example, consider one that only calculates average power consumption when a machine is switched on.

    With general databases, developers must write custom code to implement features specific to their data set. Different data workloads require different database solutions; one size does not fit all. For time-series data, no matter the size of your data set, a purpose-built time-series database is the best tool for the job.

    Why Are Time-Series Databases Becoming Popular?

    Two years ago, time-series databases were the most popular database type in the business, owing to its expanding number of use cases. It is especially useful if you’re executing sophisticated transactions like advertising, e-commerce, supply chain management and so on.

    However, because of the rapid expansion of the IoT, they are becoming even more popular. As more devices become internet-connected and send data—time-series, of course—to the cloud, many industries are interested in purpose-built time-series databases.

    Time-series data is proving to be an important tool, not just for decision-making and optimization but also for industrial applications. Finally, IT infrastructure has been growing rapidly, with everything from servers, containers, network devices, apps, and microservices being monitored to create massive quantities of time-series data.

    Older time-series databases, on the other hand, are often closed systems employing antiquated structures that do not scale to handle the increasing amount of data. A million time-series data points used to seem like a lot, but now millions and even billions of data points are commonplace.

    Finally, integrating old (previous) time-series solutions with popular data analysis tools like artificial intelligence and machine learning platforms is difficult, if not impossible. These legacy systems cannot be transferred to the cloud without a lot of work and their licensing terms are no longer adequate for current apps.

    Final Thoughts

    The expanding market and the constraints of previous time-series databases are making room for a new generation. These new time-series databases are built from the ground up to take advantage of the cloud, big data and AI/ML technologies.

    If your business or application deals with time-series data, you need a purpose-built time-series database. The advantages of using such a database are too great to ignore, and the disadvantages of using an older, less capable database are becoming more and more apparent.

  • The Rogers Outage of 2022: Takeaways for SREs

    The Rogers Outage of 2022: Takeaways for SREs

    When, eight years from now, folks are creating lists of the top IT incidents of the 2020s, there’s a good chance that they’ll include the Rogers outage of 2022. The failure, which made internet and cellular network service unavailable for more than 12 million users across Canada, was one of the most significant outages in memory, in terms of both the number of affected accounts and the level of service disruption.

    For site reliability engineers, there’s a lot to learn from this incident, from both a technical and a crisis-management perspective. Here’s a look at our main takeaways.

    The Rogers Outage: What Happened?

    Before diving into lessons for SREs, let’s go over what happened during the Rogers outage and what we know about the cause of the incident.

    Rogers is a major telecommunications provider that serves around 12 million customers in Canada. Early on July 8, 2022, it suffered an incident with its infrastructure that made internet and wireless phone service unavailable for most or all of its customers. The outage was so complete that users couldn’t even make 911 calls or withdraw money from ATMs–which means the failure brought down not just ordinary services, but also what you might consider resources that are essential for normal daily life.

    Connectivity was restored for most users within about 15 hours. However, some customers reported continuing disruptions for several days after the beginning of the outage.

    What Caused the Rogers Failure?

    At first, Rogers did not reveal much technical detail about what triggered the incident. A statement attributed to the company’s CEO the day after the outage said that a “maintenance update in our core network … caused some of our routers to malfunction.” But that didn’t explain specifically which type of maintenance update the company was trying to apply, why the update didn’t go as planned or which types of routers were affected.

    It wasn’t until several weeks later, when Rogers submitted a response to questions from the Canadian Radio-Television and Telecommunications Commission, that more technical detail came to light. Rogers explained to the CRTC that, during an update process, the company ran code that deleted a routing filter. It’s still unclear exactly how the bad code made its way into the update routine.

    Lessons for SREs

    For SREs, the Rogers outage is an example of a worst-case scenario. It was an extremely disruptive incident that took a relatively long time to mitigate, at least when you consider how severe the outage was. That’s why SREs should use Rogers’ experience to assess what SRE teams can do to avoid issues like this–and manage them more effectively when they do happen.

    Test Your Updates

    From a technical perspective, one of the biggest takeaways from the Rogers outage for SREs is that you should test updates thoroughly before applying them. If engineers had tested their update routine more carefully, they would presumably have detected the bad code that deleted the routing filter before the deletion actually took place.

    To be fair, Rogers is hardly the first company to experience a major outage due to a bad update. Atlassian, for instance, had a similar issue earlier this year, when a buggy migration process caused a long-lasting outage.

    If you’re an SRE, you don’t want your company to be the next to make headlines when something that should be a simple change, like an update, triggers a massive outage. Test, test and test again before bringing your changes into production.

    Build in Redundancy

    It also seems reasonable to conclude that Rogers didn’t have redundancy built into its infrastructure in the form of backup routers that could take over when the primary routers crashed due to the bad update.

    This is another reminder to SREs about why it’s so critical to avoid putting all of your eggs in one basket or creating single points of failure. Redundancy may cost more, and it may not totally prevent disruptions. But it can make them less severe and easier to recover from. It’s not hard to imagine the Rogers incident being a lot less severe if even just 20% or 30% of its traffic could have been automatically redirected to backup routers. That might have been enough to keep critical services operational, at least.

    Have an External Crisis Communications Plan

    Rogers has received more than a little criticism for not communicating well with its customers during the outage. The company waited hours before releasing an initial statement about the failure. It was not very responsive to media queries early on, and, as noted above, it took several weeks before Rogers finally released technical details about what went wrong (and even then, the details emerged only because of a government inquiry into the incident).

    The lack of rapid, transparent communication probably wasn’t deliberate. We have to believe that Rogers didn’t want to leave customers in the dark. Instead, the slow communication likely reflects the absence of a pre-planned strategy for external incident communication in a situation where Rogers’ own networks were non-operational. Without a communications plan, Rogers presumably struggled to decide which information to share, or how to share it, in the midst of a major crisis.

    If Rogers had a playbook in place for communicating with the public and media about this type of issue before its network went down, it probably would have been able to share information more effectively–and it may not have faced as much of a backlash from customers and regulators.

    Internal communications, at least, presumably were smoother. The company did manage to get service back up for most customers within about fifteen hours, which probably would not have happened if its engineers lacked a means of communicating with each other as they worked to resolve the incident.

    Conclusion

    For SREs, the Rogers outage of 2022 is a lesson in the importance of testing, designing redundant infrastructure and preparing for transparent communications before crisis strikes. While more careful planning in these respects may not have totally prevented the incident, it would likely have at least reduced its severity and mitigated the backlash that Rogers now faces from angry customers and regulators.

  • Open Standards Are Key For Realizing Observability

    Open Standards Are Key For Realizing Observability

    Observability has quickly become a major focus of enterprise DevOps. Yet tool sprawl and complexity can hold some observability initiatives back. The 2022 CNCF Microsurvey on observability trends found that 23% of organizations use between 10 and 15 tools for monitoring, metrics and gathering logging and tracing data. A full 50% also noted that engineers and teams using multiple tools presented a top challenge to observability.

    I recently met with Dotan Horovits, principal developer advocate for Logz.io, to get his take on improving how modern development teams handle observability. According to Horovits, open standards and open source projects are the principal drivers to realizing interoperable observability industry-wide. Below, we’ll consider the issues facing observability and introduce OpenTelemetry, a CNCF project quickly gaining traction to unify telemetry data from various sources.

    Open Standards For The Win

    Although proprietary monitoring agents have a lot of intelligence, they inherently use tight coupling between telemetry generation and the backend storage analytics, said Horovits. This can cause vendor lock-in and lead to formatting variations in telemetry data between toolsets. “The open standards have the power to converge our industry to prevent vendor lock-in, even the playing field, and allow us to get any telemetry for any source,” said Horovits.

    Without an open standard, supporting the evolving suite of technologies at the heart of modern polyglot software stacks is challenging. Organizations constantly add to their tech stack, from SQL to NoSQL, Kafka, API gateways and serverless cloud services, and these modules produce varying data formatted in unique styles. What’s more, companies are typically using more than one cloud vendor, further necessitating the need for a decoupled method to represent multi-cloud metrics. “As soon as your IT stack becomes this complex, there is a need to aggregate all telemetry into one unified observability standard,” explained Horovits.

    Introducing OpenTelemetry

    OpenTelemetry is described as a “high-quality, ubiquitous and portable telemetry to enable effective observability.” Using OpenTelemetry (sometimes abbreviated as OTel), engineers can generate, collect and export metrics, logs and traces. OpenTelemetry works with many popular libraries and is completely vendor-neutral and supported by a rising number of tools in the space. OpenTelemetry was accepted to CNCF on May 17, 2019, and is an incubating project.

    The OpenTelemetry registry boasts hundreds of libraries and plugins to export and transform telemetry data from many different environments, from legacy to contemporary sources. By decoupling the data collection and export process, OpenTelemetry is “one agent to rule them all,” Horovits explained. Of course, he added, vendors who adopt OpenTelemetry may still provide their own logic behind it.

    The Benefits of More Open Telemetry Data

    At the time of writing, live public repository data shows OpenTelemetry is the second-most active CNCF project, behind Kubernetes. So, why is the observability community so excited about OpenTelemetry? Well, for one, the interoperability it affords could equate to efficiency and time savings. “OpenTelemetry gives you the freedom to integrate throughout the stack,” stated Horovits. Part of this is because its open source roots make it a force multiplier. “You can never achieve such vast coverage with closed source as you would with a wide-open community,” Horovits said.

    Arguably, an open telemetry toolset has also become necessary due to the sheer complexity of cloud-native architectures. Most organizations are shifting to a distributed microservices architecture supported by a dynamic, elastic container-based infrastructure. In this world, a single service might span multiple clusters. Therefore, engineers might need to consider clusters, pods, nodes and namespaces when filtering for data.

    In this complex reality, there are many dimensions, multiplying what needs to be monitored. And in general, traditional monitoring systems become ineffective at this level. Tool sprawl and lack of consolidation are problems that arise when trying to cater to many environments simultaneously. Thus, a unified, decoupled observability platform makes a lot of sense.

    Future of OpenTelemetry

    We’ve seen vast adoption from cloud vendors and solution providers around OpenTelemetry. And, the project has been moving swiftly—OpenTelemetry has had a distributed tracing specification released for about a year. The project almost has a release candidate for metrics and one for logs is in the works, said Horovits.

    Beyond that, one feature that Horovits is excited about is continuous profiling, which could help users continuously analyze a function and track usage trends across instances. Another exciting potential development is eBPF. Since it’s language-independent and runs at the kernel level, eBPF could extract observability signals without requiring application-specific hooks or integrations. This could increase support for more environments.

    For more information, you can visit the OpenTelemetry site to peruse the documentation or read up on the feature roadmap for OpenTelemetry. Horovits also gave this helpful presentation on OpenTelemetry at KubeCon 2022.

  • Why App Dependency Mapping Is Critical for Cloud Migration

    Why App Dependency Mapping Is Critical for Cloud Migration

    Software dependencies are a crucial part of efficient, component-based programming. At the same time, they can be a hurdle for fast-paced agile development teams, because they can make it more difficult to deploy, update and migrate software applications. Many applications have dozens or hundreds of dependencies, each with its own transitive dependencies, making the problem worse.

    Dependencies are components that provide required functionality that a main component depends on. They can be incorporated into code using package managers like npm or Maven, Git-based code repositories like GitHub and container image registries like Docker Hub. 

    Application dependency mapping involves discovering and identifying dependencies and interactions between application components, their dependencies and underlying infrastructure. Creating a map of application dependencies is an essential part of gaining visibility over complex application structures and understanding the impact of deploying them in different environments.

    Why Application Dependencies Are Critical for Cloud Migration

    Application dependency mapping ensures you have identified all the components you must migrate to the cloud. You may not need to migrate all components to the cloud, but you do need to ensure all dependencies are identified and migrated together. Otherwise, your application may suffer performance issues because important dependencies remain on-premises.

    For example, if you move an application server to the cloud but keep the application’s database on-premises, your application will experience severe performance degradation and may also cause related applications to fail. Once a dependency is disrupted, all related applications suffer. Therefore, when migrating applications to the cloud, you must include all associated dependencies. 

    How Application Dependency Mapping Tools Can Help

    Application dependency mapping can help you avoid poor application performance and service outages. It is a critical component of the migration process, but is also difficult without automated tools. Application dependency mapping tools check an application and help with the following:

    • Modeling all inter-server relationships 
    • Identifying inbound and outbound connection latency
    • Determining the necessary TCP ports
    • Detecting running processes 
    • Application performance monitoring over time

    Cloud vendors offer specific application dependency mapping tools developed for their environments. For example, Amazon Web Services (AWS), Google Cloud and Microsoft Azure offer proprietary tools to help manage this process. However, these tools are tied to each provider, which means you should use the vendor tool that matches your chosen target cloud environment.

    Alternatively, you can use open source tools that provide similar services but are vendor agnostic. Use these tools if you want an assessment that is not specific to a single cloud vendor environment.

    Mapping Application Dependencies to Prepare for Cloud Migration

    An application consists of a hierarchy of dependent APIs and tools. The hierarchy starts with the application’s interfaces and then goes down through platform tools. Dependency management helps identify combinations of related versions and ensures that the development team recognizes new application dependencies when changes occur.

    Versioning application components

    The first step in this process involves versioning application components. If you intend to deploy a specific software component independently, you need to assign a version number for every revision and then track the dependency chain for this version. 

    This technique ensures you know the specific platform tool versions associated with each application version. If you need to roll back, you know what additional components you need to roll back to ensure version compatibility.

    Changing platform components

    This technique requires changing some platform components, such as middleware, which also requires synchronizing each application’s platform version. You should start from the top of the application dependency chain.

    Each application is designed to use specific operating system and middleware features, expecting “version Y or later” of the tools. As a result, you must validate each tool with a specified version against the tool’s dependencies and continue validating each of these dependencies until you reach the bottom of the dependency chain.

    However, some application dependencies are not as obvious and explicit as others. You may encounter problems if you expect to run a guest operating system in a virtual machine (VM) on a hypervisor platform in a different source. 

    You may also encounter issues if a cloud stack release, like OpenStack, requires a certain scripting language version. You can mitigate this issue by testing every dependency chain against the standard middleware and operating system combination, ensuring that all dependencies are recognized.

    Rebuild the dependency tree

    Once you are ready to move to the cloud, you need to supplement the dependency tree for applications and include all the references to the cloud provider’s APIs and features. Make sure to determine how the provider notifies of changes to tools and APIs and prepare to validate new dependencies these changes can create.

    For a multi-cloud or hybrid deployment, you must compare cloud dependency trees across all the components and applications planned to migrate beyond a cloud platform boundary. Note that having a different dependency tree for each provider can result in problems when scaling or failing over between providers’ platforms. You can avoid this by synchronizing components in advance.

    You should also rebuild your dependency trees whenever you change software platform components. A basic update could undo all the mapping and work you have done, and you can easily overlook a change to a small middleware component. You can mitigate this by setting up a life cycle management process to ensure dependency problems do not arise when incompatible elements are introduced. 

    Conclusion

    In this article, I explained the basics of application dependency mapping, and showed how application dependency tools can make cloud migrations safer and easier:

    • Versioning application components—Understanding what version your application expects for all dependent components.
    • Changing platform components—Determining the impact of switching out certain platform components. 
    • Rebuild the dependency tree—Use the map of your existing dependency tree to rebuild a matching dependency tree in the cloud environment.

    I hope this will help you evaluate the use of dependency mapping and make your next migration a success.

  • What Are the Seven Layers of the OSI Model?

    What Are the Seven Layers of the OSI Model?

    Working in the software field, it’s not uncommon for engineers to refer to different “layers.” Perhaps you’re working with protocols at the “networking layer” or evaluating a solution that sits at “Layer 4” or “Layer 7.” Although these concepts are obvious to some, not everyone knows what these layers refer to. To get you up to speed, below, we’ll define the seven layers of the OSI Model.

    What Is the OSI Model?

    The Open Systems Interconnection model (OSI model) is a conceptual model to help visualize a computerized network. Made up of seven layers, it represents the distinct levels that make up an end-to-end computing system. In the spirit of promoting open interoperability, the model is intended to represent a universal standard that’s agnostic to a specific technology or vendor.

    The OSI Model was co-conceived by International Organization for Standardization (ISO) and Internet Engineering Task Force (IETF) and first published in 1984. It has since been redefined as ISO/IEC 7498-1:1994. Although the OSI Model has been around for decades, it’s still referenced quite often.

    The Seven OSI Model Layers

    The OSI Model is split into seven abstraction layers: Physical, data link, network, transport, session, presentation and application. You can think of the bottom one, Layer 1 (the physical layer), as the closest to the most rudimentary electrical connections. The farther up you rise, the closer you get to Layer 7 (application layer), which is the most user-facing category.

    1. Physical Layer

    The physical layer, Layer 1, represents the lowest level of data transmission. This category refers to raw unstructured data bits—and the process of converting them into electrical signals for a device to read. “Physical” hardware like network hubs, modems, adapters, cabling and network controllers work in this layer. Bluetooth, Ethernet and USB describe physical layer specifications.

    The data link layer deals with data packaged into frames. Technologies in this layer assist in node-to-node data transfer, and the protocols here describe how to establish and terminate a connection and how the connection should flow. The data link layer is further divided into two sublayers: Media access control (MAC), responsible for how nodes connect with one another, and logical link control (LLC), which checks for errors and orchestrates the frame flow and synchronization. An example data link layer protocol is Point-to-Point Protocol (PPP). MACsec also can work to apply encryption at this layer.

    3. Network Layer

    The network layer is responsible for sending and receiving data frames structured in packets. Technologies in the networking layer use routers to send packets to nodes on different networks. The network layer might split or fragment the message if it exceeds the maximum network packet sizes. You’re probably familiar with the network layer if you’ve looked up your IP address—the network layer uses the IP protocol (or other logical protocols) to find locations. Such protocols specify a message along with the address of the intended recipient node.

    4. Transport Layer

    The transport layer deals with the transmission of data segments, which are sequences of data at variable lengths. The goal of the transport layer is to optimize data transmission by performing data segmentation and adjusting the size of the package or the rate of transmission. Transport layer protocols include TCP and UDP. Tunneling protocols also operate at the transport layer. The OSI Model also defines five classes of the connection-mode transport protocol.

    5. Session Layer

    The session layer handles the communications between two or more computers. Protocols here are used to create a “session” between entities, which is common in applications that use remote procedure calls. The session layer handles the connection and authentication between a client or server, including actions like logon, look up, log off or session termination. DNS, along with name resolution protocols, operate in the session layer.

    6. Presentation Layer

    The presentation layer is about data translation and formatting. In this layer, protocols handle things like encryption, decryption, compression and decompression. The goal of the presentation layer is to transform data in such a way that it can be sent over a network in syntax that fits the constructs specified by the application layer. For example, technologies that serialize data structures to XML or JSON can be thought of as working for the presentation layer. This sort of data translates into what is graphically displayed for the end-user.

    7. Application Layer

    The application layer is the top layer of the OSI Model and sits closest to the end-user application. User-facing software directly interacts with the application layer through functions like file sharing, message handling or database access. High-level protocols such as HTTP and FTP are used in this layer to share resources. Web browsers and email clients are examples of applications that interact with the application layer.

    The OSI Models: A Helpful Taxonomy

    There you have it—a primer or refresher of all seven OSI Model layers. It’s important to note that the model does not intend to serve as an implementation specification itself—it’s purely a conceptual framework. But it does help provide context into where tools operate and how they interact with elements of a distributed computing system.

    In the context of DevOps technologies, Envoy-based service mesh is often described as working on Layer 4 (the transport layer) and Layer 7 (the application layer). Or, since eBPF filters network frames, it can be said to work at Layer 2 of the OSI Model.

  • RapidAPI Adds Tool to Make Building APIs Simpler

    RapidAPI Adds Tool to Make Building APIs Simpler

    RapidAPI today added a free RapidAPI Studio offering to its portfolio that makes it easier for developers to build, consume, manage and monetize application programming interfaces (APIs).

    Wade Wegner, senior vice president and head of product at RapidAPI, said RapidAPI Studio, available in beta, makes it simpler for developers to move between design and development, testing and management without losing context.

    It is also designed to be slipstreamed into developer workflows by either being accessed via a browser, as a desktop application, a macOS-native application or as an extension to the VS Code tools provided by Microsoft, added Wegner.

    That approach also makes it easier for teams of developers to collaboratively work on API projects, he noted.

    APIs are foundational to every modern application development initiative. Each microservice included in a modern application has its own API. As a result, the number of APIs—internal and external—that organizations are building and exposing has increased exponentially, especially as the number of digital business transformation initiatives continues to expand. While the bulk of APIs employed today are internally facing, the number of external-facing APIs being employed tends to rise sharply as more business processes become digitized.

    Less clear is who within the organization is responsible for managing those APIs after they have been deployed. In some cases, developers assume responsibility for them along with every other component of an application. Developers, however, tend to move on to other projects; DevOps and cybersecurity teams are being tasked with managing, securing and updating APIs. This is part of a larger effort to better secure software supply chains in the wake of a series of high-profile breaches.

    Unfortunately, documentation of those APIs has been somewhat lax within many organizations. In many cases, organizations are not even sure how many APIs have been deployed. Many organizations also discover, to their chagrin, that so-called zombie APIs previously abandoned by developers can still be exploited by cybercriminals to exfiltrate data.

    At the same time, the management of APIs is becoming more complex. Along with existing REST APIs and legacy web services based on XML, organizations are now deploying GraphQL APIs. The result is a mix of API schemas that require different levels of expertise to manage and secure.

    It’s not clear how the increased focus on securing software supply chains might impact API management. However, it’s apparent that cybercriminals are becoming more adept at discovering insecure APIs through which they can exfiltrate data. The challenge organizations face is that it’s not feasible to secure those APIs without first having a framework in place for managing them.
    One way or another, an API management crisis will soon come to a head as it becomes even simpler to build and deploy APIs. The only thing that remains to be seen is to what degree organizations will get ahead of that issue before it spins out of control altogether.

  • Dynatrace Embeds DX Monitoring in Observability Platform

    Dynatrace Embeds DX Monitoring in Observability Platform

    Dynatrace has enhanced its analytics capabilities to enable digital experience monitoring, including session replays, based on the logs, metrics and traces it collects via its observability platform.

    Steve Tack, senior vice president of product management at Dynatrace, said this update makes it simpler for DevOps teams to track user journeys across the multiple components and microservices that make up a modern application environment. Rather than trying to, for example, determine the impact of an issue by parsing through log data without any context, Tack said Dynatrace’s analytics capabilities will automatically fire up a session replay to enable DevOps teams to view the exact nature of the application issue being experienced.

    Previously, DevOps teams would have needed to deploy the Dynatrace session replay offering separately. The overall goal is to remove much of the friction DevOps teams encounter when trying to fix an application issue by providing context in a way that is more accessible, he added.

    In general, as observability continues to advance, it’s now become much less about the amount and types of data that can be collected, said Tack. The data is table stakes for achieving observability, he added. The focus now needs to be on the richness of the actionable analytics being surfaced once that data is collected, he said.

    In fact, as more organizations start to build and deploy applications based on highly interdependent microservices, the richness of the analytics is now more critical than ever, noted Tack. While microservices-based applications are more resilient than monolithic applications, troubleshooting them is more challenging whenever application performance degrades, he added.

    It’s not clear at what rate DevOps teams are embracing observability. The concept has always been a core DevOps principle, but, for the most part, more mature DevOps teams have only been able to achieve what amounts to continuous monitoring of pre-defined metrics. Observability platforms promise to surface anomalies indicative of IT issues before they escalate; DevOps teams can then launch queries to better determine their root cause and potential severity.

    Dynatrace is making a case for an observability platform, accessed as a cloud service, that employs the company’s Davis AI engine and machine learning algorithms to surface the root cause of application issues.

    As observability platforms continue to evolve, the days when IT teams determined the root cause of an IT issue by process of elimination are finally coming to an end. DevOps teams will be gaining more visibility into their application environments in ways that surface more meaningful actionable intelligence. That intelligence is especially critical for externally-facing applications driving digital business transformation initiatives.

    There is, of course, no shortage of options when it comes to observability platforms. In some cases, IT organizations are opting to rely on cloud services while others are extending DevOps platforms they built themselves and continue to maintain. Regardless of the approach, the issue will no longer be how to collect data as much as what to do with it once it’s collected.

  • Deciphering the Observability Market

    Deciphering the Observability Market

    Observability has become one of the most overused buzzwords in IT and cybersecurity. Today, the term is used by vendors to refer to everything from application performance to network monitoring, cybersecurity and data and analytics.

    While the term’s ubiquity has created confusion for everyone from end users to journalists, startups in the space have also attracted over $2 billion dollars in venture capital investment over the past two years. This momentum prompted TechCrunch to ask if this area, specifically data observability, was effectively recession-proof.

    This observation triggers two questions. First, what is observability? And second, what are the different kinds or variants?

    What is Observability?

    Observability, as a broad practice or capability, originated in the 1960s as part of industrial control theory. The idea is that by watching the output of a system, you can figure out what’s happening inside the black box. This sounds a lot like monitoring, a popular practice for IT and security teams. However, observability takes a different approach than monitoring’s alert-based methods; it allows you to ask questions about a system that dig deeper than the pre-defined thresholds from your monitoring system. 

    Observability requires collecting massive amounts of data from systems, networks and applications to feed its discovery process. As systems and applications become more complex, figuring out what went wrong and why becomes much more challenging. At a high level, these are the kinds of challenges many startups in this area are addressing. Each takes a different approach and solves a different part of the observability problem.

    Types of Observability

    The most frequently mentioned type is data observability, which is concerned with the health and quality of data passing through data pipelines for analytical use cases like feeding a data warehouse. Vendors like Acceldata, Bigeye or Monte Carlo Data typically target their products to data engineers tasked with building and operating analytical data pipelines. 

    Another group of companies addresses applications. These firms, like Honeycomb and Observe.ai, collect data from applications to help site reliability engineers (SREs) understand performance issues and aid them in debugging and troubleshooting. 

    There are also companies specializing in machine learning observability, and these are concerned with the performance and drift of models in production. These companies, like Arize and WhyLabs, target data scientists. 

    Finally, there are companies in the observability data space which is comprised of logs, events, metrics and traces essential for all other forms of observability and monitoring to work. Companies in this space, like Cribl, use specially-designed pipelines to connect the sources and destinations of observability and security data. These pipelines allow companies to route data to multiple destinations, enrich data in flight and reduce data volumes before ingestion.

    As you can see, there are dozens of companies in the space. Each category of vendors addresses a unique challenge in today’s sprawling IT and security landscape.

    Beware Observability-Washing

    As the hype around this topic grows, companies hoping to differentiate themselves from their competitors may suddenly rebrand as observability companies. This is already happening. Legacy monitoring and alerting companies are calling themselves observability companies, hoping to attract the same kind of attention as others in the space that actually do have differentiated products and positioning.

  • Deloitte Aligns with Dynatrace for Observability

    Deloitte Aligns with Dynatrace for Observability

    Deloitte has expanded its cloud observability practice to include the Dynatrace Software Intelligence platform.

    Jay McDonald, managing director and co-chair for modern delivery at Deloitte, said the IT services provider will now provide traditional DevOps and observability consulting expertise along with a set of instances of the Dynatrace Software Intelligence Platform that it will manage on behalf of organizations.

    Deloitte opted for the Dynatrace platform because it employs a single agent for collecting metrics, logs and traces that are then analyzed by its Davis artificial intelligence (AI) engine to enable observability, said McDonald.

    While there is a growing appreciation for observability to enable DevOps teams to identify the root causes of IT issues before they become major problems, deploying an observability platform requires a lot of specialized expertise. Many organizations are opting to rely on third-party organizations such as Deloitte to manage that task within the context of a larger DevOps workflow, said McDonald.

    Managing those workflows has become especially challenging for many organizations to deploy and maintain as they move to build and deploy microservices-based applications, he added.

    Michael Allen, vice president of worldwide partners for Dynatrace, said there are now already more than 4,000 organizations using the Dynatrace Software Intelligence Platform, a number that will continue to increase as observability platforms are consumed as a service.

    It’s not clear at what rate DevOps teams are embracing observability. The concept has always been a core DevOps principle, but, for the most part, more mature DevOps teams have only been able to achieve what amounts to continuous monitoring of pre-defined metrics. Observability platforms promise to surface anomalies indicative of IT issues before they escalate; DevOps teams can then launch queries to better determine their root cause and potential severity.

    One of the benefits of the Davis AI engine is that machine learning algorithms surface the root cause of many of those issues without requiring DevOps teams to know exactly how to construct the right query, noted Allen.

    However, the biggest challenge many DevOps teams encounter is they lack the skills required to deploy and manage observability platforms. In fact, many of them are now relying on third parties to manage a wide range of DevOps platforms so they can concentrate their efforts on actual workflows. Hiring and retaining DevOps professionals also has become more challenging as the total number of applications being built and deployed expands beyond DevOps teams’ abilities to effectively manage.

    One way or another, the days when IT teams determined the root cause of an IT issue by process of elimination are finally coming to an end. DevOps teams are gaining more visibility into their application environments in ways that surface more meaningful actionable intelligence, regardless of who sets and manages the observability platform. The opportunity to employ that intelligence to improve overall application resiliency and drive a wide range of digital business transformation initiatives is well within DevOps teams’ reach.

  • Why More Incidents Are Better

    Why More Incidents Are Better

    Ask most SREs how many incidents they’d have to respond to in a perfect world, and their answer would probably be ‘zero.’ After all, making software and infrastructure so reliable that incidents never occur is the dream that SREs are theoretically chasing.

    Reducing the number of actual incidents as much as possible is a noble goal. However, it’s important to recognize that incidents aren’t an SRE’s number-one enemy. What matters more than the number of incidents you experience is how effectively you respond to each one.

    Plus, there’s value in incidents. They are a learning opportunity. If your business never experienced them, it would arguably be facing more risk, not less.

    We know: These ideas may sound a little counterintuitive. You might even accuse us of being “pro-incident”—which we sort of are. Allow us to explain.

    The Silver Lining

    In many respects, incidents are inherently bad. When one occurs, it means something broke. That’s bad. It may also mean that users were disrupted, operations halted or money was lost. Those things are even worse.

    On the other hand, incidents aren’t all bad. They actually benefit SRE teams for several reasons:

    • Learning opportunities: Incidents are opportunities to figure out what went wrong and prevent it from recurring. They can also help teams learn how to react more quickly or efficiently the next time something fails.
    • Get ahead of bigger issues: Sometimes, working through one incident means you can avoid another that’s even worse. Perhaps one server fails, for example, and your response revealed that the failure was due to a larger issue that would have eventually caused a worse outage if left unaddressed. But thanks to the incident, you detected the larger issue before it triggered a more massive failure.
    • Reinforce team culture: Nothing breeds camaraderie or a spirit of collaboration like working alongside other engineers to respond to a crisis in the middle of the night. Although being in this setting may not be anyone’s first choice, it does often have a positive impact on your team’s culture and esprit de corps.
    • Demonstrating value: Assuming you handle them well, incidents are an opportunity for SREs to prove how valuable they are to the organization. If incidents never happened, it’s a safe bet that some bosses would start to wonder why they need SREs in the first place. (It would be a flawed train of thought, of course, because SREs would deserve credit for preventing incidents, but it’s a thought that may float around some C-level brains nonetheless.)

    We could go on, but the point is clear: Although incidents cause problems in some respects, they actually create value in others.

    Focus on Response, Not Avoidance

    This is not to say that you should welcome incidents with open arms. Obviously, any decent SRE should focus first and foremost on being proactive and preventing incidents from happening whenever possible. They should use chaos engineering to identify problems that could be lurking unseen in production environments. They should leverage IaC to minimize risks. And so on.

    That said, what ultimately matters more than incident frequency is the effectiveness of incident response. It’s better to experience ten incidents that you resolve in under an hour each than one incident that takes mission-critical systems offline for a week.

    So, in addition to investing in tools and processes that mitigate the risk of incidents, SRE teams should place equal emphasis on ensuring that they can react quickly and effectively when an incident happens. This means having the ability to share information efficiently, define clear roles, know what to prioritize when working through complex incidents and have clear plans in place that spell out how you’ll handle a problem as soon as you detect it. Without these abilities, you’re at risk of letting incidents that should be small turn into major outages.

    ‘Zero Incidents’ is not Realistic

    It’s important to recognize, too, that while it can be fun to imagine a world where zero incidents occur, the reality is that such a world will never exist. If it could, we wouldn’t see new records set each year for the number of security incidents that businesses collectively suffer, for example.

    Nor would we see headlines about major outages at huge enterprises like Facebook or AWS on a recurring basis. If those companies, which have world-class reliability teams and virtually endless resources at their disposal, can’t reduce incidents to zero, neither can anyone else.

    Conclusion

    The bottom line: There is no such thing as total incident prevention, no matter how hard you try. And even if there were, that wouldn’t actually be a good thing, for the reasons explained above.

    So, by all means, undertake reasonable proactive efforts to prevent as many incidents as you can from happening. But don’t let investment in prevention cause under-investment in response. Being prepared to handle incidents when they happen—which they inevitably will—is what matters most.