Category: Application Performance Management/Monitoring

  • US DoJ Makes PyPI Give Up User Data ¦ Tape Storage: Not Dead

    US DoJ Makes PyPI Give Up User Data ¦ Tape Storage: Not Dead

    Welcome to The Long View—where we peruse the news of the week and strip it to the essentials. Let’s work out what really matters.

    This week: PyPI complies with a “string of subpoenas,” and LTO continues to grow, despite predictions of its demise. (more…)

  • AWS Identity and Access Management (IAM) Roles and How to use Them

    AWS Identity and Access Management (IAM) Roles and How to use Them

    Amazon Web Services relies on the AWS IAM service to govern who is authenticated and authorized to use AWS resources. It plays a hugely significant role in AWS security–and so do its various identities. Let’s look in more detail at IAM roles and how they work.

    What is the Purpose of IAM?

    When you create an AWS account, it has a single sign-on (SSO) identity called the AWS account root user, which has complete access to AWS services and resources. To avoid potentially disastrous issues arising from this kind of unfettered access, IAM allows shared access to the AWS account via identities.

    Based on the security-first model, AWS restricts all actions for all identities by default–apart from the root user. This restriction enables us to manage granular access for identities following the principle of least privilege, so permissions are granted only as required to execute daily tasks.

    What are IAM Roles?

    AWS users, federated identities and AWS roles are the three major categories of AWS Identities. IAM roles are conceptually similar to AWS users, but whereas a user is uniquely associated with a principal(users/apps/etc.), a role can be assumed by anyone who needs it.

    Rather than passwords or access keys, temporary security credentials are provided to whoever assumes the role. This eliminates the overhead of managing users and their long-lived credentials.

    Any of the following entities can assume a role to use its permissions:
    ● AWS user from the same account
    ● AWS user from a different account
    ● AWS service
    ● Federated identity

    Structure of an IAM Role

    The two essential aspects of an IAM role are the trust policy (who can assume the IAM role) and the permission policy (what the role permits the user to do).

    Trust policies use specified conditions to define and control which principals can assume the role. They prevent the misuse of IAM roles by unauthorized or unintended entities.

    Permission policies define what the principals assuming the role are allowed to do.

    Types of IAM Roles

    AWS IAM roles fall into one of the following major categories based on their trust policies:
    ● Service role
    ● Service-linked role
    ● Web Identity role
    ● SAML 2.0 federation role
    ● Custom IAM role

    Service Roles
    AWS services are the trusted entity type for these roles, which are created to allow AWS services to perform actions on the user’s behalf. They do this by inheriting the permissions assigned to the service role.

    You might wonder why AWS services need permission to access each other. It’s because, by default, even AWS services have no access to the resources in the AWS account. However, service roles allow AWS services to access resources based on their requirements.

    Most AWS services rely on service roles for optimal functioning. For example, they allow Cloudformation to create and delete resources on the user’s behalf based on a YAML or JSON file. The Amazon EC2 IAM role is a special type of service role. EC2 relies on an instance profile as a container for an IAM role, which is then assumed by the applications running within the EC2 instance to perform actions the role allows.

    Service-Linked Roles
    A service-linked role is a unique kind of IAM role linked to an AWS service. It simplifies the process of setting up a service by automatically adding all the required permissions for a service to perform actions on the user’s behalf. It is predefined by the service. Most service-linked roles do not permit changes to trust or permission policies.

    Web Identity Roles
    A user assumes a web identity role when they log in to AWS using an identity provider (IdP) such as Amazon and Facebook. Users do not have an identity within AWS itself; in exchange for an authentication token, they get temporary security credentials in AWS that map to an IAM role that is authorized to use the resources in the AWS account.

    SAML 2.0 Federation Roles
    These roles are assumed by users who are included in an external user directory, typically within organizations. This enables federated single sign-on (SSO), so that organizations can give users access to the AWS console and CLI without having to create a separate IAM user for each person. An organization that manages its employees in Microsoft Azure Active Directory could connect to AWS directly and give its users access to the AWS console and CLI based on the permissions provided to the SAML 2.0 federation role.

    Custom IAM role
    Custom IAM roles do not fit any of the other categories and support scenarios other than the ones listed above.

    What is the Time Limit for Assuming a Role?

    An entity can assume a role for as long as the IAM role’s session duration property dictates. For example, if the session duration for a role is set at 12 hours, the temporary credentials will expire 12 hours after they are issued.

    If you assume the role using the assume-role* command, you can specify a value for the session duration using the duration-seconds flag, ranging from 15 minutes to the maximum session duration permitted for the role. The default duration is one hour, but it can be extended to up to 12 hours. The session duration can be adjusted by editing the role.

    What About Role Chaining?

    One role can assume another role in a process known as role chaining. Role chaining cannot extend beyond an hour on the user session; it doesn’t follow the role’s maximum session duration field.

    And Cross-Account Access?

    IAM roles are often leveraged to enable cross-account access – the process of leveraging a principal in one account for access to resources in a different account. Organizations often maintain multiple AWS accounts to segregate environments such as development, staging, test, UAT, and production, and they grant identities from one account permission to access resources in another.

    This could be to process data in the production account and then anonymize and copy it to the UAT account, so that the frequency and attributes of data are the same. This keeps the UAT environment as close to the production environment as possible.

    How do IAM Roles Differ From Resource-Based Policies?

    Both identity-based policies and resource-based policies are permission policies, but identity-based policies are attached to identities such as users, roles, or groups, and resource-based policies are attached to resources such as S3, Amazon SQS queues, or VPC endpoints.

    Identity-based policies specify the actions the identity can take and on which resources, whereas resource-based policies determine who is allowed to access the resource to which they are linked and specify the actions they can perform. Resource-based policies must be inline and not managed. This list provides details of which resources support resource-based policies.

    Where both identity and resource-based policies are present, AWS evaluates them together. An action is permitted if it is allowed in either or both policies. If either policy contains an explicit DENY, it overrides the ALLOW.

    You can provide cross-account access using just the resources policy rather than using a role as a proxy. This is achieved by attaching a resource-based policy on the resource you want to share. This resource policy comprises all the principals allowed to access this resource, in contrast with an identity-based policy, which specifies the resources a principal has access to.

    Using a resource-based policy for cross-account access rather than roles has advantages.

    When a user assumes a role, they cede the permissions they have in the trusted account so that they can secure permissions in the trusting account (account with shared resources). With a resource-based policy, the user keeps their permissions in the trusted account, and they also secure access to the shared resources in the trusting account. The principal can access both accounts.

    AWS IAM Roles Anywhere

    AWS IAM Roles Anywhere is a kind of service role that permits on-prem machines or workloads external to AWS (such as servers, containers, and applications) to access resources on AWS by acquiring temporary security credentials. This completely eliminates the headache of managing long-term credentials.

    Workloads are required to have X.509 certificates issued by a certificate authority, which must be registered with IAM Roles Anywhere as a trust anchor. Though a relatively recent use case, it marks a move toward minimizing the usage of long-term credentials.

    Benefits of Using IAM Roles

    ● No management of long-term security credentials
    ● Support for Single Sign-on (SSO) using SAML 2.0
    ● Support for web identity federation, allowing users to log in via popular external identity providers (IdPs), such as Amazon, Facebook, and Google
    ● Enhanced security posture because you don’t have to rotate and replace short-term credentials (An expiry is already associated with them).
    ● Support for use cases such as cross-account access, Identity federation, AWS IAM Roles Anywhere

    Key Takeaways

    IAM roles are key pillars supporting, not just the IAM service, but the entire authentication and authorization flow. Their value extends beyond applications such as cross-account access, AWS IAM Roles Anywhere, and identity federation, being important for enhancing any organization’s security posture substantially. When used to their full potential, they can be extremely useful in IAM.

  • FinOps Foundation’s FOCUS Aims to Standardize Cloud Billing

    FinOps Foundation’s FOCUS Aims to Standardize Cloud Billing

    Cloud spending is on the rise. As organizations shift more applications to the cloud, there is a growing need to optimize workloads and reduce spending when possible. But a lack of visibility into complex cloud bills often impedes these goals. This is compounded by a lack of data standardization across the multi-cloud landscape, with cloud vendors producing lengthy billing reports in various formats. With things so decentralized, FinOps practitioners often struggle to make sense of it all.

    FinOps Open Cost & Usage Specification (FOCUS) is set to change this status quo. FOCUS, an initiative sponsored by the FinOps Foundation, a Linux Foundation technical project, aims to create a standard, vendor-neutral specification for cloud cost, usage and billing data. A common schema for cloud billing should help unify cost data across cloud vendors and greatly aid FinOps objectives in areas like allocation, chargeback, budgeting and forecasting.

    I recently met with Udam Dewaraja, FOCUS lead at FinOps Foundation, to understand the context that led to creating the specification. Below, we’ll dive a little deeper into FOCUS, examining the details of the proposed specification and considering what benefits it will bring to cloud consumers and cloud service providers (CSPs).

    Why the Cloud Needs a Common Billing Schema

    The complexities of today’s cloud services can often create headaches around cloud billing. For instance, AWS alone has hundreds of services and potentially thousands of SKUs. Cloud service providers each have their own data formats and terminologies, which are reflected differently in detailed bills. And with the granularity of cloud metering, enterprises could be staring at hundreds of millions of rows of cost data. Now, extrapolate that across multiple cloud service providers, SaaS and APIs, and you can see the dilemma that the modern CFO faces.

    “There are many nuances that FinOps practitioners need to understand,” said Dewaraja. “And a lack of a common definition can be a scalability challenge as you add more things.”

    It’s relatively easy to spend money on more and more cloud-based consumption models. As Dewaraja said, “Everyone has the company credit card.” Of course, you can use tools like AWS Cost Explorer, Cloudability or Datadog to garner some insights to reduce spending. But to really get the fuller picture, you need to rationalize cost metrics across many areas. Plus, in 2023, The State of FinOps report found the widest variety of cloud cost management tools in use in the history of the report, indicating tooling sprawl.

    Overview of FinOps Foundation

    First, for those unfamiliar, the FinOps Foundation is a Linux Foundation-supported group that aims to help practitioners understand the discipline of FinOps. Representing a community of over ten thousand FinOps practitioners, the FinOps Foundation routinely shares best practices and oversees a FinOps framework. “FinOps is more of a cultural practice that should be adopted,” explained Dewaraja. As such, the initiative aims to oversee training, meetups and events and encourage executive buy-in around the FinOps concept.

    Understanding the FOCUS Specification

    So, what exactly does FOCUS entail? Well, FOCUS is an open source royalty-free specification for cloud billing data. It will provide standard dimensions, attributes, and terminologies for cloud billing data sets. To formalize the specification, the project contributors are actively considering the essential FinOps capabilities and what requirements these cloud bill data sets have in common.

    The specification itself has yet to be officially released and is still being formalized at the time of writing. Microsoft and Google have joined the initiative, as well as numerous cloud finance and cost analysis vendors. And while the group waits for additional CSPs to join, the community is building out open source validators and converters, said Dewaraja.

    Benefits of a Standard Cloud Billing Spec

    Cost management solutions are already available, but what’s not readily available is a method to make business decisions based on broader richer datasets, explained Dewaraja. This is where standardized cloud billing can greatly aid efforts by creating a more holistic picture.

    Specifically, this can aid objectives like allocation, chargeback, budgeting, and forecasting. Doing so is often an incremental process, explained Dewaraja—as you gather insights into data, you iteratively use this visibility to make improvements over time.

    Some other benefits of adopting FOCUS will likely be:

    Putting finance in engineering terms. In the cloud, finance is quite dynamic. And interestingly, the FinOps discipline is becoming engineering’s responsibility as part of the shift left approach, said Dewaraja. Engineers now need to understand this problem and how to architect their solutions cost-effectively. As such, bringing specifications into the mix will help put financial responsibility squarely on the shoulders of engineering terms.

    Bringing accountability to cloud consumption. A standard specification should empower essential FinOps capabilities that rely on usage segmentation like internal cost reporting, showback and chargeback. Because if invoices are just one big pie, there is no accountability for individual departments, and you can’t drive the behaviors you need, said Dewaraja.

    Win-win for practitioners and cloud vendors. With FOCUS, cloud providers will drive more trust in cost data, providing more visibility. And corporations using the cloud can make their adoption more efficient, likely influencing them to bring on more workloads. “Driving trust and understanding of this cloud data is mutually beneficial,” Dewaraja said.

    FOCUS Set to Improve FinOps

    Dewaraja, who previously spearheaded FinOps practices at Citigroup, knows all too well the strains of managing disparate cloud billing data. He oversaw a broad, dynamic cloud-native platform of distributed cloud and SaaS and sought to embrace platform engineering principles. In this environment, you need a holistic cost picture that associates various expenses to make the right calls, he said. “You really see success if you can plug this into the workflows of those developers and into business processes themselves,” he said.

    Theoretically, FOCUS could deliver on these goals. But although the project is gathering momentum, it’s still in an early stage. For the standard to gain a foothold, it will require greater awareness and more contribution from other CSPs. The specification is community-driven but is being led from the practitioner down, adopting corporate contribution licenses instead of individual developer contributions. Thus, cloud vendors and cost management vendors will likely lead the initial charge.

    FOCUS looks great in the long run but will require a little bit of work upfront to get folks on board. That said, it should be on the radar for 2024 roadmaps, said Dewaraja. To learn more about FOCUS and how to contribute, software providers should check out the FinOps Foundation materials.

  • Red Hat Enhances Insights to Simplify RHEL Management

    Red Hat Enhances Insights to Simplify RHEL Management

    This week at its Red Hat Summit event, Red Hat introduced an expanded set of capabilities to Red Hat Insights that simplify management of Red Hat Enterprise Linux (RHEL).

    Managing Linux across the hybrid cloud is complex, and it’s made even more so because of current macroeconomic conditions and the challenge of finding and retaining highly skilled DevOps and IT professionals, said Gunnar Hellekson, vice president and general manager, Red Hat Enterprise Linux, Red Hat.

    The enhanced Insights capabilities add a layer of abstraction to help reduce enterprise Linux complexity across the hybrid cloud and make RHEL management easier by lowering skill level requirements for managing Linux server estates, the company said.

    Using a single UI, entry-level systems administrators and IT support teams can access Red Hat Insights’ predictive analytics to identify potential bugs, misconfigurations or security vulnerabilities and remediate them. Systems administrators of all skill levels can detect, assess and push fixes for potential problems without a deep understanding of Linux management systems like Red Hat Satellite Server and without having to interact with the command line, the company said.

    These enhancements help organizations maintain services and systems even in uncertain macroeconomic times when hiring is constrained. Red Hat research also claimed that these enhanced capabilities enabled IT organizations to find IT issues up to 90% faster and remediate them nearly 66% faster across the hybrid cloud.

    A Red Hat Insights image builder service also helps lower the skills level needed to manage Linux estates by enabling IT teams to build standardized, optimized operating system images, including the creation of their own ‘golden images’ that adhere to organization-specific needs for security and compliance. These images can be deployed across public clouds, hybrid clouds, data centers and the edge using the same console.

    As the economic downturn continues, organizations are looking to do more with less and identify ways to make the most of the resources, talent and skillsets they already have. Red Hat is making the case for adding a layer of abstraction to allow less-highly-skilled IT and DevOps team members to contribute to managing Linux estates across hybrid, multi-cloud and edge deployments.

    “The demands of managing Linux across the hybrid cloud are stretching the ability of resource-constrained IT departments to keep pace,” said Hellekson in a statement. “CIOs need to be able to extend existing skillsets and lower the skill barriers for overseeing Linux estates, which is exactly what the management capabilities of Red Hat Insights are designed to do.”

  • ServiceNow Adds Observability Platform to SaaS Portfolio

    ServiceNow Adds Observability Platform to SaaS Portfolio

    At its Knowledge 2023 conference, ServiceNow launched an observability service that runs on its software-as-a-service (SaaS) platform.

    Based on the observability platform the company gained with its acquisition of Lightstep, the ServiceNow Cloud Observability platform makes it possible to collect and analyze data collected from software running on infrastructure and also from workflows created using ServiceNow applications.

    Pablo Stern, senior vice president and general manager for technology workflows products at ServiceNow, said the goal is to make it simpler to investigate IT issues involving applications and workflows that have multiple dependencies. That makes it possible to quickly diagnose an issue, determine the root cause and then dispatch the right site reliability engineering (SRE) team to troubleshoot it, he added.

    ServiceNow Cloud Observability uses both open source OpenTelemety agent software to collect logs, metrics and traces in addition to a log management tool that ServiceNow gained last year with its acquisition of Era Software.

    In addition, ServiceNow is employing a Service Graph Connector that makes it simpler to pull data from, for example, OpenTelemetry and Kubernetes clusters into the ServiceNow IT Operations Management (ITOM) without requiring any additional tooling. ServiceNow is also making available its Unified Query Language to make it possible for IT teams to query the data residing in ServiceNow Cloud Observability.

    As application environments become more complex, the need for entire IT teams to have access to an observability platform is growing. While observability has always been a core tenet of DevOps best practices, Stern said it’s clear that these capabilities also need to be made available to IT administrators. In fact, as observability becomes more democratized, it will drive the rise of services operations that will enable IT teams to be better aligned around specific processes, he added.

    Most IT organizations today employ a wide range of tools that enable them to monitor a set of pre-defined metrics. Observability platforms, in contrast, enable IT teams to collect data from systems in a way that can be investigated using a query tool.

    It’s still early days as far as the adoption of observability platforms is concerned, but as the collection of logs, metrics and traces becomes more unified, an opportunity to better converge IT practices based on DevOps principles and the ITIL framework that many IT administrators rely on to manage IT will emerge.

    The challenge, regardless of what methodology is used, is that IT teams need to be able to frame the queries they use to discover the root cause of an IT issue before any potential disruption occurs. In the long term, however, it’s expected that machine learning algorithms will leverage the data collected by observability platforms to automatically identify those issues.

    Regardless of how an IT issue arises, the metric that matters most is how quickly it can be resolved—no matter who in the IT team is tasked with resolving it. The issue is finding a way to provide IT teams with better tools to achieve that goal.

  • Nobl9 Adds Tools to Make SLO Queries More Efficient

    Nobl9 Adds Tools to Make SLO Queries More Efficient

    At its SLOconf event today, Nobl9 added a Query Tester tool that validates whether queries made against the data its service level objectives (SLO) management platform collects will deliver meaningful results. Nobl9’s SLO management platform collects data in real-time from data sources such as New Relic, Datadog and Dynatrace to help customers meet SLOs.

    In addition, a Query Checker tool will now automatically put any new SLO into a testing state to automatically validate that the query against that SLO will run correctly.

    The company has also improved the precision of its calculations to provide DevOps teams with more accurate, reliable and actionable insights into the performance of services and added a Metrics Health Notifier toolset that provides insights into anomalous data being collected from various sources as well as what data might be missing. Nobl9 has also added a Custom Query Delay tool that gives DevOps teams more control over when data is pulled along with access to data source logs.

    Kit Merker, chief growth officer for Nobl9, said the overall goal is to improve the overall resiliency of its platform for managing SLOs.

    Finally, Nobl9 also announced that its platform is now available for the Google Cloud Platform (GCP) in addition to existing support for Amazon Web Services (AWS). As part of that GCP support, Nobl9 made available an educational site that provides example of how DevOps team can take advantage of the large language models (LLMs) that Google is making to improve SLOs or generate an SLO that is automatically converted into a set of YAML files.

    A survey of 314 IT professionals conducted by Dimensional Insight on behalf of Nobl9 found 69% have implemented SLOs, with more than three-quarters of those respondents crediting SLOs for preventing business disruptions. A total of 80% also noted their organization has an increased focus on system reliability, with just over a quarter (27%) saying SLOs saved their organizations a half million dollars or more. A full 95% said SLOs are also driving better business decisions.

    However, 97% also reported it is difficult to manage SLOs because of a lack of reporting, lack of application and environment support and interoperability issues.

    Nobl9 is trying to spur greater adoption of SLOs by making available an open source SLO specification that defines a common interface for constructing SLOs across a Git-based workflow. Nobl9’s platform based on that specification has now been integrated with 24 data sources, said Merker.

    SLOs, of course, are not a new idea. They have been employed as a metric to track the performance of IT services for decades. However, as more microservices-based applications are built and deployed, it’s becoming more challenging to maintain SLOs across applications that have many more dependencies than legacy monolithic applications.

    Ultimately, each IT team needs to provide some sort of objective benchmark that assesses their overall effectiveness at delivering application services. SLOs-as-code are intended to make it simpler to gather the metrics that confirm whether service levels are being achieved. Those SLOs are then tracked and integrated across any set of application services a DevOps team cares to track.

  • New Relic Taps Generative AI to Simplify Observability

    New Relic Taps Generative AI to Simplify Observability

    New Relic has added a tool that takes advantage of generative artificial intelligence (AI) to make it simpler to query the telemetry data collected by its observability platform via a natural language interface (NLI).

    Jemiah Sius, director of developer relations for New Relic, said New Relic Grok will make the New Relic One platform more accessible to a wider range of IT professionals, especially those that haven’t mastered New Relic’s proprietary programming language that enables DevOps teams to identify the root cause of an issue.

    New Relic Grok automatically pinpoints code-level errors in integrated development environments (IDEs) in addition to analyzing code, stack traces and production telemetry to suggest fixes. New Relic is providing this capability via integration with the large language models (LLMs) created by OpenAI, which is providing the foundation for a broad range of generative AI capabilities through Microsoft. New Relic has a long-standing alliance with Microsoft, added Sius.

    Using plain language, anyone can generate a system or app health report complete with anomalies, issues and recent deployments in a way anyone can understand without having to create and filter dashboards, he noted.

    New Relic Grok also makes it simpler to identify instrumentation gaps that can then be addressed using the instructions it provides. It can also set up missing alerts and automate alerts using the open source Terraform infrastructure-as-code (IaC) tool.

    Finally, IT teams can also use New Relic Grok to manage accounts, users and user access, data retention rules, usage, billing and other administrative tasks.

    New Relic previously invested in machine learning algorithms to automate observability tasks. These AI extensions to the New Relic platform are part of a larger effort to extend the reach of the New Relic One observability platform further left toward developers and IT administrators rather than just site reliability engineers (SREs), said Sius.

    However, New Relic doesn’t anticipate that generative AI will eliminate the need for SREs to use its programming tool to solve more complex problems, he added.

    The overall goal is to provide development teams with frictionless access to observability data at every stage of the software development life cycle to reduce mean-time-to-detection (MTTD) and mean-time-to-resolution (MTTR) of issues.

    In general, most organizations are still in the early stages of achieving full-stack observability. A recent New Relic survey found only 27% of respondents have achieved full-stack observability and only 5% claimed they have a mature observability practice in place. A third (33%) of respondents also said they still primarily detect outages manually or based on complaints, the survey found. On the plus side, the survey also found nearly three-quarters of respondents said C-suite executives in their organization are advocates of observability, and more than three-quarters of respondents (78%) saw observability as a key enabler for achieving core business goals. However, more than half (52%) of respondents said they experienced high-business-impact outages once per week or more and 29% said they take more than an hour to resolve those outages.

    AI, of course, promises to improve reliability. At the very least, it could make it easier to identify issues long before a major disruption occurs.

  • The Hidden Configuration Tax Affecting Uptime SLO

    The Hidden Configuration Tax Affecting Uptime SLO

    This article is a preview of a talk by Greg Arnette for SLOconf 2023 on May 15 – 18. To watch this talk and many more like it, register for free at sloconf.com.

    Tax season was just a few weeks ago, and that got me thinking about how frustrating it is to get hit with a surprise tax bill and how that can relate to uptime: Configuration sprawl. DevOps practitioners constantly strive to maintain uptime service level objectives (SLO) while working with complex and ever-evolving infrastructure and applications. Configuration sprawl is a significant challenge many teams face, negatively impacting system performance and reliability.

    My company’s recent survey of 900 DevOps and platform engineering leaders revealed a startling (but not too surprising) conclusion: Most teams struggle managing secrets and configs at scale for infrastructure and applications.

    Where Does Configuration Sprawl Come From?

    As organizations scale their operations, the number of configuration files and settings grows, making it increasingly difficult to manage and maintain consistency across the board. They may have multiple environments, such as development, staging and production environments, often requiring different configurations, leading to duplicated settings and increased complexity.

    Configuration sprawl may also lack standardization as teams adopt different naming conventions, file formats and storage locations without enforced configuration standards, resulting in fragmentation and confusion. Relying on manual processes for configuration management increases the likelihood of errors and inconsistencies.

    Grinding to a Halt

    So you’ve got a lot of configuration–why should you care? Sure, it means more files to manage, but does it really impact your environment? There are a few places where configuration sprawl will negatively impact you. You might see an increase in error rates because as complexity grows, the likelihood of errors and misconfigurations increases, leading to system instability and downtime.

    Increased configuration complexity also means slower troubleshooting. Identifying the root cause of issues becomes more challenging and time-consuming when configurations are disorganized and fragmented. Ultimately this means less agility as you are forced to slow the deployment of new features and updates, impeding the organization’s ability to respond to market demands.

    How to Overcome Configuration Sprawl

    Don’t worry–you can tackle configuration sprawl, keep your SLOs on track and reduce the efforts required to manage and maintain configurations. To mitigate the harmful effects of configuration sprawl, consider adopting these best practices:

    Don’t repeat yourself (DRY)
    DRY config means variable names are declared once. Then flexible variable values are injected across build, deploy and runtime phases as defaults, inheritances and overrides. DRY avoids errors by allowing single changes to propagate where they are needed.

    DRY solves the common problem of forgetting to update critical secrets or parameters and failing to synchronize related changes (e.g., client and server port numbers or rotating a database password.)

    Decouple and abstract
    Decoupling and abstracting config means externalizing config from source code, separating how code and configuration interact, and isolating the specific values (log level) from the configuration interface (how the log level value is fetched.)

    This approach allows the same code to run in different environments using parameterized config. Abstracting config supports a best practice of implementing independent “code and config change” life cycles. Teams that frequently need to change config independently from code should decouple and abstract.

    Centralized config
    Consolidate configuration complexity into a single location because it’s easier to manage complexity within a single location versus multiple locations. Syncing config to the edge (where it is consumed) mitigates downtime. For example, it’s easier to fetch the configuration for service X versus coding all the logic needed to generate the configuration for service X.

    Standardize your configuration management by enforcing consistent naming conventions, file formats and storage locations across teams to reduce confusion and streamline configuration management. Centralizing the configuration and generating that configuration with transformations allows all the places that consume the configuration to be simple. Simple equals easy to understand and troubleshoot.

    Version control & monitoring
    Use a version control system like Git to track config changes to configuration files, making identifying and resolving issues easier. And continue to monitor changes and identify potential problems before they impact system performance.

    Train your team
    Ensure all team members know your best practices and patterns for configuration management. The proliferation of microservices and teams operating in parallel, establishing their unique naming schemes and patterns, exacerbates this problem for DevOps teams.

    Cutting the Uptime Tax

    By understanding the causes of configuration sprawl and implementing best practices to manage and reduce its impact, teams can improve their uptime SLOs and maintain a more reliable, high-performing infrastructure. Remember, investing in the proper tools and processes for configuration management is not only a best practice but also a critical factor in ensuring the long-term success of your DevOps initiatives. Those are tax cuts I can get behind.

  • State of Developer Experience Report Finds Growing API Reliance

    State of Developer Experience Report Finds Growing API Reliance

    Web APIs continue to grow in interest among developer users. APIs can empower new customer experiences and help engineers avoid rebuilding common functions. The technology is also powering microservices and headless architectures that we’ve seen gain more traction in recent years as enterprises become more composable.

    On the provider side, a web API strategy can enable co-creation in partner ecosystems and even open new revenue opportunities for the business. Yet, like any software-as-a-service (SaaS), APIs require great developer experiences to create quick onboarding journeys and easy ongoing maintenance.

    Nylas recently released its inaugural State of Developer Experience Report, detailing the key trends, technologies and priorities that are molding the modern developer experience. The study found increasing reliance on APIs and hopes to increase investment in API-driven technologies. I also met with Isaac Nassimi, SVP of product, Nylas, to explore the reasons behind some of these findings and to get his perspective on the API economy at large.

    Rising Importance of APIs

    As I’ve covered before, the number of APIs in the market has ballooned as more development teams have come to rely upon APIs to power new application functions. The largest companies, those with 10,000 or more employees, have more than 250 internal APIs, according to Rapid’s 2022 State of APIs report.

    Similarly, the Nylas study underlined the growing reliance on APIs. A full 98% of developers said they view APIs as a key contributor to helping them and their team get their work done. And 86% of developers said they expected their use of APIs to increase in 2023.

    According to Nassimi, APIs are something that’s become more and more prevalent over time. For example, in 1998, setting up a web server was pretty cumbersome. But nowadays, a junior developer can accomplish the task (and much more) with a few lines of code, he said. APIs abstract complexity and help leverage external infrastructure so you aren’t constantly reinventing the wheel. “They add more functionality and help outsource labor, thought and cognitive load,” explained Nassimi.

    APIs Bring Developer Experience Benefits

    Another possible reason to shift toward APIs is to handle escalating tool usage. Almost half (48%) of developers said they are either always or often overwhelmed by the number of tools they use daily. Simultaneously, 98% of developers said APIs would lessen the number of workplace tools they use daily. The study indicates that investment in APIs can increase automation and reduce the manual headaches of crafting new features by hand.

    For example, Nassimi describes creating a video transcoding service from scratch in a previous company. The entire engineering team had to dedicate months and months to the process, using muscles they hadn’t ever exercised. After much effort, they ditched their work and ended up just using an API. “It was a really good feeling to delete 200 lines of code,” said Nassimi. “If you do that five times, you reduce all these esoteric things you have to learn how to do in your company by an order of magnitude.”

    In addition to reducing headaches, APIs can also enable speed. For example, it can take over a year for three senior engineers to build an email or calendar integration without the help of an API, the report found. With an API, this integration timeline can be minuscule, said Nassimi. As a result, 95% of all respondents said they would like to see their company invest more heavily in APIs within the next year.

    Developer Experience Improvements

    Developers find speed to be the number one benefit when working with APIs. And to grant this speed, API providers must create a streamlined developer experience (DX). The speed of implementation can make the difference between a good DX and one that is not so good, and a significant contributor to this speed is familiarity. One DX hangup is that the API you use must function like the code you use in your own environment, said Nassimi.

    Although solid documentation and naming conventions should exist in the background, Nassimi shared that forcing developers to learn terminology about the thing the technology is abstracting is a lousy developer experience. Instead, installing SDKs in the unique language of choice, like TypeScript bindings, and using autocompletes to understand the SDK can grant a much better experience.

    Workflow automation and AI can also enhance developer experience, as it frees up time for engineers to be more productive. In fact, two out of three developers would like their company to invest in AI for workflow automation and to curate better user and customer experiences. And 72% of developers said they or their organization are currently using AI for data analytics and making sense of their data.

    However, when it comes to generative AI like ChatGPT and Bard, developers are less enthused—only 14% of developers reported it as a useful area that their companies should invest in over the next year. Although the headlines proclaim that generative AI will disrupt most aspects of modern work, the technology is still new and can produce errors and introduce potential security repercussions.

    Looking Through the Micro Lens

    Developers integrating with APIs will have various backgrounds and priorities. And although speed is an important consideration, providers should also understand where their workloads lie and what particular features they should be driving, said Nassimi. “You want to show off all endpoints and features, but if you narrow down to show users what they need to get started, you will close more deals,” he said.

    For more details, you can download a copy of Nylas’ State of Developer Experience 2023 behind an email gate here.

  • Spotify Adds More Plugins for Backstage Developer Portal

    Spotify Adds More Plugins for Backstage Developer Portal

    Spotify this week added additional plugins for its open source Backstage platform that is used to build developer portals. The new plugins make it simpler to address role-based access and access Insights, a tool from Spotify that tracks Backstage usage trends.

    In addition, Spotify is also enhancing a Soundcheck plugin for Backstage that is used to visualize and track development of software components. Forthcoming capabilities include a no-code interface that makes it possible to programmatically create checks of code without writing any code.

    DevOps teams will also be able to view, export and understand how teams and components are doing compared to established best practices and monitor trends, graphs and historical views, and receive notifications when levels change.

    Finally, Spotify is committing to adding additional integrations with third-party tools and platforms such as Snyk, Sonarqube and the open source Argo continuous delivery (CD) platform.

    There are now five plugins available via a Spotify Plugins for Backstage subscription service, and Spotify said more are planned.

    Meg Watson, a group product manager at Spotify, said Backstage has been gaining traction in cloud-native application environments that are especially challenging to develop using multiple tools. Because of Backstage, developers are not only more productive but turnover has also been reduced because developers are provided with a better experience, she noted.

    There’s a lot more focus than ever on developer productivity, which Backstage addresses by creating a catalog of blueprints that developers can readily consume rather than requiring them to build these capabilities themselves multiple times over. The goal is to create scaffolds that developers can consistently reuse across multiple application development projects.

    Backstage was donated to the Cloud Native Computing Foundation (CNCF) and is now being advanced by contributions from multiple vendors. Much of that focus is on lowering the barrier of adoption for the platform, said Watson. Specifically, Spotify is working on a QuickStart for Backstage edition of the platform that is simpler to install, she said.

    In general, Spotify is trying to strike a balance between centralizing the management of DevOps workflows and the need to enable developers to define workflows that are natural to them versus ones that have been imposed on them, added Watson.

    It’s not clear to what degree Backstage will help fuel the adoption of a shift toward platform engineering to centralize the management of DevOps tools and platforms. The concept of a portal through which developers can self-service their own needs may not be new, but as an open source project, Backstage has made it simpler for many organizations to achieve that goal.

    Just about every DevOps team is now being asked to find ways to help improve developer productivity at a time when many organizations are trying to do more with fewer resources. Regardless of the motivation, however, it’s apparent that DevOps workflows are continuing to evolve and mature.

    In the meantime, DevOps teams would be well-advised to eliminate as many bottlenecks as possible before they negatively impact developer productivity.