Tag: CD

  • We Will Control the World!

    We Will Control the World!

    One thing about centralizing control is that there always has to be someone in charge. That can be a very good thing or a very bad thing. In IT, it doesn’t even get far enough to be a terrible thing.

    Generative AI and increasing integrations have renewed the insistence from the crowd that it is certain they should control every little thing in your (virtual) data center. The idea that a given vendor should be the configuration and management point for a wide variety of tools – their own, competitors and nearby market tools – has been around forever. Vendors love the idea that they have that level of input into your infrastructure and, let’s be honest, many of us love the idea of having it taken care of for us.

    But there are several questions to ask here. First, what is the likelihood that vendor X will have a great level of support for the tools from other vendors in your infrastructure? I mean, before we even get into the idea that this desire for control is not at all altruistic (we’re not talking Stalin here, but we certainly aren’t talking Washington, either), there is the question of how AI will be able to learn about and configure/manage a competitor’s products. Watching wire traffic? Well, if the competitor’s systems are super-well-known and the format of configuration APIs/files are well known—maybe? But this leads us to the real problem.

    My advice to vendors would be to get their own products under control first. We work in a high-complexity field where the tools require a lot of specific domain knowledge to manage, and even more to manage well. Vendors have generally taken for granted that this level of knowledge would be available. Before making that level of knowledge unnecessary for competitors or spiderwebbing into competitor installs, make your tool(s) damn simple to use. That doesn’t mean you have to simplify it—you have the power of AI and are going to focus on configuration and management. So nearly eliminate configuration issues with your own products.

    Example: “Great Security Tool, I just added an API at https://staging-devopsy.kinsta.cloud/rocks. Please lock it down so it is accessible only to users in the WeRule group, then generate a WAF rule to keep it from being used to access more than three records a day, and only allow GET and POST calls.” Most vendors across the industry aren’t even close to this level of simplicity, so why in the world would you trust them to configure/control other vendor’s products?

    There is a trend to make products cover more of the space they are in. This is very big in security right now, but you also see it in DevOps. CI/CD/CDD/VC/kitchen sink (KS)—I’m okay with that, as it does offer us simplified configuration and better cross-app communication. I think this is a better route than trying to serve as command-and-control for an entire tech stack that the vendor does not own. I still want point products and best-of-breed available, but a lot of us just don’t have the time and resources to implement five different products to get the job done—even if they’re “free,” as we all know, they are not free—so one vendor doing an entire chain is great for that use case, and often the shared information about the app, security, the network, etc. can work together better than the sum of its parts. So let’s stay in the “umbrella app” lane and not get carried away with trying to drive in the “manager of the entire world” lane.

    We’re busy, we’re tired and many of us are jaded. So don’t tell us how you can manage other vendors’ tools, tell us how you’re making it easier to manage yours. Thanks, from all of us in IT.

  • Monitor Toolchain Health, Too

    Monitor Toolchain Health, Too

    A funny thing happened on the way to being agile. We adopted a ton of tools to make us more responsive. They really do help us deliver software quickly and efficiently, and they have improved IT processes along with our specific process improvements.

    They also created a toolchain that resembles areas of responsibility of old. That makes perfect sense; we still need to store source code, build applications, deploy somewhere, worry about security, etc. And we need to do all of that faster, which is where automation and the toolchain really shine.

    But you need to keep in mind the overall impact of the toolchain upon the development process. Now the organization is dependent upon those tools. We went from slow and steady to fast and with a large series of steps along the toolchain. Things are better, we’re delivering quality software at a faster pace, but in some ways, the process risks becoming more fragile. If a tool in the chain shows signs of trouble, you need to be aware of it and on top of the issue almost immediately. The price of ignoring issues in the process can be catastrophic disruptions in software delivery. Losing a CI tool is akin to shutting down customs at an international airport: The source of the problem is big, and the downstream problems are show-stopping.

    So while you are tooling your applications, make certain you are tooling the build chain, also. Know when a build that always takes ten minutes suddenly spikes to triple that. Know when a test suite delivers non-breaking errors at an increased rate. Know when space is low on a shared drive that the build tools use. Tool everything, because the software delivery toolchain is now the core of application development and deployment.

    Tooling these systems can also offer a warning of problems in a given application. It is possible that the increase in build time is because there are a lot of previously unused libraries being included … and that those libraries have not been properly vetted by security. It is possible that the disk space on that shared drive has PII test data that needs to be cleansed. There is a lot of good that can come from over-tooling the build chain, and unless you have a bit of tooling that slows the process, there is no downside. It won’t take too long to build a list of what to watch, and modern DevOps tools have enough reporting and APIs that it won’t take long to begin watching those items, either.

    That is where the people part comes in. Don’t ignore the tooling once it is in place. Build a process to regularly review results and take action when necessary. No tooling is useful if the results are ignored; sadly, we have a long history of spectacular failures that started with, “We saw indications, but didn’t think it was important at the time …” So act on results, even if that just means flagging them for review in the next reporting cycle—just don’t ignore them.

    And keep rocking it. The organization thrives on the software you build, deploy and support. It is a massively complex system that has IT staff at its heart, so keep it thriving and watch for indicators of illness in the core toolchain.

  • Identify Needs First

    Identify Needs First

    In several projects of late, I have noted a trend in otherwise rockstar IT departments: They identify an area that they wish to implement or approve and then set out to address it. So, “We need a new CI tool,” spawns a project to choose and implement a new CI tool, for example.

    But there are so many details that these projects are missing from the start that should have been put right into the charter/project kickoff documents.

    Let’s start with, “What CI tools do we currently have in place?” This is the 2020s. How is it possible we are still starting these projects as if there was no solution in the virtual building already doing the job? We need to start from, “Here are the tools that can do this job that are currently running in our environment.” Only then should we move on to, “Here are other options to consider.” But I will humbly suggest that if there is a single tool that can solve your issues, it should be used–and if, of all of the tools running in the organization, none of them can solve all of your issues, then a new product should be chosen that can.

    The same is true for things like alerting. When I am analyzing a given vendor’s solutions, I ask about how many places they can raise issues–it is important that a vendor asking you to purchase their product can report and alert to whereever the organization needs it to. So that could mean Jira or ServiceNow, Slack or email, centralized logging—all of the above and sometimes more (SIEM for security events, for example) need to be supported. But when a given organization is implementing? Choose a reporting method.

    The most common is, “We do Slack, but we also do email. And, of course, we coordinate though issue management like Jira …” Stop. Seriously. This is not about, “Oh, they made it easy, so we just implemented every option.” This is about, “We need to keep track of conversations and resolutions.” If two people are discussing an issue over Slack while a separate team is discussing over email, you’re failing to communicate. Even if there is linkage to ServiceNow or Jira and that is where conversations should happen, we all know that final resolution—but not troubleshooting steps—is what ends up there. So offer one robust and widely available alerting mechanism and let people run with it.

    That type of redundancy is all over IT. Organizations that would not implement per-team email solutions (not servers, solutions) will happily allow for per-team DevOps toolchains. Why? What, exactly, is different? The purpose of a mail server is to deliver mail. Choosing which mail vendor/service to run with is based upon how a given product does so and what additional services come with it. How is that different than, “The purpose of the DevOps toolchain is to deliver applications, and selection is based upon what vendor/product best serves the way our organization does so?” There isn’t a difference; we imagine one. We latch on to one feature that we must have to keep our current favorite, and the compromise is that the next team over can keep their tools, too, because they must have some other feature.

    Evaluate what the organization as a whole needs, identify tricky bits like must-have features and determine if that really is a requirement. If it is, ask why it isn’t a requirement for every project not using that tool.

    And keep moving those apps along! You are building and deploying more apps more securely than ever before. Continue to sanitize the build and delivery process, and continue to crank out apps that make the business hum. But don’t make it part of your job to support fifteen tools that all do the same job. Streamline and improve. And keep rocking it.

  • Harnessing AI in Continuous Delivery and Deployment

    Harnessing AI in Continuous Delivery and Deployment

    In my recent article Revolutionizing the Nine Pillars of DevOps with AI-Engineered Tools, I explained that the continuous delivery pillar is an approach where code changes are automatically built, tested and prepared for release to production; this may be followed by continuous deployment where an approved release is automatically deployed to the production environment. The practices of the CD pillar aim to make releases and deployments less disruptive and more frequent, increasing speed and efficiency.

    In this article I explain how AI can help continuous delivery and deployment processes.

    Continuous Delivery Use Cases

    Orchestration of Test Environments: Teams need a reliable way to prepare and manage various test environments. This includes setting up and tearing down environments, managing versioning and configuration and handling data and resources. AI tools that can help include Harness, and IBM’s UrbanCode, which can automatically spin up and down environments for testing purposes and handle configuration as code.

    Execution of Various Tests: This includes system acceptance tests, regression tests, performance tests, user acceptance tests and security tests. The AI tool that could help here is Parasoft’s SOAtest. Its AI can generate test scenarios, ensure full coverage and automatically adapt to changes in the code base.

    Performance Monitoring of Release-Candidate-Artifact: Tools need to assess the performance of release candidates, identify bottlenecks and compare performance against previous versions. Dynatrace is an AI-powered tool that provides automatic and dynamic baselining of application performance, helping teams to recognize when performance is lagging.

    Automation of Policies for Artifact Acceptance: Rules and policies need to be defined and enforced regarding when an artifact is ready for release. Tools like Harness use machine learning to analyze past deployments and test results to predict the success of future deployments and automate the approval process.

    Continuous Deployment Use Cases

    Orchestration of Deployment Environments: Teams need to manage deployment environments, including setup, versioning, configuration and data management. Spinnaker is an open source tool that uses AI to manage multi-cloud deployment strategies.

    Execution of Deployment Strategies: This includes executing strategies such as Blue-Green, Canary and Feature Flag gradual rollouts. A tool like LaunchDarkly can utilize AI to manage feature flags, analyze usage data and roll out features gradually.

    Testing in Production: Testing the new release in the live production environment is crucial. A tool like Sentry can use AI to monitor application use in real-time and identify and alert teams to issues faster.

    Restoring Production to a Prior Version: When a problem is detected with a new release, teams may need to revert to a previous stable version quickly. StackStorm is an event-driven automation platform that uses AI to define workflows, allowing you to quickly roll back to a previous version in case of failure.

    Updating Documentation and Configuration Management Databases: It’s crucial to keep records up-to-date. A tool like Guru uses AI to automatically update and keep track of documentation, ensuring that it’s always accurate and up-to-date.

    Challenges When Transforming CD to use AI

    Here are some challenges organizations might face when transitioning their existing CI/CD pipelines to incorporate continuous delivery, continuous deployment and AI tools, along with possible solutions:

    Lack of understanding of AI capabilities and uses: One of the first challenges that organizations might face is the lack of understanding of what AI can and cannot do. Some team members may have unrealistic expectations or fear that AI will replace their jobs. Regular training and workshops can help teams understand how AI can enhance their roles rather than replace them. Clear communication about the goals and benefits of incorporating AI can also help allay fears.

    Data privacy and security concerns: Using AI tools often involves handling a large amount of data, some of which may be sensitive or private. There are also concerns about the security of AI tools themselves. Establish a robust data governance policy that includes procedures for handling sensitive data. Also, when selecting AI tools, prioritize those that are known for their robust security measures.

    Complexity in integrating AI tools with existing systems: Not all AI tools will seamlessly integrate with existing systems and infrastructure. This could lead to technical debt and instability in the pipeline. Prioritize AI tools that have wide integration capabilities. Engage the tool’s support team or hire consultants if needed to ensure the integration is smooth and stable.

    Resistance to change: The transition to a new way of working can be met with resistance, especially if team members are comfortable with the existing systems and processes. Change management strategies can help ease the transition. This can involve regular communication about the benefits of the change, training and providing support during the transition.

    Cost of AI tools and training: AI tools can be expensive, and there’s also the cost of training team members to use these tools. Start with a cost-benefit analysis to understand if the transition is worth the investment. Look for AI tools that offer a good balance between cost and features. Consider phased implementation to spread the costs over time.

    Dependence on vendor support: Given the complexity of AI, there may be a high dependence on the AI tool vendor for support. This could become a problem if the vendor is unresponsive or ceases to support the tool. Choose vendors who have a good track record of providing strong customer support. Also, consider tools that have a strong community around them, which can provide additional sources of help and support.

    Summary

    Navigating the landscape of continuous integration, continuous delivery and continuous deployment (CI/CD) pipelines can be complex, but it’s a vital process that’s being revolutionized by the introduction of AI tools. From orchestrating test and deployment environments, managing testing and deployment strategies, to automating delivery tasks and quality checks, AI is playing an increasingly critical role in streamlining and enhancing these processes.

    However, as with any significant change, transitioning to AI-enhanced CI/CD pipelines is not without its challenges. This includes understanding AI’s capabilities, data privacy and security concerns, technical integration difficulties, resistance to change, cost considerations and vendor support dependencies. By identifying and addressing these challenges head-on, organizations can effectively leverage AI’s capabilities, bringing increased efficiency, reliability and value to their software delivery pipelines.

    With proper planning, continuous learning and strategic use of AI tools, organizations can not only overcome these challenges but also thrive, gaining competitive advantages in their software delivery processes. The future of CI/CD is undeniably entwined with AI, and embracing this evolution is an exciting journey that holds the promise of transformative benefits.

  • Forget Change, Embrace Stability

    Forget Change, Embrace Stability

    DevOps is about change. It was introduced to accelerate change, and it has opened up a whole new world of tools and possibilities that are all based on the premise of changing IT. Starting with Agile, we now have DevOps variants and subprocesses that run the entire software development life cycle (SDLC) process.

    Want to make more frequent updates that can regularly show demonstrable progress to business owners? Agile and CI! Need to get the app settled in its inevitable home with automation? You need CD! Want testing and zero-day protections? Start with DevSecOps built into the CI/CD process and add in application and API security! Want to fully automate deployment in the whay builds are automated? Grab a GitOps tool and let’s goooo! The list goes on and on. Literally. We’ve almost got more tools than we have names for spaces at this point.

    And all of those tools introduce change. Most of them introduce a massive amount of change. You don’t get to GitOps without breaking some of your existing processes and reworking them. We have fallen in love with change in IT.

    But there is a funny thing that’s happened along the way; I’ve mentioned it before, but never from this perspective. All great changes require time for consolidation. Be it stock market swings or revolutions, afterward, there is a period of consolidation. In IT, containers are a good example. When they became capable of actually running workloads, IT experienced a period of time when making all the various machines run in VMs was the thing. People weren’t (generally) looking for the new big change in the process, they were just glad that those one-off machines that had custom hardware could now be moved to something more manageable and that spinning up servers could be done in days or weeks instead of the months it used to take when you had to order hardware. But that consolidation created organizational standards and stability that the organization enjoyed for years – even after containers rose to challenge VMs and some organizations even today.

    So every once in a while, take a breather, consolidate gains, create standards that have meaning for your organization, then continue on with whatever the shiny new change is. Make sure that once you have identified, for example, the need for code scanners that can work within or that are complementary to the integrated development environment (IDE), that you have one scanner—or one scanner per programming language—to maintain and monitor instead of one tool per project. That saves time and creates stability, both of which will serve the organization moving forward.

    Change is here to stay; it’s the only constant. IT is caught in a cycle of change followed by change, with the next wave driven by generative AI. But in times of change, seek stability and, in times of stability, seek changes that can have a positive impact. So go standardize and organize before the next wave hits.

    Mostly, the changes will serve the organization by freeing up your time and letting you spend more of your time rocking the IT world. Those systems stay up because of you. More time is more innovation, and more stability means less burnout. All of which is good for the organization.

  • The Fallacy of Continuous Integration, Delivery and Testing

    The Fallacy of Continuous Integration, Delivery and Testing

    We know that continuous integration and continuous delivery (CI/CD) have become a DevOps best practice. And many have learned that by adding continuous testing (CT), they can create a virtuous loop, ensuring perpetual code quality and security. They’re not wrong.  Yes, testing continuously is good practice, but some incorrectly equate the concept of continuous with monitoring, thinking they’re able to see everything that could eventually affect a customer’s experience (CX) with their software application. That’s a fallacy. Many of the most glaring blind spots that wreak havoc on CX exist outside the environment that CI/CD/CT is able to “see.” CI/CD/CT, even when supported by application performance monitoring (APM), doesn’t and cannot monitor for and alert to blind spots across the internet stack. And there are many. Only internet performance monitoring (IPM) can take CI/CD/CT as input streams and shift to continuous resilience as a measurable outcome.

    How CI/CD/CT Typically Works

    As customer requirements evolve, software development teams write code that triggers a new software build. In the continuous model, each new build moves to a runtime environment for integration and quality assurance and, ultimately, is deployed to end users across public or hybrid clouds. This is an application-specific view of the world.

    Since quality assurance is part of the typical continuous process, monitoring is an assumed benefit or outcome. This assumption might be because many now complement their CI/CD/CT processes with APM. What is achieved with APM in this approach is testing, not monitoring. While testing provides a better picture of performance at the application level, it doesn’t provide full visibility into the blind spots that exist between the application layer and the end user.

    APM Isn’t Enough

    While APM tools focus on code—including everything that negatively affects an application such as database wait times and inefficient code—IPM focuses on the internet, which has become your network. IPM monitors what impacts your customers, workforce or application (API) experiences across the internet.

    While APM and IPM may seem similar, they are nothing alike. They may share common tools, such as using synthetic agents, but the difference is in what they monitor and how. In this age of highly distributed applications, that difference means everything.

    Where APM can “simulate” what a user may see from cloud location to cloud location, which can confirm application performance, this simulation isn’t true or comprehensive monitoring because it is blind to network-related problems with end-user experience, which are vast. Customers and employees each have a unique internet performance fingerprint and access applications from myriad devices across multiple locations and network infrastructures, not through the individual cloud locations that APM synthetics simulate.

    IPM approaches synthetic testing more holistically, replicating multiple user journeys across a much more complex Internet infrastructure to proactively flag performance issues before they impact end users. This true monitoring takes place 24/7, not as point-in-time tests.

    Forensics vs. Monitoring

    With APM and similar observability solutions that embrace real user monitoring (RUM), it’s important to focus on what monitoring really means. APM passively collects data that is a proxy for user experience, enabling postmortem assessment of causes and possible fixes, but after the fact. Observability looks at logs, metrics and traces and provides an assessment of whether users are having a good experience or why they did not. Neither solution by itself will prevent poor experience, and many fixes won’t take place until a postmortem analysis feeds the next iteration. The poor experience has already done its damage, and the costs are already sunk.

    With thousands of potential blind spots across the globe that can slow or disrupt applications, from ISPs and wireless carriers to CDNs, it’s not hard to grasp the increasing and differentiated importance of IPM. Where the combination of CI/CD/CT with APM can proactively prevent deployment of poor code that would affect user experience, only IPM is continuously replicating – even more valuable than monitoring–millions of experiences across a dynamic multi-node network where the unexpected is now expected.

    Internet Resilience Vs. Retrospection

    Businesses today understand the importance of resilience: It’s a board-level concept. Any business that has lost market share and market cap during sustained downtime or poor user experience understands what’s at stake: Millions if not billions of dollars. IPM is squarely focused on “internet resilience,” and its pillars are availability, reachability, performance and reliability. For businesses reliant on the cloud and employers operating with highly distributed workforces, retrospection alone is a risky proposition. Now that the internet is the new corporate network, internet resilience is a new organizational mantra.

    The Takeaway

    Whether your organization embraces CI/CD/CT already or is rethinking its approach to DevOps, this article should give you pause. Your job–perhaps as part of a larger team–is to catch performance issues and potential disruptions with your application before client impact is realized. Without IPM, only part of that job is being done.

    Why is this so important? Even conservative estimates put the cost of an outage at $6,700 per minute (Gartner) and the “cost of slow,” where something causes an application to perform poorly and it goes unnoticed, is likely significantly higher because it happens all the time. The fallacy isn’t that there’s value in CI/CD/CT–there is. The fallacy is that you’re in control of the entire CX of your internet-deployed products without monitoring the internet stack and ensuring internet resilience. Only by adding IPM can you have this assurance.

  • Raise Those (Feature) Flags

    Raise Those (Feature) Flags

    I’ve written about feature flags before, but I think it’s time for a solid bit of advice: “If you are not yet using feature flags for DevOps, it is past time to reconsider.”

    I say this because feature flags have been a slow-growth item, with a few spikes here and there in usage, but mostly on the development side. While many of you are using feature flags to do dynamic toggling of features–definitely an advanced use case, but still very development heavy–few organizations are using them for deployment options–blue/green without separate code, environment changes for different targets (deploy this to the cloud provider, but do this for local installation type stuff), etc.

    Feature flags are starting to be supported by DevOps tooling for this precise use case, and it is glorious. While the ability to say, “run this code,” and if the new code fails, change something like an environment variable to say, “stop running that bit,” is huge. We have enough time with this process that the industry is well aware of the need to clean up as soon as the correct code options are known, and that is the biggest weakness; it makes spaghetti code that must be cleaned up. But the benefits are well worth a bit of cleanup at the end of a project, even if far too many of you don’t bother.

    Now we can use them in deployment, and in full life cycle toolsets, we have them available for both–so a set of source code can be gated behind a feature flag and deployment can use the same feature flag. This enables the ability to say, “If our FF environment variable is ‘TestDeployment,’ connect to ‘Test Database’ … then use the same variable to say, “If our FF environment variable is ‘TestDeployment,’ spin up the container with test data instead of real data.” This simple example is something that could be (and has been) done manually, but doing it this way makes it part of the process. Automated testing can then follow without the need for human intervention, and from “commit” to “test results,” IT can focus on other things instead of spinning up the right containers and making sure the right DB is accessible.

    After all, we are in a time of massive automation, which should not worry us. There are globally too many IT positions and too few people to fill them; automation should be exercised as much as possible to address that simple fact. And there is a good chance that you are already using tools in the DevOps toolchain that support feature flags–a growing number do support them. If not, it is easy to research and find a good tool to do so and while it doesn’t offer as much functionality, a stand-alone feature flag product will still offer more options than not having any.

    And keep rocking it. The applications you wrote years ago are still powering the enterprise, even if you are at a different organization today. We leave a legacy of applications behind us that is a testament to quality. Keep leaving quality, and use all of the tools you reasonably can.