Author: Mitch Ashley

  • DevOps Chat: Mayhem Testing With ForAllSecure

    DevOps Chat: Mayhem Testing With ForAllSecure

    Secure software depends on people finding vulnerabilities and deploying fixes before they are exploited in the wild. This has lead to a world of security researchers and bug bounties directed at finding new vulnerabilities.

    As dedicated as security researchers are, there is a vast ocean of software in existence, waiting for someone to find and exploit the next security vulnerability for profit or nefarious uses. With autonomous vehicles on the horizon, is there an autonomous solution to finding and fixing software vulnerabilities?

    Enter DARPA Cyber Grand Challenge winner “Mayhem,” created by a team of researchers from Carnegie Mellon University who spun out security startup ForAllSecure. And they have a BHAG (Big Hairy Audacious Goal). “Our vision is to check the world’s software for exploitable bugs so they can be fixed before attackers use them to hack computers.” Mayhem has moved on from capture the flag contests to observing and finding vulnerabilities in DoD software and is working its way to corporate systems.

    In this episode of DevOps Chats we talk with David Brumley, For All Secure co-founder and CEO, and CMU professor about the technology behind Mayhem, how it observes software as it executes and injects changes to effect and observe new and potentially exploitable behaviors. More information about Mayhem is also available at www.forallsecure.com.

    As usual, the streaming audio is immediately below, followed by the transcript of our conversation.

    Transcript

    Mitch Ashley: Hi, everyone, this is Mitch Ashley with staging-devopsy.kinsta.cloud, and you’re listening to another DevOps Chat podcast. Today, I’m joined by David Brumley, CEO at ForAllSecure. David is also a professor at CMU. He’s currently on leave. The topic that we’re talking about today is called Mayhem behavior testing—it ties back to David’s work at CMU. So, David, welcome to DevOps Chat.

    David Brumley: I’m happy to be here.

    Ashley: Great, we’re happy to have you on the podcast. Would you start by just introducing yourself a little more fully—your background, maybe a little bit about some of the research and how that led you to form ForAllSecure?

    Brumley: Yeah, absolutely. When I got out of undergrad, I was a computer security officer. My job was to chase intrusions on the Stanford network and try to help people fix them. And at the time, I got pretty frustrated with the idea that we were always behind attackers, that I couldn’t find vulnerabilities first or get the fixes deployed.

    I actually went back to grad school, got a Ph.D. and really made that my work since 2003 on how do we go about finding vulnerabilities before attackers and just as importantly, how do we get those fixes fielded? Because it’s not just about finding the vulnerabilities, it’s about how quickly we can get those in place.

    I’ve been working on that, I’m a tenured professor at Carnegie Mellon, and for the last three years, been working on commercializing it.

    Ashley: Excellent. Well, an entrepreneur and a professor, and it’s great to see that your research has led you to kinda bring this out to market. That’s exciting. So, tell us, what is Mayhem? What is Mayhem or behavioral testing?

    Brumley: Yeah, what Mayhem does, we call behavioral testing. So, behavioral testing is about watching an application as it executes and learning from that execution and then trying to come up with a new input that would cause the application to do something different. And you do this again and again and again, like hundreds of times per second, with the idea that if you can learn the behaviors of a program and you can start driving it to new behaviors, you’ll, one, come up with a test suite for the program, that’s important for the DevOps part and how you get things out quickly, and second, things like vulnerabilities and exploits that trigger vulnerabilities are just triggering a particular behavior. So, you can automatically generate inputs that trigger these. At one point, people were saying we were automatically generating exploits with this technology.

    Ashley: Interesting. I’m pretty sure that it’s not the same thing, but it sounds like a more thoughtful and learned approach to what we know as chaos testing from the Chaos Monkeys from Netflix, et cetera. But you’re really picking specific behaviors, building a profile of the application, and then looking for new behaviors that you can test against it?

    Brumley: That’s what we’re doing, and so, there’s really kind of two technologies people may have heard of. One is called fuzzing. So, fuzzing is about running the application again and again and again. And it uses heuristics to pick those inputs, and one of the things we built is a way to monitor that application as it runs and to inform the fuzzer on how to come up with new inputs that would get different behaviors.

    The second that we did is, we really took a page from surprisingly formal verification. So, the formal verification was trying to prove a program was safe. What we tried to do is use those same techniques to prove where a program is unsafe, and our proof creates an input that triggers the unsafe property.

    So, we use these two techniques—fuzzing, symbolic execution and a few others and a portfolio approach to try to come up with behaviors that are exploitable.

    Ashley: Interesting. Do you do this in a production environment, in a test environment, in both? How do you approach this?

    Brumley: For a product, we always do it in testing. We think it’s important to have these techniques out there in testing so that you can find them before attackers. We have done some simulations in production. One of the things we participated in was, DARPA had a big autonomous Cyber Grand Challenge. So, we fielded it there, but we think the market would be in the testing environment.

    Ashley: Mm-hmm. It would definitely make sense—makes more sense there. Say a little bit more about fuzzing and how that exactly works. Is it the software that’s doing the variations, or are you introducing variations to that from what you learn?

    Brumley: We’re introducing the variations, so the idea is, an exploit for a program is just an input.

    Ashley: Mm-hmm.

    Brumley: And, you know, test cases are just inputs. So, you run the program on an input and you watch how it executes. Like, you see it executes this system call, that system call. It covers these branches. And you use that to learn a new behavior or guess a new input. And then the fuzzer will use a heuristic to pick that new input, with the idea it should trigger a new behavior.

    Ashley: Step back just for a moment, you’ve talked a little bit about Mayhem and the technology behind it. You’ve stepped out of CMU to set up this company, ForAllSecure. How did you get that started? Are you working with a particular private sector, government sector? Sounds like something that might be interesting to the government side of things, too.

    Brumley: Yeah, we got started by really participating in this, as I said, DARPA project. So, they have Grand Challenges every so often. One was the self-driving car contest. This was, like, a self-driving car for a computer security contest. We won that and that gave us our first $2 million, so that was, like, seed funding from the government.

    Ashley: Excellent! Wow, that’s awesome.

    Brumley: Yeah. And so, we’ve been working to bring it into the government so it’s used in pretty much every service and in the IC right now to try to look for vulnerabilities in weapons platforms as well as in software the government might be using to try to fix it before attackers can break into it.

    Ashley: Mm-hmm.

    Brumley: About, oh, six months ago, we then went and raised money from NEA to help transition this from just a government tool to something in the enterprise section as well.

    Ashley: Mm-hmm, so you’re also beginning to work with the private sector, then?

    Brumley: Yeah, we’re working—starting to work with the private sector. It’s really kind of an interesting difference between the two.

    Ashley: Oh, there’s a lot of difference. [Laughter] Yes. Highlight what you see the difference as. I have some experience with both, also.

    Brumley: Well, I think that the private sector, it’s harder for them to put a value on finding a vulnerability and fixing it quickly while, in the DoD, it’s really easy. That mission didn’t succeed, which is a big deal.

    Ashley: Mm-hmm.

    Brumley: So, I think that’s one difference. I think the second difference as we see going to market is, in the DoD, they care a lot about checking legacy systems, because they still have to maintain them. And if you think about it, lots of vehicles, airplanes, trains—things like that. And in an enterprise, they care about things that were developed in the last two or three years only. And so, they’re just kind of products in different parts of the life cycle.

    Ashley: Yeah, interesting. You also have many layers of different parties involved in the government sector—contractors, subcontractors, primes, et cetera. So, you’re dealing with lots of layers of companies that are working as part of the government team.

    Brumley: Yeah, working with the government is definitely, it’s a chore until you get into the groove and you realize a lot of the, what you perceive are barriers, are actually just regulations put up to protect taxpayers from complete abuse. So, it takes a long time, it takes about nine months to get anything really going from start to scratch in the government. But once you kind of get that cycle, it goes forward pretty quick.

    Ashley: Yeah, that was my experience, too. It’s an investment, but once you get going, it’s great.

    Brumley: Yeah.

    Ashley: So, talk a little bit about where you are in the product life cycle? Do you have product in the market? Is Mayhem publicly available? Are you alpha, beta? Give us an idea of that.

    Brumley: Yeah. So, within the DoD space, we have Mayhem that you can buy starting in July this year—so, pretty close. So, we have a number of beta installs, people are happy, and we’re gonna be switching over to general availability pretty quick.

    Ashley: Mm-hmm.

    Brumley: In the commercial market, we’re taking it a little bit slower, because some of these differences I’ve discussed. They use different application stacks, they’re often concerned with much newer software than old software, they often use different languages.

    And so, that, we’re looking for design partners. People who wanna take Mayhem, what’s working in the DoD, what’s working in places like aerospace, and see if it works for them and figure out what we need to do to make it a really good fit.

    Ashley: Mm-hmm. So, you’re spending a lot of time with customers and potential customers really learning what the commercial sector is looking for or can benefit from.

    Brumley: We’re learning what they’re looking for. We’re also trying to understand, like you said, when you go to market, it’s kind of interesting in the DoD who the buyer and who the user are and trying to understand that in the commercial space. Like, it’s kind of fascinating to me in DevSecOps, it’s almost always the security team that’s the buyer but it’s the development team that’s the user.

    Ashley: Mm-hmm. Yeah, exactly. And that’s an interesting dynamic there, it’s, you know, the DevSecOps conversation, one that we helped facilitate bringing together through a lot of the activities at staging-devopsy.kinsta.cloud.

    I’d love to hear a little bit about, so you said you’re not commercially available yet in the commercial or private sector, but you are entertaining companies to work with? Is that true?

    Brumley: That is. So, we’ve worked with a large aerospace vendor that makes airplanes. Can’t say much about that. We’ve been able to really help them find some flaws in some software that they use as well as build up test suites for them so that when the developer does push a fix, that can be more rigorously tested, working with a Fortune 100 company, and then some IT infrastructure people.

    Ashley: Mm-hmm. I could see, for example, data center providers being another place that it’s gonna really benefit from.

    Brumley: Oh, absolutely. Anyone who has something that’s extremely high value like a web server that, if it gets compromised, your entire business is at stake.

    Ashley: I’d love to hear your thoughts about how you fit this into the DevOps or the DevSecOps pattern or cycle. How does this happen? Where do you build it into?

    Brumley: Yeah, so, from what we can see in research right now in a lot of DevSecOps, it’s really kind of a primitive stage where people want to run a scan before they field software. To me, this is insane. I don’t know any attacker who runs a commercial product and a scan and then says, you know, “Here’s the new zero day” or, “Oh, better not go after that, the scan was okay.”

    Instead, what they do is, they’re always trying to break into that software. They’re always trying to learn from what hasn’t succeeded and come up with attacks and succeed. So, Mayhem is like that. So, the idea is, it’s more of asynchronous testing. You’re gonna push as part of your DevOps cycle through, you know, things like making sure you’re not using old versions of libraries. But then, you’re gonna launch a process that, for the lifetime of that product, tries to hack it in the background.

    I think that’s really kind of the conceptual shift that we think needs to be made is to stop thinking of security as a scan you do once, and actually expect it to work, to something that’s always going on in the background.

    Ashley: I’m gonna shift just like we do with testing, we’ve moved into continuous improvement, continuous testing do the same for DevSecOps where this is running maybe on multiple versions of the product and test environment all the time.

    Brumley: Absolutely, as well as all the dependencies. If you look at Google—so, Google runs fuzzing, one of the technologies that I talked about earlier, on Google Chrome. And they found 12,000 new vulnerabilities in the last three years, and each one was accompanied by a test case, so zero false positives. And one of the things this has allowed them to do is get ahead of attackers, because they’re finding those flaws so quickly with their automated, always on infrastructure that I’ve talked to people who try to participate in bug bounties. And by the time they report it, it’s already fixed.

    Ashley: Mm-hmm. I mean, aren’t the bad guys also doing fuzzing themselves, so you’re at a disadvantage if you’re not?

    Brumley: Absolutely. Every top notch hacker I know uses fuzzing extensively as one of their techniques. I mean, this is assuming you kind of do the base, like, did the person forget to set a password? Always check for that first, right? But after that, after that base level, you’re gonna do fuzzing, especially after high value targets.

    Ashley: Talk a little bit about where you see things going next. You’re kinda testing with the commercial market, you’re heavily engaged with the federal market. What happens next in terms of the product roadmap?

    Brumley: Well, the next question that we have in terms of the product roadmap is, how do we close that life cycle from finding a vulnerability to having a fix fielded. So, some app site companies, what they wanna do is always add on that next language support so that they can support a larger and larger set of languages.

    Ashley: Mm-hmm.

    Brumley: For us, we’re really looking at C, C++, Java, Python—the core languages used by infrastructure Go and Rust. And we want to be able to do deep analysis in these languages and then be able to help automatically suggest fixes and actually, for DARPA, we proved that the computers could automatically patch. So, it could automatically patch binary programs, and it could assess whether that patch would have any business impact, and field it, if not. And so, that’s really what we’re trying to do, and we’re going language by language as opposed to trying to cover all the languages at once.

    Ashley: Interesting. What’s your top pick in terms of languages you’re looking at first?

    Brumley: So, the top one that we started with is C and C++. And that’s been a bit surprising to people, because it’s an older language. I can tell you why we picked it.

    Ashley: Okay.

    Brumley: The reason that we picked it is, regardless of what language you choose, you’re gonna be calling out to C and C++ code. In Java, you have JNI calls. In Python, you have, you know, as part of TensorFlow, if you’re doing machine language, everyone goes to C/C++ or a compiled language when they want performance.

    And so, if you wanna analyze these applications completely, even if it’s Java or Python, you have to be able to analyze those components as well. So, we’re kind of working our way up.

    Ashley: Yeah, I was gonna say, it sounds like you’re working way up the stack into the scripting languages.

    Brumley: Absolutely. The second reason is, when you look at really critical infrastructure out there, it’s still primarily developed in C and C++. So, you know, I, like many people, have probably put up a Flask or Python website, it’s really quick to do, it’s online—you definitely don’t want it hacked. But when you start talking about a car or an airplane or a power plant—these are also on C/C++. And so, just from a safety point of view, you have to cover those.

    Ashley: Interesting. Do you foresee, looking kinda down the road—I’m not asking you to announce anything officially, but do you see that you might be working with either the patch management companies, the vulnerability scanning companies as a complement to them? Would this be something that potentially displaces them? How do you see that future unfolding?

    Brumley: Yeah, I think people who detect known vulnerabilities are complementary to us. So, the way I think of it is, there’s really two types of security companies out there. There are those who find new vulnerabilities, and there’s those who check for old vulnerabilities.

    So, you have, for example, Tenable, which is a network scanner that goes and looks for known vulnerabilities. You have software component analysis, which looks for known vulnerable versions of libraries and other things, right? So, those are pattern matchers in some sense.

    We’re not one of those companies. That’s someone that you would partner with. What we’re doing is, we’re finding new vulnerabilities. And so, this is more like a typical SaaS or DaaS type solution.

    The other sort of people that we’re looking to partner with are the patch management. So, one of the cool things that we can do is, when we find a vulnerability, we automatically are building a test suite as well, just part of the process. So, when there’s a patch we can replay it, and we can tell you things like, “Hey, have any of the previously passing test cases stopped working? Has performance dropped?” as well as, of course, where the bug’s fixed.

    Ashley: I definitely can see the partnership alignment there. If you had to paint a future of what success looks like for, both for you and ForAllSecure, what would that be?

    Brumley: I think we’re really motivated by this vision that we wanna automatically protect and check the world’s software. And so, we wanna close that cycle and make it autonomous. And we’re making, actually, pretty careful design decisions as we do that.

    So, I think a lot of people talk about, you know, big startups, billion dollar businesses—that’s not really our goal. Our goal is, we want to make it so the time from when we detect a vulnerability to there’s a patch that’s tested and ready to field is seconds. That, to us, is success.

    Ashley: The billions will come later, right? [Laughter] That’s if you’re passionate about that core problem.

    Brumley: I think so. Because this problem, we were talking about DevSecOps. Like, DevSecOps, the problems facing DevSecOps face every system administrator if we go back in the world. In Silicon Valley, we think of SREs and DevSecOps, but there’s this huge nation full of just system administrators who wanna know, “If I update it, is it gonna break? Do I need to update it?” And solving those problems for those people is what’s important for us.

    Ashley: Do you have any demonstrations coming up, either online or you’re gonna be at some conferences? Is there a way for folks to see this yet?

    Brumley: Yeah, so, we’re gonna be at Black Hat and we’re happy to demo it there. We’ll have a booth and a room. We also have a number of online videos. If you want to see the fully autonomous, in its splendor system, there’s definitely videos of the DARPA Cyber Grand Challenge as well.

    Ashley: Going back to the original challenge.

    Brumley: That was the original vision. I mean, when I talk to people about this technology, what really lights up their eyes is not finding security bugs, it’s that there’s not enough people to do the work. And by being autonomous and focusing on that value proposition, you’re really focusing on the pain point, which is, how do I do things that there’s just not enough highly skilled people to do.

    Ashley: Well, I applaud you for what you’re doing. You know, there’s yet another scanner, yet another anti-virus something. It’s great to see folks that are really doing some innovation, kinda thinking about the problem differently and really trying to solve it in a more systemic way, and that’s one of the things that I see that you’re doing. It isn’t just, we’re trying to test for these kind of behaviors, but it’s also in this ecosystem of how patches happen and things can happen dynamically, automatically.

    Brumley: Absolutely. I mean, we’ve had a number of customers who want us to support 10 different languages, but there’s no product in the world that can add 10 languages to support and also be really good at each one of them. For us, our main value proposition is, for everything we do, can we have zero false positives, can we make sure it’s actionable before we report it at all?

    Ashley: Awesome! Well, we’re—[Laughter] the time always flies by on these podcasts. Is there any last parting thoughts that you wanna share with us before we wrap up?

    Brumley: Well, thank you for having me and, as I said, we’d love to talk to you at Black Hat 2019.

    Ashley: Great, great. Well, I’ll be there as well, so we’ll get to meet up there. We’d love to have you on a future podcast as things develop and hear more about how things are moving along for you and ForAllSecure.

    Brumley: Thank you very much. Have a great day.

    Ashley: Well, another DevOps Chat podcast has flown by, as they always do. I’d like to thank David Brumley, CEO, at ForAllSecure for joining us today. Thank you, David.

    Brumley: Thank you.

    Ashley: And I’d like to thank you—you, our listeners—for joining us, of course. This is Mitch Ashley with staging-devopsy.kinsta.cloud. You’ve listened to another DevOps Chat. Thank you for joining us today.

    — Mitchell Ashley

  • DevOps Chat: Chaos Testing With xMatters

    DevOps Chat: Chaos Testing With xMatters

    Introducing chaotic, unpredictable test software into your methodical testing regime is a good idea, right? Yes, it’s a branch of testing called Chaos Engineering, or Chaos Testing. Netflix’s Chaos Monkey famously introduced many of us to the idea that resilient systems, networks and software become more resilient and less brittle if we use chaotic testing methods to find their weak points before customers do.

    An innovative engineer at xMatters launched a new open source chaos testing tool named after H.P. Lovecraft’s nightmarish character Cthulhu. Cthulhu is designed to test across multiple cloud providers, initially supporting Google Cloud with plans to support Amazon Web Services. It’s open source, free to use, looking for more contributors, and is available on GitHub.

    We are joined on this DevOps Chats by Tobias Dunn-Krahn, CTO, and Gabrielle Gasse, lead engineer on the Cthulhu open source project, both at xMatters.

    As usual, the streaming audio is immediately below, followed by the transcript of our conversation.

    Transcript

    Mitch Ashley: Hi, everyone. This is Mitch Ashley, with staging-devopsy.kinsta.cloud, and you’re listening to another DevOps Chat podcast. Today, I’m joined by Tobias Dunn-Krahn, who is CTO at xMatters, and also Gabrielle Gasse who is lead engineer of a software project we’re gonna talk about called Cthulhu, and that’s our topic today is Cthulhu.

    Now, if you’ve read the H. P. Lovecraft books, you know at least the fictional character that we’re talking about, but we’re gonna get into this software and find out a little bit more about what it does.

    Tobias, Gabrielle—welcome to DevOps Chat.

    Tobias Dunn-Krahn: Thank you.

    Ashley: Well, let’s start by introducing yourselves. How about you, Tobias, if you would start, just introduce—tell us a little bit about yourself, what xMatters does, what you do there as CTO.

    Dunn-Krahn: Sure. So, right, my name is Tobias Dunn-Krahn, CTO at xMatters. What that entails is purview for development operations, quality and product strategy. So, that’s my role.

    What does xMatters do? xMatters is a digital services availability platform. What that means in a practical sense is that xMatters helps teams that are responsible for digital services provide a high level of up time and reliability for those services. So, what that means is filtering unwanted signals for those services, engaging resolvers when those signals are relevant, facilitating collaboration during the resolution of an incident and, as well, integrating tools in a tool chain to eliminate any manual effort that’s involved in resolving an incident so that it can be done in a very timely fashion.

    So, that’s what xMatters does in a nutshell, and what I do there.

    Ashley: Interesting, yeah. I’m interested to see how Cthulhu fits into that. Gabrielle, would you introduce yourself and tell us a little bit about what you do at xMatters?

    Gabrielle Gasse: Yeah, of course. So, I’m the lead engineer working on Cthulhu, our chaos testing tool. On a day to day basis, I’m also just a Java developer working on the system itself.

    When I came into the company about a year ago, I was tasked to verify the resiliency of our system as we moved it to Google Cloud. And so, I did some research on existing tools and realized that we need something a bit different that was available already. And so, that’s how I came to start building Cthulhu.

    As—this is not the core product of xMatters, and we do want to promote the use of chaos engineering as a day to day practice for everyone, we decided to release the tool itself in open source.

    Ashley: Great. Maybe we should start with that, for folks that don’t know what chaos engineering or the theory behind it is, do you wanna say a little bit about that, Gabrielle?

    Gasse: Of course. So, the idea behind chaos engineering, if you think of a distributed system made of multiple microservices, you may end up with a fairly large amount of small VMs or containers running in parallel. And, as things are deployed in the cloud, sometimes failure happens. It always happens.

    And so, chaos engineering takes that as a premise, and it says that if you expect to fail all the time, then you will build your software with resilience in mind so that failure is not a problem. And so, this is where the idea of chaos engineering comes from. And so, a tool like Cthulhu is something that will run in our system that introduced outages so that we can verify that our system detects those issues and in some case—in most cases, ideally—recover, programmatically or thematically from those failures so that we don’t need to page one of our engineers at 3 a.m.

    Ashley: Tobias, is this a project or an idea that you came up with or you and Gabrielle did it together, Gabrielle came to you, or how did this all get started at xMatters?

    Dunn-Krahn: No, I will claim no credit for this project. [Laughter] I will hand that completely to Gabrielle. But another way of looking at that is, as we decomposed our monolith over time into microservices, what we really wanted to do was reduce the scope of subject matter expertise and knowledge amongst our teams to focus that on a smaller part of our system. So, that reduces complexity for individual teams and being able to keep those services up and running and highly reliable.

    But one of the things that it also introduces is, it chases that complexity into the space between the services—different failure modes of those services, different complexities around the dependencies.

    So, during our digital transformation we recognized that early on, and Gabrielle came up with the idea of introducing chaos testing for this and also that the existing chaos testing tools were not suited for our purposes, and I’ll let her explain how Cthulhu was different.

    Ashley: Okay, great. Gabrielle, why don’t you pick it up from there?

    Gasse: Yes. So, a year ago, when we started working on chaos engineering, there was a few tools that was available. Chaos Monkey is a very popular one. At the time, Chaos Monkey was working for Amazon Web Service primarily, and also it’s in deployment in Spinnaker. But it didn’t work for us, and so, we’re on Google Cloud.

    We came up with this tool that can connect to lots of different cloud services—currently, we support Google Cloud and Kubernetes, but we’re adding other services like Amazon Web Service, for example. And so, it can run those failure scenarios in concert together, and also, we wanted to have the ability to version control those tests so that when we find a particular scenario that does cause a failure, we could have a way to produce it as a file, if you like, that we could then attach to a bug that we could give to the engineer to work on.

    Ashley: Let me ask, then, you mentioned Chaos Monkey, which to me, was sort of a natural comparison for what you’re doing—does Cthulhu run in the background, kinda running all the time, creating its havoc, much like Chaos Monkey does of getting things to go down and see if they recover in a production environment? Is it similar for Cthulhu or are you taking a little bit of a different approach?

    Gasse: It is similar, although there is a different way to run it. In usual scenarios at xMatters, we are running it on a need basis. So, we have a small scenario that we’ll run immediately. It’s also possible to run Cthulhu, as you mentioned, in the background and it will pick up tests and run them, essentially.

    Ashley: Mm-hmm. So, run it in a targeted control manner, or let it run in the background and do its thing, it can go either way?

    Gasse: Yeah, exactly. The thought was very much to have this ability to run it constantly like Chaos Monkey, but also give a tool for an engineer to perform a scenario locally or in a development environment to try and reproduce and understand failures.

    Ashley: And you mentioned, of course, Chaos Monkey was created by Amazon Web Services, functions very well within an AWS environment—was multiple cloud services the main differentiator for Cthulhu, or was also getting into Kubernetes, Docker, Containers something that you kinda took on as a special added functionality that Cthulhu does that maybe another tool doesn’t? What are all of the differentiators that you’ve created Cthulhu for?

    Gasse: Chaos Monkey was actually created by Netflix, who have their stack on Amazon Web Services.

    Ashley: Oh, you’re right. Yeah, my mistake—thanks for correcting me.

    Gasse: But, as you mentioned, Chaos Monkey works on a virtual machine in Amazon Web Service exclusively. There are other tools like Pumba that works at the Docker Container level. Powerful Seal came out last year around the time that we were looking at that, and Powerful Seal works on Kubernetes deployment at the pod, but also at the node level, which is pretty interesting.

    But there’s no one system that will do—that will test across all of those platforms. You need to use Powerful Seal in concert with Pumba, and there was nothing that was running directly on Google Cloud at the time. Now, Chaos Monkey does support Google Cloud.

    So, the main value that Cthulhu brought for us was the ability to have a single tool that could perform complex failure scenarios across both our Kubernetes deployments and our VMs. And—yeah, and we wanted, also, the ability, as I mentioned earlier, to version control some of those tests so that we could pass them on to our developer where Chaos Monkey and those other tools, you run them with some parameters to filter out certain machines that you don’t want it to break. And then it just rolls from there. There’s nothing—there’s no way to build a scenario from it that is repeatable.

    Part of what we test with it is that our system is able to detect failures. So, as we run a scenario, we’re expecting errors to be logged in Splunk alerts to be sent out to us. And so, the tool itself simply introduced failures, and then it’s up to the engineer to look for what is broken.

    Ashley: Mm-hmm, mm-hmm. Very good. Tobias, how does Cthulhu fit into the strategy of xMatters?

    Dunn-Krahn: So, we would—you know, our personal experience or our corporate experience of going through our own digital transformation was—and a continuing process, I don’t wanna make it sound like we’re done, we’re never done. But it was very instructive, and we learned a lot, and we’d like to share that. And many of our customers are going through the same thing, so we would like to provide them the tools that they need to be successful in these transformations.

    So, Cthulhu in particular aids in the transition process more than what you would—the situation you’d be in if you were a Cloud Native company. So, we wanted to provide that support to our customers and to the community at large. We have long been  users of open source software, and we wanted to contribute back to the community as well as get all the advantages of having an extended developer community contributing to Cthulhu and we can all make it better together.

    Ashley: Great. So, Cthulhu is an open source project. Do you have it hosted at GitHub or your own servers? Where do people find it?

    Dunn-Krahn: Gabrielle, you can confirm, but I believe it’s on GitHub.

    Gasse: That’s correct, yes.

    Ashley: Great. You know, sometimes open source projects are open, but largely, most of the work is done by a lead developer or someone who created the idea. Sometimes, it’s a very vibrant community where people are testing or creating new features. You know, and part of that is also not only the style of the project but also the maturity of it.

    Where are you in that? Is it mostly you, Gabrielle, who’s doing the work? Do you have—what kinda activities do you have from the outside from other people?

    Gasse: Yeah, the publication of Cthulhu on GitHub is fairly recent. We’ve had a fair amount of people local to Victoria when we promoted it so far that are following the project, but we have yet to see people who are really picking up or starting the contributing modules for it.

    Ashley: Mm-hmm.

    Gasse: I think one of the next additions that will really help motivate people to use Cthulhu but also contribute additional modules will be us supporting Amazon Web Services in particular.

    Ashley: Mm-hmm. And how long has it been out? I should’ve asked that originally.

    Gasse: It’s been out for a few months, only.

    Ashley: Okay. So, still pretty early in its life. So, it’ll be interesting to see, as more people use it, that’s usually where someone will get interested and say, “Well, I wanna put some stuff in here. I wanna try to make some changes and contribute some things.” What’s Cthulhu written in?

    Gasse: Cthulhu is written in Java. We’re using the Spring Boot platform, and as part of that, there is an easy way to write modules that plugs into the platform that adds functionalities. So, it’s easy to add new ways to break things in Cthulhu without having to understand the entire code base.

    Ashley: What are other things that we should know about Cthulhu?

    Gasse: I can talk briefly about the next item on the roadmap, if—

    Ashley: That’d be great, yeah.

    Gasse: —yeah.

    Ashley: Sure—love to hear that.

    Gasse: Mm-hmm. So, after the support for AWS is done, my next task, if you like, is to start adding a functionality to run commands through SSH and sending files through SCP on target hosts. And that opens the door to a whole new range of tests. For example, one could pause a process on the target virtual machine to simulate a dead live or use tools like Traffic Control to introduce noise on the network, lost packets and all that.

    And so, once we have that support, we’ll really be able to have lots of interesting types of failure that go beyond simply just, a virtual machine has shut down and is no longer available.

    Ashley: Tobias, I appreciate very much you handing all the credit where credit is due to Gabrielle. Also appreciate your thoughts on where do you see this going from your perspective, either strategically or aligned with product plans, product strategy—what do you see as the future for Cthulhu?

    Dunn-Krahn: Sure. Well, as I mentioned earlier, the xMatters product is built to help teams support digital services, and part of that is making sure not only that you plan for an incident but that you can simulate real world incidents and do some form of practice or maneuvers.

    So, the way I see this being integrated into the product in the future is facilitating those types of practices and then evaluating the performance of the team or that all of the configuration in xMatters is correct to solve problems as quickly as possible, et cetera.

    Ashley: Excellent. Well, you know, we’ve probably just barely scratched the surface, and I’m excited to see what you all do and what the community does with Cthulhu. I’d love to thank you both for being on the podcast. I thank Tobias Dunn-Krahn and also Gabrielle Gasse from xMatters, any additional information might be available at xMatters, is that true, Tobias?

    Dunn-Krahn:  That’s a good question. Yes, there’s certainly some material on the website, or you can search for it on the web if you can figure out how to spell it.

    Ashley: [Laughter] Well, let me give everybody a head start, it’s C-T-H-U-L-H-U, so, just the way H. P. Lovecraft spelled it. Well, I’d like to thank both of you, Tobias and Gabrielle, for joining us today on the podcast. Time has flown by again, and hopefully, we can have you back another time as this evolves and you can tell us some more stories about where you’ve taken it and what people have done with Cthulhu.

    Dunn-Krahn: Alright. Thanks very much for having us.

    Ashley: You bet. Thank you. This is Mitch Ashley with staging-devopsy.kinsta.cloud and you’ve listened to another DevOps Chat.

    — Mitchell Ashley

  • DevOps Chat: Global Intelligence for AWS GuardDuty, With Sumo Logic

    DevOps Chat: Global Intelligence for AWS GuardDuty, With Sumo Logic

    Wouldn’t it be helpful to know if other cloud users are seeing the same or similar attacks that you are? Security intelligence about cloud applications beyond just those you own and operate as an enterprise opens up a new dimension in attack visibility against even large sets of cloud apps.

    During AWS re:Inforce 2019, Sumo Logic announced it is extending its machine analytics and intelligence platform to include AWS GuardDuty. Dubbed Global Intelligence Service for Amazon GuardDuty, the new service is more than just a data aggregation and reporting play. The new service provides additional context around GuardDuty data by reporting attack information across multiple Sumo Logic customers using AWS GuardDuty. Essentially, it’s a “crowdsourcing” approach to reporting threat intelligence across the cloud.

    In this episode of DevOps Chat, David Andrejewski, senior engineering manager at Sumo Logic, joins us to talk about this new, more expansive threat intelligence service. More information about Global Intelligence Service for Amazon GuardDuty is available in the press release and website.

    As usual, the streaming audio is immediately below, followed by the transcript of our conversation.

    Transcript

    Mitch Ashley: Hi, everybody, this is Mitch Ashley with staging-devopsy.kinsta.cloud, and you’re listening to another DevOps Chat podcast. Today, I’m joined by David Andrzejewski, Senior Engineering Manager at Sumo Logic, and our topic is a real meaty one—Global Intelligence Service for Amazon Guard Duty. David, welcome to DevOps Chat.

    David Andrzejewski: Thanks so much. Thanks for having me on.

    Ashley: Well, we really appreciate you being on the podcast today. Would you start by introducing yourself, just tell us a little bit about what yourself and what you do at Sumo Logic, and maybe just a brief overview of Sumo Logic the company?

    Andrzejewski: Yeah, fantastic. So, Sumo Logic, what we offer is a cloud-based machine data analytics platform, which is a bit of a mouthful, but what that specifically means is a cloud-based, multi-tenant service where you can send your logs and metrics and data to our service for helping you monitor, troubleshoot, and secure your applications, platform, environment, infrastructure for your own application or your IT infrastructure.

    So, we offer this as a kinda cloud-based web app service. The idea here is, you’re sending us all of that data and we can kind of manage the scale, reliability, and analysis for you and kinda help you focus on running your business and kind of running your app and everything like that.

    Specifically at Sumo Logic, I manage the Advanced Analytics Team, and we’re focused in particular on ways that we can do especially kind of useful or interesting things to help customers get more out of their data. So, some of the features in the past we’ve worked on are log reduce, which kind of applies clustering techniques to your logs to kind of give you a more comprehensible snapshot as well as outlier detection and things like this. But our current focus right now is this kind of Global Intelligence Service that we’re really excited to be announcing at AWS Reinforce.

    Ashley: Great. Well, we are a DevOps and also a security podcast—talk a little bit about how Sumo Logic fits into the DevOpsSec world.

    Andrzejewski: Yep, absolutely. So, ultimately, in any of these situations, you really need some kind of—you need the ground truth actual data to really kind of know what’s happening in your environment, both from what kinds of experience are your customers getting, how is the, kind of the health of your machines or what resources are available, what is the performance of various services, you know, events are being kind of logged and audited and then you’re kind of pulling data together from all up and down the stack into one place.

    And that’s kind of where we come in as sort of one kind of one-stop shop for sending all of these very kind of heterogeneous logs into one service such that the teams that are responsible for up time from kind of a DevOps SRE perspective and security from all sorts of aspects of that, which increasingly is, like you said, kind of this DevSecOps. The developers themselves are often increasingly deeply involved in all of that, where they can kind of pull together the data from their own kind of custom application logs, the infrastructure, the actual machines, the operating system, the firewall, the databases, the third party services and kind of have it all in one place so that when you can sort of help detect when things do go wrong and then kind of identify, dig in, troubleshoot what’s going wrong and then once you fix it, you know, verify that it actually is corrected.

    Ashley: Mm-hmm.

    Andrzejewski: So, you would interact with the service by sort of operating queries against your data, creating dashboards and visualizations and using these dashboards or potentially setting up alerts to notify members of your team that some condition has been triggered that requires immediate attention or action to generate sort of scheduled kind of reports on these kinda things and in general, kinda analyze and interact with the data, the raw data being kind of sent out by your software, your machines and everything like that to really understand what’s going on in your app.

    Ashley: And are you oftentimes the system in the SOC, or do you connect into a different monitoring system, alerting system in the SOC?

    Andrzejewski: So, this is kind of a—as far as the SOC aspect, some customers are kind of using it for tracking these different kind of logging events, others are sort of, we have a lot of kind of integrations with other services to kind of consume data kind of emitted by those services. So, for example, AWS Guard Duty, you know, alerts like that.

    So, it all kinda depends on, there’s sort of a wide variety of ways that customers have kinda made use of Sumo Logic. So, yeah, in some cases, I think that that’s a reasonable use case that the [Cross talk]—

    Ashley: It’s probably, I imagine, where they’re starting from if they have systems all ready for you to integrate in, or maybe you become the main single pane of glass, if you will.

    Andrzejewski: Yeah, so, in the security use case, I think—again, it depends on kind of which sort of person within the organization might be most interested in that. Developers often, you know, given that this is a place, a single sort of solution that you can collect the logs for both these kind of monitoring SRE troubleshooting as well as some of the security use cases.

    Ashley: Mm-hmm.

    Andrzejewski: That ends up being kind of an appealing property that you can kind of have one service to collect the data, analyze the data for various use cases, right?

    Ashley: Okay, great. Well, this is the week of AWS Reinforce, and I feel like we’ve kind of laid out enough teasers about your announcement that you’re making this week. [Laughter] Tell us about Global Intelligence Service for Amazon Guard Duty.

    Andrzejewski: Yeah, yeah, we’re really excited about this. So, the basic idea here is that we’re going to add some additional context around the Guard Duty alerts where we can essentially take your alerts and kind of recontextualize them in terms of how frequently we’re kind of seeing them or they’re being sent or observed across others.

    So, if something is a very kind of rare threat that very few other organizations are kind of getting, seeing in their environment, then that might be something that you would want to prioritize more highly, right, if you are—you know, there’s going to be some, in any kind of alerting system like Guard Duty or others, there’s often going to be a very kind of high just baseline of sort of very generic, general threats that everybody kind of sees at some base rate all the time, and this is kind of one way that you can sort of kind of do some noise reduction on those and really kind of prioritize and focus on the ones that are a little bit more idiosyncratic or unique to your environment and more, perhaps more interesting and more relevant for kind of immediate investigation and triage.

    And this is kind of being exposed to the customers via an app, which is kind of, in the Sumo Logic scenario, it’s kind of some pre-defined analyses and dashboards against your Guard Duty data. So, your customer would sort of consume this service by installing the new Guard Duty with Global Intelligence Services app, and they would immediately get access to these new visualizations, analyses, and dashboards that kind of place their own Guard Duty alerts in this modified context, this more—you know, bringing in a little bit broader information about how often these are being seen in other environments, what are the general most common threats, threat purposes, various threat types seen across sort of the general customer population at Sumo Logic.

    And this will sort of kind of serve these multiple purposes for the user—on one hand, helping you kind of prioritize rare events, on the other hand, kind of giving you sort of this bird’s eye view survey of what’s generally out there, right? And that can help potentially in prioritizing or allocating efforts at improving your own security posture or efforts in, for example, you see that a particular kind of threat is especially prevalent, you know, elsewhere, that might be something that you say, “Okay, well, what do we have to be doing to make sure that we’re, our bases are covered with respect to that ourselves,” right?

    So, ultimately, we want to kind of show you your AWS Guard Duty threat data, but in this kind of enhanced context that’s hopefully going to be more meaningful and more relevant to you and help you make better decisions and kind of run your business more securely.

    Ashley: Okay, so, the benchmarking aspect of this, is that something that happens in real time? You know, we’re looking at the console, we’re getting alerted to say, “This attack is happening, and by the way, it’s happening on 20 percent of the other AWS customers” or however your present it, or is it more a retrospective looking at past data of what’s occurred over the last month, last quarter, last year? Or is it both?

    Andrzejewski: Yeah, so, this is something that we’re kind of, in some sense, as we are having conversations with customers, seeing what’s going to be most valuable for their use cases. Right now, the real time—is it real time, like, up to the second? You know, is that really something that’s going to be valuable, or is that going to be kind of too noisy of a signal?

    Ashley: Mm-hmm.

    Andrzejewski: So, right now, it’s derived from sort of recent history in some sense, and yet, potentially, that’s an interesting question of whether more very long term stuff, like, month over month is going to be useful for people, but you know, is it—right now, we’re not kind of focused on the use case of, you know, up to kind of real time real time.

    Ashley: Mm-hmm.

    Andrzejewski: Really, your data will still be real time, real time, but the question of whether that’s—you know, does the context need to evolve that quickly, or is that just going to kind of introduce more noise?

    Ashley: Right.

    Andrzejewski: Part of the advantage of kind of using a little bit of a recent historical window is looking at—you know, it gives you a little bit broader pool of data to kind of contextualize your real time activity against, right? So, again, it’s all about sort of your data and kind of seeing how—you know, how it kind of looks in context of this broader view and evolving the broader view in real time. It’s an interesting thought.

    Ashley: Well, that makes sense, too, because you may be having a false positive or something. Not everything that is being highlighted by Guard Duty might be as relevant as other things, so if you have a chance to filter some of that information out to say, “This is what’s relevant” versus, if everything’s real time, you could introduce some noise, maybe some things that aren’t as helpful, too.

    Andrzejewski: Yep. Yeah, exactly. So, it’s—that’s definitely a potential tradeoff, there.

    Ashley: So, what would you learn—because now, there must be some kind of an opt-in where you’re looking at events that are happening across different customers, AWS users—what would you learn by knowing that a whole other population of people are seeing this same kind of attack versus if you were only looking at your data, if you had known that you would do something differently?

    What would be a scenario like that where this would be super helpful?

    Andrzejewski: Right, so, the—you know, one kind of particular case sort of going the other way of a rare event might be that, you know, again, so, some of these threats are going to be very kind of common, and it ends up that that’s not something quite as potentially actionable or something that you need to jump on right away, but there’s still kind of part of your Guard Duty feed and they’re going to kind of potentially be a little overwhelming when you’re looking at everything. Whereas this more kind of like rare threat type, this—by kind of knowing, “Okay, this is a little bit out of the ordinary in that lots of customers are not seeing this,” that kinda helps it pop to the top of your queue and maybe you are going to allocate your very limited time and attention to focus on that.

    It might be a little more challenging to prioritize or see that kind of rare event pop, because you don’t necessarily know that it’s rare, you just sort of see it amongst all of the other kind of more common alerts and more common threats.

    Ashley: Right. This is a rare event, something that’s unusual, you should take a look at this versus something that’s happening normally across multiple AWS accounts.

    Andrzejewski: Exactly. That’s exactly right. And so, that potentially becomes a valuable tool to kind of—like I said, ultimately, teams everywhere are ultimately struggling with the scarce resource of your time and attention.

    Ashley: Mm-hmm.

    Andrzejewski: And anything that can kind of bring in some additional context information to help you allocate that time and attention more efficiently, more effectively is potentially, hopefully going to give some good value for our customers.

    Ashley: Well, many resources are precious resources, and certainly security folks are definitely in that precious category. So, being able to assign them to something you know is something we need to look at versus kinda chasing things that we really don’t need to be paying attention to would be super valuable.

    Andrzejewski: Absolutely.

    Ashley: So, obviously, Sumo Logic does a lot of other things in your products. You have operational analytics, business analytics, those kinds of things. Does this Global Intelligence Service tie into those in some way, to give you some additional broader insights, or this focused really—just really heavily on the security aspect of what’s happening in Guard Duty?

    Andrzejewski: No, absolutely—that’s a really good point. So, right now, kind of, very sort of aligned with Global Intelligence is some of the stuff that we’ve been doing a few years now around the modern app report and similar where kind of giving these sort of horizontal views of technology choice trends and various kind of software choices, infrastructure, cloud provider choices, kind of adoption of various tools and technologies across the industry, and that definitely kind of extends into the, kind of the DevOps SRE space. And again, hopefully the idea there is that these tools kinda give you this sort of broader context about actual hard data where various—you know, where the market in general is on their journey to cloud adoption or their migration to Docker or Kubernetes or other kind of related technologies.

    Ashley: Mm-hmm.

    Andrzejewski: So, that’s one place where there’s already something in place at Sumo Logic that’s kind of bringing that broader sort of horizontal context to bear to kind of help teams kinda make more effective, more informed decisions and, again, prioritize and allocate their attention and efforts in kind of using that extra information to kind of know, actually, from the hard data where others are actually at.

    Ashley: Mm-hmm.

    Andrzejewski: So, right now, this particular announcement at AWS Reinforce is focused on Global Intelligence Services for AWS Guard Duty in particular which, as you said, is absolutely a very kind of security focused use case.

    Ashley: Mm-hmm.

    Andrzejewski: But in general, the general sort of Global Intelligence Services, you know, kind of platform and approach is something that we’re also actively thinking about how can we help in the operational analytics use cases? Where are there places where that broader context is going to help our customers make better decisions or solve their problems more quickly?

    So, right now, we have this Guard Duty announcement, definitely focused on security, but in general, we see Global Intelligence Services as a very powerful kind of horizontal enabler across the SRE DevOps kind of side of the house as well.

    Ashley: Mm-hmm. Yeah, it seems like an area that would be ripe for machine learning—yeah, the statistical analysis that you could do across multiple areas of the analytics that you’re doing. You know, that’s what data’s all about, right? What insights and learnings can you get from that that brings additional value benefit to the business?

    Andrzejewski: Definitely, yep.

    Ashley: Well, good. Is this product available now? Is it coming out some time down the road? How soon can we get our hands on this?

    Andrzejewski: So, as of the kind of announcement, it should be available to Sumo Logic customers at the enterprise level, and again, it would be a matter of installing the app and kind of start—obviously, in this case, you would need Guard Duty as well. I can provide some links to go out with this, if that’s helpful—

    Ashley: Sure.

    Andrzejewski: – that would have kind of links to the specific resources that would be involved in getting started and getting set up with this, but yeah.

    Ashley: Okay, great. It sounds like a pretty easy thing to do to get started. If you’re already running Sumo Logic, installing the app is not a heavy lift, it’s pretty straightforward.

    Andrzejewski: Absolutely. And the other aspect, of course, having Guard Duty in place as well. Which, again, is—in the links I’ll send, should be, is a pretty painless process.

    Ashley: Excellent! I also want to mention, too, that Sumo Logic is located in Booth 714 at the Reinforce conference. So, definitely recommend folks stop by and maybe see a demo of this, get to talk with folks about the new announcement.

    Well, our podcasts have no problem flying by quickly, because we talked about great things and I appreciate you doing that with us. David, I’d like to thank you, David Andrzejewski, Senior Engineering Manager at Sumo Logic, for joining us.

    Andrzejewski: Well, thank you so much for having me. It’s been a really great conversation—super, super interesting to talk about.

    Ashley: Great. It’s good stuff, and I wish you and Sumo Logic the best at the show this week.

    I’d like to also thank you—you, our listeners—for joining us. This is Mitch Ashley with staging-devopsy.kinsta.cloud, and you’ve listened to another DevOps Chat. Be careful out there.

    — Mitchell Ashley

  • DevOps Chat: Enterprise DevOps Adoption, with Micro Focus’ Mark Levy

    DevOps Chat: Enterprise DevOps Adoption, with Micro Focus’ Mark Levy

    Adopting DevOps means changes from how you build software to organizational processes and cultural behaviors. Large enterprises have even greater challenges to overcome, including supporting a large number of applications and addressing technical debt while also taking on new development initiatives like digital transformation.

    What can we learn from experienced DevOps practitioners?

    In this DevOps Chat we spoke with Mark Levy, director of strategy, software delivery at Micro Focus, covering seven keys to successful DevOps implementations at large enterprises.

    As usual, the streaming audio is immediately below, followed by the transcript of our conversation.

    Transcript

    Mitch Ashley: Hi, everyone, this is Mitch Ashley with staging-devopsy.kinsta.cloud, and you’re listening to another DevOps Chat podcast. Today, we’re tackling the topic of large enterprise adoption of DevOps—no small task, for sure.

    I’m very happy to welcome Mark Levy from Micro Focus. Mark, welcome.

    Mark Levy: Thank you, Mitch, and glad to be here.

    Ashley: Would you just tell us a little bit about yourself and what you do at Micro Focus?

    Levy: My name’s Mark Levy. I am a director and working in the product marketing organization, and I focus a lot on strategic marketing initiatives as an evangelist, a DevOps evangelist. I develop and write content and commentary about the DevOps market. I’ve had a 25-plus-year career, and like many people, I’ve had many positions. I started out as like a sysadmin and was a developer and managed dev teams.

    But, in 2007, I started to focus exclusively on application delivery. And so, that has been—this whole area of application delivery, software delivery is something that I’ve been really focused on for the last 10 or so years.

    Ashley: Well, tell us briefly, a little bit about Micro Focus. Not all of our listeners might know what you all do.

    Levy: So, Micro Focus has also gone through this major transformation over, I would wanna say, the five to eight years or so. You know, we’re a global enterprise software company. We were established in 1976, we have 14,000 employees and 40,000 enterprise customers, about 4,800 software engineers, and we really focus on large enterprises, helping them run and transform their business. We have sort of focused on four key areas of how do we support our customers through their digital transformation journey, and enterprise DevOps is one of them.

    Ashley: I imagine that our folks, listeners that are from large enterprises are gonna be particularly interested, and I’m sure there are learnings and information that will be useful to other folks as well.

    So, you talked about playing in the space with large enterprises. I imagine that’s probably some global international companies, maybe some that are both European or U.S.-based as well. How would you describe how you differentiate what you do as Micro Focus as compared to other folks? There’s a lot of companies, of course, working to try to help enterprises adopt DevOps.

    Levy: Yeah, of course. And that’s sort of what I’d like to do first maybe is just sort of set the stage, because I always look at this, especially in the enterprise IT world, it’s so interesting. You know, Gartner made this sort of—you know, their strategic assumptions. And they said something like, through 2021, at least 80% of the DevOps initiatives will fail.

    Ashley: I have not seen that, no. Interesting.

    Levy: I don’t know if you’ve heard that, but—yeah. So, it’s interesting. What does that tell me—tell us, I believe? One is, the business is already in the fight. The enterprise, the business side is already in this digital economy and fight. We see that on the streets.

    And the other thing is, you know, IT is sort of behind the curve. And the reality is that our customers, the IT customers, the IT organizations have this incredibly complex hybrid landscape, it spans mainframe, mobile, on-prem, off-prem, and while the principles—the main DevOps principles—are similar to what the unicorns, so to speak, are delivering, that it’s really a whole different challenge.

    Because what’s critical is that there’s this above the line mandate to drive new digital revenue, but at the same time, you know, optimize the enterprise. And, you know, I sort of—I like one of my colleagues, his reference. It’s like, redesigning a jet plane while you’re flying it, you know? And so, I think that’s how, you know, what’s critical with large enterprises is that they’re trying to transform themselves and at the same time run the business.

    Ashley: Yeah, I would almost describe for a large enterprise, it’s like trying to build or transform the entire fleet of planes while you’re building it.

    Levy: Yes! Exactly.

    Ashley: Talk a little bit about—I mean, we all, I think, can appreciate in every business, especially large enterprise, they’re saddled with technical debt; they, of course, are working on operational efficiencies; they have applications that they’re trying to support and maintain.

    How do they take on something like a DevOps initiative successfully, as you were pointing out earlier, so that it doesn’t just stop at one group and kind of fall off from there, they really wanna be able to transform their part of their business? How do you do that?

    Levy: Right. So, I think one of the key things that we focus on enabling our customers to build on what already works, okay? They have all these core investments, and how can they take their core investments and leverage the ROI that they’re already having today, but implement new technologies and DevOps practices and capabilities? And I think that’s critical. I mean, how do they implement and scale DevOps practices, for example? It’s very, very important, and it’s hard.

    And so, we look at it from the basis of, how can our solutions and our products help them really get from point A to B? What is their current state now? Which is what they really need to understand. And it’s amazing how many companies really don’t have a current understanding, detailed, of their current state.

    Ashley: Mm-hmm, or a common one. [Laughter]

    Levy: Or a common one, because, you know, it changes a lot, but how do they deliver value to the customer? You know, what are their value streams, and what are the processes and the flow of artifacts? And so, they need to understand their current state.

    And then, they need to do a number of things that I know the people that are helping the DevOps community, the DevOps community have been evangelizing for a while—I think critical, especially in large organizations and enterprises, you know, they need to first really create this culture of continuous improvement. I think that’s critical.

    Ashley: Okay.

    Levy: You know, they need to enable, motivate and empower the teams, and the leaders really need to get on board, right? Without the leaders, it just will not work.

    Ashley: Yeah, I can speak from personal experience—instituting a continuous improvement culture is no small task in and of itself.

    Levy: Right, and you know, the thing is, the leadership style must also change, right? I mean, command and control really doesn’t work, decisions—it’s about empowering your product teams to drive as much decision making as possible. And that sort of has gone contrary to, you know, the organizational structure of the large enterprises.

    Ashley: Mm-hmm.

    Levy: So, that, I think, you know, is one of the critical challenges that these companies are faced with. Edward Deming made a great quote, “A bad system will beat a good person every time,” right?

    Ashley: Mm-hmm.

    Levy: And when I think of a bad system, I think, in its largest sense, the system. And this is where I think Micro Focus can help quite a bit is, anyone wanting to change this culture really needs to define the actions and behaviors that they desire and then design sort of the work processes that are necessary to reinforce those behaviors.

    So, you know, systems, processes, tools, help with and enforce good culture. But you know, if you have a bad system, it’s not gonna work. So, I think, you know, these are areas that we can help on the cultural side. I come back to thinking about the idea that, you know, sort of the general friction between dev and ops teams, where ops teams are really, initially, only measured by mitigating risk, and dev teams are only measured by delivering change. You know, there’s the conflict that has to be resolved. That’s critical.

    Ashley: You’re talking about culture; the other thing that occurs to me is, there can also be inherent or built in incentives that are counter to the culture that you’re trying to create. That friction is a good example for one. But, you know, there may be past punishment for failures or, you know, emphasis on down time, so the answer is always minimize change or whatever it might be. You’ve gotta take those things head on.

    Levy: Exactly, yeah. And so, I think, you know, you ask sort of where people start. I think, you know, you need to start—and everyone, a lot of people talk about this, and it is true. You really need to start instituting this culture of change, because DevOps drives a lot of change. It’s a transformational journey in the enterprise.

    Ashley: When you said, the number one thing you said was, “Leverage what works.”

    Levy: Yes.

    Ashley: What do you look for? What kind of things are you trying to say, “That’s a practice you do well, and let’s keep that, or let’s build upon that”—what do you look for?

    Levy: Well, you know, I think about, there’s a lot of legacy systems, core investments that these large enterprises have had and invested lots of money and time into it. And take, for example, COBOL, okay? That language, I think people are not necessarily aware how invasive it is. There’s still about 250 billion lines of COBOL around, right?

    Ashley: Mm-hmm, mm-hmm.

    Levy: It drives 70%, they say, of all business transactions—still, in the world.

    Ashley: That’s amazing.

    Levy: It is amazing.

    Ashley: [Laughter]

    Levy: And so, you know, I’ll give you an example of sort of what we’ve done and we’re, you know, we are very good at managing and understanding those environments. We have, you know, that’s a big part of our business. So, we can, you know, lift—if we can lift and shift, you know, COBOL workloads off of, you know, sort of the more legacy platforms into, let’s say, AWS, you know, that’s a big win for our customers.

    As an example, we had one customer that basically had—it was a major global insurer, had about eight different IT environments with a back end COBOL app sitting on the mainframe and then front end and WebOps. And basically, each environment cost about $3 million a year. And, to request that environment, to provision that environment, took six weeks.

    Ashley: Mm-hmm, mm-hmm.

    Levy: And so, we built, through our solutions, we were able to lift and shift that workload to AWS and reduce the cost from 3 million to 120K and reduce the provisioning time from six weeks to two hours.

    Ashley: And I wanna mention, too, to our listeners, too, we just had a webinar in the last week talking about mainframe APIs and frameworks specifically to mainframe environments to make it a little bit easier to adopt WebOps. So, though it doesn’t get the kind of sex appeal that necessarily new technology does, it is a huge, huge issue for enterprises and how do you bring the mainframe environment with the rest of the cloud and everything else.

    Levy: Right, it’s—you know, back to the theory of constraints. I mean, in your value stream, if—you know, you can develop and you can be an agile team, but if it takes six weeks to provision, you know, your mainframe environment?

    Ashley: Where do you suggest people get started? Is it go pick a new, from scratch greenfield project, or pick something that’s, you know, reasonably mission critical so we’ll get the support?

    Levy: Yeah. Well, I think one is, I like, initially, to set up a lab, like taking a greenfield. You need a place to learn. A lot of these things aren’t easy, or they’re not, people are not, the teams are not used to. So, like, I love the idea of what a lot of the community’s doing with DevOps dojos and building sort of a team level, maybe CI/CD, have a place to experiment and take that.

    But then, you know, ultimately, in large enterprises, they’re gonna come up. They might get easy wins early, right, but the challenge is, once they scale, they’re gonna run into some of the more challenging times. And that’s where, back to what they have to really start doing is taking this broader, system level view of the enterprise.

    So, you know, if they understand how they’re delivering value to their customers—and in enterprise, there’s many ways to do that—then they need to understand and prioritize according to the business objectives. Let’s say one business objective might be, we need to recapture value. Maybe our IT budget isn’t really growing, but we need to drive for more innovation. So, we need to—our objective is to recapture that value and redeploy it into delivering value and innovation. And so, that would drive a lot of your decisions in where you optimize.

    Ashley: Speaking of budgets being flat or not growing, digital transformation is all the rage, that’s what many, many companies are taking on, not just in customer experience, but across the enterprise, bringing them into the digital world. How does that fit on top of this if you’re trying to introduce DevOps?

    Levy: I don’t see how they can really run their digital transformation without DevOps. I mean, software today—you know, I talk about this a lot where business innovation and software innovation are generally becoming one and the same, right?

    Ashley: Mm-hmm.

    Levy: And going quick, innovating faster with less risk is critical. And so, IT is, you know, it’s imperative to really deploy DevOps practices to support the digital transformation of the business. So, I think, you know, the IT organizations that implement DevOps, those people are the drivers.

    Ashley: Well, and it’s hard to imagine how you were gonna transform systems and applications you already have as well as ones you might leverage as SaaS services or create yourself without an environment to integrate those altogether due to continuous integration testing, etcetera, as well as be responsive to the business.

    Levy: We’ve seen enterprises, they will typically start out, right, with their team level implementation and they might do really well, and it might be a greenfield app. Maybe it’s a new growth app, which is great, and they may get some early wins, which is super, but it’s still not their core business.

    Ashley: Mm-hmm.

    Levy: And so, they have to, ultimately, really transform their core business in order to be competitive in the market.

    Ashley: Mark, what’s the one thing organizations should not do when they’re looking to adopt DevOps—don’t make this mistake?

    Levy: I think the idea that they can just automate their way out of it, I think, is something that could get them in trouble. Because, you know, really leaving out the whole cultural aspect of it, it really is—you know, it’s interesting, if you read Jeff Bezos’ letter to shareholders every year, he just provides some incredibly insightful, you know, insights on how Amazon runs. He talks about this accelerated decision making process that they run, and they make these decisions, really, based on only 70% of the information that they need, basically. Whereas, large enterprises wait until they typically have 90%.

    And so, they are able to make decisions earlier because of the agility that they’ve built into their system and their culture of change. And this allows them to not only, you know, get earlier into the market, but it allows them to—everyone, it allows them that if they make a mistake, they can recover from their mistakes faster, and the cost of making a mistake actually goes down.

    Ashley: Mm-hmm.

    Levy: But this whole idea of organizational agility and providing that—there’s more than just delivering value faster to the customer, because it’s never a straight line, right, to success. So, we make a lot of mistakes and wrong turns in our journey, and being able to identify that quickly and have the agility to correct that quickly is actually more of a differentiator than just delivering fast.

    Ashley: It’s certainly required for any kind of agility. It’s not just getting things and getting decisions made faster, but changing when things aren’t headed in the right direction—or change.

    Levy: Right, right. And so, this is where, you know, if people are looking at DevOps, they’re saying, “Okay, well, we’re just gonna automate the deployment pipeline and we’re done?” No, there’s no “done,” it’s about really, ultimately, improving as an organization, right, and having that continuous improvement, the Kaizen kind of mentality of the organization. And that, to me, is what is so transformative as a company.

    Ashley: Going back to what you said in the beginning about having a culture of continuous improvement.

    Levy: Yeah, absolutely. And that, it would be, you know, frankly, the most powerful thing leaders can do—is that change. Because it really just changes the whole business. The business, then, gets more confidence, which is critical, right? With all this change, sometimes they’re not confident to make business decisions, but if you can provide them the confidence to go out there and really be aggressive, you know, that’s—you’re gonna see, you know, a whole bunch of positive results on the revenue side, on the public side of the business that’s been driven from, traditionally, the back office IT organization.

    Ashley: Well, Mark, I’d like to thank you. I’d like go for another 30 minutes, [Laughter] but I’d like to thank you for the time on the podcast. You know, I walk away with several things from our conversation. The things you shared that I heard were leverage what works, define your current state, create a culture of continuous improvement, command and control doesn’t work—you’ve got to empower product teams. Start with a lab or a dojo, some learning environment. I think your premise of, you can’t do digital transformation without DevOps is pretty intriguing. And last but not least is organizational agility, being able to make decisions quicker, get out ahead of things without having to have 100% of the information, kinda the Amazon model.

    Well, and that should wrap things up. I’d like to thank Mark Levy from Micro Focus for joining us. This is Mitch Ashley with staging-devopsy.kinsta.cloud, and you’ve listened to another DevOps Chat.

    — Mitchell Ashley

  • The Failings of Multitasking

    The Failings of Multitasking

    I often find myself coaching colleagues (and reminding myself) about managing down the number of simultaneous tasks or work assignments. Everyone in IT knows that demand always outstrips the capacity of IT to take on work. IT is in demand by all parts of the organization, often at no direct or hard cost to departments or individual staff, which can balloon workloads and exacerbate the growing backlog of requests.

    This capacity limit, or constraint, to do work isn’t the only reason IT continues as the big bottleneck in organizations. I believe the “silent killer” is the wasted cycles, or “overhead”, that comes with taking on too much concurrent work. Said another way, one of the greatest challenges of managing work is managing work-in-progress within IT. This is one of the reasons the Kanban approach brings great benefit to IT organizations struggling to stay above water.

    The Context Shift Tax

    We frequently pride ourselves on our ability to multitask but there is a cost, or tax, we pay when juggling too many work items. Each time we pause on one task and shift to another, time and effort is required to reset, reload and return to a point we can productively make progress on that work. Think of this as a context shift that functions much like paging between memory and swap space in a computer operating system. If you’re multitasking 5 items, each swap to a different task requires another knowledge context shift. Expand the list to 12 or 15 work items, add the interruptions we all experience, and we’re potentially spending more time during the day making those mental shifts than doing actual work.

    A legitimate argument can be made that over-multitasking increases complexity and decreases quality. As our mental bandwidth is split across too many work assignments, focus and concentration pays a price. We also experience a form of memory loss — some details are forgotten over time, even what the original problem was or why we are doing the task.

    IT organizations function much in the same way — shifting from project to project, task to task, as priorities change, new needs arise and existing work is delayed. As the list of priorities and projects increase, the organization and its people expend more and more on this context shift tax, wasting energies multitasking, losing focus, extending delivery times and decreasing customer satisfaction.

    Kanban – Keeping a Lid on the Multitasking Dragon

    Kanban’s secret sauce isn’t just about organizing work differently, it’s about thinking about work differently. Kanban provides a visual image of work in any system, how it flows, and where bottlenecks appear. The true secret sauce is in applying a limit or constraint to work in progress. If you’ve tried kanban, you know what I’m talking about. Much has been written about the profound effects of limiting work in progress.

    Using Kanban we can decrease multitasking, both by individuals and the organization. This reduces the context shift tax, freeing up cycles to maintain focus, perform work and complete tasks. I like to think of it as keeping a lid on the multitasking dragon.

    Limiting work in progress means work items are completed, in shorter time frames, with fewer context shifts. Imagine spending 5% or less during the day making context shifts, versus a heavy multitasking load that could consume 20% or more overhead when shifting between work. Less or no memory loss is experienced, decreasing complexity and reducing the opportunity for quality to suffer. By limiting multitasking we also free up cycles to focus on measurement and improvement, essential ingredients to any truly successful IT organization.

  • The DevOps Journey

    The DevOps Journey

    Maintaining a constant thirst for learning is essential for anyone dedicating their career to the IT, computer and networking industries. It’s more than just learning about the latest technologies, rather, it’s imperative we embrace new methodologies and paradigms for doing work, thinking about problems and creating solutions. Enter DevOps, which is a healthy mix of all the above.

    You might find things a bit confusing if you’re new to DevOps. There are many perspectives on what it means. For myself, there were two significant milestones on my journey to understand and embrace DevOps;  1) rethinking how we design large, dynamic applications and manage the underlying infrastructure in cloud environments such as Amazon AWS, and 2) reading the book The Phoenix Project by Gene Kim, Kevin Behr, and George Spafford.

    At first blush, I thought DevOps suffered the “definition bloat” of so many other industry buzzwords (for example, “cloud” comes to mind.) Depending on who you talked to or what you read, DevOps definitions frequently encompassed one or more of the following:

    1. embedding development resources within business units for better alignment and reduced cycle time,
    2. high frequency software releases into production (multiple per day, hour, etc.)
    3. automating infrastructure and system administration tasks vs manual provisioning and sysadmin
    4. embedding operations functions within development organizations
    5. operations performed by developer, even doing away with traditional operations in some cases

    If these seem like a somewhat confusing, loosely assembled list is ideas, don’t fret — they struck me the same way at first too. These ideas actually fit very well together around a few core ideas essential to DevOps.

    Quicker Time To Value – Being Fast and Agile

    DevOps is about changing traditional paradigms, enabling us to more rapidly deliver production updates, new capabilities and improved operations. Outcomes include reducing software releases from months to multiple, micro releases per day, and even hours. To do this requires much more than can be achieved through traditional operational efficiencies. It requires tight alignment, coordination and execution by product, development and operations functions. Embedding multiple disciplines together helps gain this alignment, shorten communication paths, setting priorities quicker, enabling lightning fast coordination. DevOps promotes new practices like embedding development and operations functions within other parts of the business, breaking down long standing barriers between IT / Ops and development / business groups.

    Automated Provisioning, Operations and SysAdmin

    From app development to infrastructure operations, DevOps is about automation. DevOps is about faster software testing and integration, quality improvement, and working in much more dynamic and rapidly changing operating environments. For example, rather than building system images by hand (stick built), automation is used to rapidly generate and spin up new virtual environments on a small or large scale. Systems wake up, check in to learn their personality, install, configure or update code (OS, database, web server, application code revs, etc) and join the network dynamically. This can occur with a single instance or with hundreds simultaneously. Initial upgrades and feature updates can occur on smaller or isolated sections of the network, then are rapidly rolled out across the network, and rolled back quickly if necessary. Sysadmins, be they developers or ops people, use scripts, code and automation tools to manage much larger environments than would ever be possible using manual methods. Automation is an essential differentiator in any DevOps environment.

    Culture of Open Collaboration, Communication and Failing Fast

    DevOps breaks down silos and places a premium on collaboration and communications amongst all players involved. It’s no longer us vs. them or this department vs. that other department. The blame game is gone and the focus is on making work happen, and fast. Communication lines are drastically shortened and conflicts in chains of command are collapsed. It’s about how “we” get things done, no longer about what “they” didn’t do. Failures are viewed much differently. Failures happen and from them we learn and develop newer, better, more innovative solutions to solving problems and prevention problems in the future. Because of the speed DevOps can bring, recovery from failure is much quicker. If a new release into production is causing issues, we quickly flip back to a previous known-good version running on other virtual instances. Updates can also be rolled out much more quickly.

    I hope these and coming views on DevOps are valuable to you as a reader. I know our conversation will help me and the teams I work with, as we learn and grow our expertise in DevOps. Please share your ideas, disagreements and learnings as we move along this journey together into the fast moving and quickly changing world of DevOps.