Author: Mitch Ashley

  • DevOps Chats: The Future of Farming with Aerobotics

    DevOps Chats: The Future of Farming with Aerobotics

    Spinnaker Summit 2019 Preview: Never underestimate an entrepreneur looking to solve a real-world problem. Aerobotics is bringing drones, imagery, mapping, containers, Kubernetes, Spinnaker, data pipeline processing and machine learning to agriculture. Sounds interesting, right?

    Aerobotics uses aerial drones and sophisticated image processing to help growers manage yield, pest and disease in their farms through artificial intelligence.

    Speaking at the upcoming Spinnaker Summit in San Diego, Aerobotics’ Head of Software Nick Coles, shares with us how his company evolved from making drones to map orchards into a software company using raw image data to create high resolution, georeferenced maps of orchards. Aerobotics technology is comprised of a map engine and tree engine. It’s both a fascinating problem space and exciting application of cloud-native technologies, including Spinnaker.

    This episode of DevOps Chats features a preview of Nick’s talk, “The Future of Farming with Aerobotics.” Nick’s talk is on Sunday, November 17 12:30 PM, at Spinnaker Summit 2019 in San Diego.

     

    Transcript

    Mitch Ashley: Hi everyone. This is Mitch Ashley with staging-devopsy.kinsta.cloud, and you’re listening to another DevOps Chat podcast. Today I’m joined by Nick Coles, head of software at Aerobotics. This is a special edition of DevOps Chat, because we’re talking about talks at the Spinnaker Summit 2019, which is in San Diego this year.

    Our next topic is the future of farming with Aerobotics, aka drones. So you’re gonna be really interested to hear about this. This talk is on Sunday, November 17th at 12:30 p.m. Nick, welcome to DevOps Chat.

    Nick Coles: Hey Mitch. Thanks for having me. Really excited to be on this podcast.

    Ashley: Happy to have you here. Tell us a little bit about yourself. Tell us about what you do at Aerobotics and a little bit about the company.

    Coles: Cool, so I guess just from the very beginning, born and raised in Cape Town, South Africa, as you can probably hear by the slightly strange accent. [Laughs]

    Ashley: Detected. We detected a slight accent there. [Laughter]

    Coles: Yeah. And but yeah, born and raised Cape Town, South Africa. Studied at the University of Cape Town, where I did a electromechanical engineering degree, which was kind of, you know, embedded systems, microcontrollers, processes and stuff. And that’s kind of why I got interested in drones and robotics. And, yeah, just over four years ago joined Aerobotics and sort of–at the time it had just been funded, so there were essentially two people, and I kind of joined in a software engineering role, not really knowing what I was doing.

    Ashley: [Laughs]

    Coles: And yeah, kind of been there ever since heading up the software engineering team. Which obviously a huge component of it has been DevOps and back-end processes. And so that’s kind of why I sort of have been exposed to Spinnaker and, yeah, kind of what led me to this today.

    Ashley: Excellent. So it’s the true startup experience, which means your job is you do whatever needs to get done, and you kind of evolve into your current role, it sounds like.

    Coles: Yeah, pretty much. I mean I joined–back in about four years ago when we joined it’s basically we were a company that was building drones, building the hardware, sort of soldering the motors to the propellers. And essentially drones kind of transformed throughout our journey, and we ended up having to move out of the hardware space and peel into the software space, because companies like DJI came up with these commoditized perfect drones. And it was quite hard to compete with them when you’re, what, five engineers in a garage soldering drones on a piece of [Crosstalk]–

    Ashley: There are a lot of drones available and you can differentiate in the software. So I read from the description of your talk that Aerobotics uses aerial drone imagery to help farmers, growers, manage their yield, pest, disease, that kind of stuff on farms, using some AI. Is that a problem that you started initially with, or did the company evolve to that kind of problem? How did you get started, and the company get started around that area, specifically?

    Coles: So, I mean, the company has actually had quite an interesting sort of journey in that regard. So the idea in the beginning was we are drones. We build drones. We use drones for agriculture. We use drones for security, game counting, mining. We tried to essentially just throw drones at any sort of problem.

    And sort of as we progressed as a company and more and more players sort of entering the field, and we sort of started honing in on our value proposition, and started specifically focusing on tree crops within the agriculture space because it’s such a high value. And so it actually then makes sense for the cost for specific tree crops.

    But originally we weren’t using any AI. We were kind of just trying to use raw sort of computer vision and object detection. The data that we were seeing is so vast, and case-by-case is so unique, that we kept having to change the parameters, slightly tweak things. And so the only sort of long-term solution could’ve been machine learning. And so that’s kind of what we embarked on about two to three years ago and sort of been, yeah, what we’ve been using ever since to try to essentially generalize the imagery we were receiving and sort of try to replicate the human behavior to the best of its ability.

    Ashley: Okay, excellent. So I know your software staff appreciate that about the company in kind of the broadened space. I know your software stack is a combination of, what, Kubernetes, of course you’re using Spinnaker, I think Circle CI for your CI/CD. Tell us a little bit about how you got involved with Spinnaker initially. Was that something that you chose or kind of evolved there kind of by accident? How did you happen down that path?

    Coles: So, I mean, it is–it did happen by complete accident. Traditionally when we were processing all this data we were just spinning up EC2 instances. And so we were essentially just sending SQS messages and spinning up instances. And instances would essentially be able to sort of, I don’t know, kind of retrieve these messages and then figure out what sort of processing to do. You know, it was an absolutely terrible solution for instance _____ a drone with a code. We’d have to SSH into the instance, patch it, figure out what was going on.

    Ashley: Oh boy.

    Coles: At this time also it wasn’t–like, DACA hadn’t really sort of like exploded. And or maybe it had in sort of like San Francisco and everything. But in Cape Town we weren’t really exposed to all these sort of advanced kind of tech solutions. So we had quite an archaic little approach going on, which was working, I guess, with R Scale.

    And then in the beginning of last year we got accepted into Google Launchpad, which was really exciting. Four of us went across to San Francisco, and we basically went on this launchpad program where they kind of helped us all the way from technical to business. And it really was inspiring to kind of go to San Francisco for the first time and sort of see like just the scale of things and just like the intensity of the tech companies there and really what they were doing.

    And we kind of got put in touch with a couple of mentors. And one of them, based on sort of what the problems we were trying to solve were and the problems that we were facing, recommended Spinnaker. And at that time we were kind of on AWS’ stack. And so we kind of then got incentivized to sort of try, set up a system on Google. And you know, we had like Spinnaker set up, Google’s Kubernetes’ engine, and kind of we just took it from there.

    And just it was an interesting journey, ‘cause there was quite a lot of iteration within like Spinnaker itself and in terms of just how the community’s progressed over the last sort of two years, but also just sort of like where we were at in terms of not really understanding how Spinnaker worked, the containers spiraling out of control, realizing we needed to stabilize our Redis pods. Like our Redis pod basically caused so many problems early on.

    But yeah, so that kind of was the journey in terms of how we got involved in Spinnaker. So it was basically just by chance somebody mentioned it, and we kind of were like, wow, this shows the problem we’ve been having.

    Ashley: Like a great solution.

    Coles: And we kind of just ended up, yeah, just going with it.

    Ashley: Fantastic. So from what I understand is the drones that–just giving our audience a little bit of context–the drones that are flying around the orchards, for example, are collecting all of this data. And that goes up into some data processing in the cloud, correct? And it does, I guess, splits into two areas. One is a map engine that creates–takes that raw imagery and creates multilayered geomaps of the orchard. And then the tree–a tree engine, which also then looks at specifically where exactly are the trees. And that’s how you do your processing. Is that a correct description of how this works at a high level?

    Coles: So it is basically quite a good description. So just to maybe give like slightly more detail on like the map engine side, we are–I don’t know if you mentioned–but we are using multispectral cameras attached to the drones. And from that data we can get information that isn’t necessarily visible to the human eye. And that kind of gives us all a health and powerful readings from the plots.

    Ashley: Is that like using infrared kind of things, or what?

    Coles: Exactly. Exactly. Infrared sensors. And so, yeah, so the map engine side is essentially like the big data processing side of things where we’re taking thousands of thousands of images and stitching them into these high resolution maps of individual orchards. I mean, just to give sort of some context on kind of what the data is there, we generally produce about like these five gigabyte high resolution images, which is of a specific orchard. And the kind of process around getting the raw imagery to that five gigabyte sort of image, we use about 64 gig RAM and 16 CPU instances. So that’s kind of just like the scale of what you’re running in. It can take anywhere between two hours to 24 hours to form that stitching process.

    And then moving onto the tree engine, this is sort of where our machine learning models are running in production. And that’s essentially taking these high resolution maps and segmenting the image to get individual trees from that image. And from that we can then calculate additional metrics, such as the height, the health, and the area of the individual trees.

    And one thing that, I guess, we left out because we are still developing it, is we now have a fruit engine, which is essentially now like kind of our new product focus area where we’re going from the drone, which is kind of taking a high level of the orchard, and now we’re going right up to the individual trees. And from those trees we’re starting to detect fruits and sizes from the fruits. And the goal here is to give the farmers an accurate sort of size distribution of the fruit in the orchard so they can plan accordingly when it comes to harvest time.

    Ashley: They’re all–it’s all about yield, of course, and quality of the yield as well. And by the way, I come from–I’m not a farmer, but I grew up in the middle of Nebraska in the U.S., so I’m around a lot of farmers. So I have some idea of what you’re talking about, at least an appreciation for it.

    Well how did–tell us a little bit about why you chose to do this talk at Spinnaker Summit and some of the things that you want to communicate beyond kind of the problem set or the solution that you’ve come up with.

    Coles: Okay, so one of the reasons why I thought it would be quite cool to speak at the summit was, I guess, just based off like chatting to some of the Spinnaker guys on the Slack channels and everything, it feels like we’re using it in quite like a unique sort of way. I mean, we’re using it very much to manage essentially our data pipelines. And I guess we’re almost using it in a similar way to how someone may use something like Apache Airflow or something.

    And I thought that the combination of the uniqueness of how we’re sort of using it, and sort of just to kind of explain how it has completely changed the way we’ve been able to develop and iterate super quickly, I thought it’d be worth kind of sharing that with the community and sort of seeing how other people were using it in comparison. And so I guess, yeah, that was kind of just the main thing, just the type of data that we’re sending through the system is quite different. It’s quite unique, in comparison to the rest of the world.

    You know, we process these large images which come up with a–which have a whole bunch of different problems with them. You know, we have to ensure that our containers are responsible, they’re not using too much RAM, they’re not spiking in terms of, you know, being out of memory because they’re loading an image all at once. And there’s just been a lot of interesting things that we’ve had to sort of develop to kind of handle our case. So I thought it’d be cool to share that with the community and take it from there.

    Ashley: You know, I think it’s super relevant, and not everyone may have an agricultural problem domain, but with a massive quantity of data that we’re collecting through sensors, whether it be drones or IoT and home or IoT and the business or any device, any systems that are collecting all this data, one of the real challenges is how do you process this much information? How do you combine it together? How do you take different sources like your map engine versus the tree engine, look at it, and then do that analysis and produce valuable, interesting results?

    I think a lot of businesses, whether they be startups or enterprises, have those kinds of projects and challenges, or at least they’re thinking about it, if they’re not already doing that. And I bet you’ve got some valuable lessons that you’ve learned over time about why you set up the pipelines the way you did or how to handle some of the production challenges.

    Coles: Yeah, you know, it’s also we had an interesting space, also, because we’ve learned so much over the course of the last two years. Like when you’re doing things in production, you also very quickly start figuring out what it should be. And so there’s even more and more of an end-goal, if that makes sense.

    Ashley: Mm-hmm.

    Coles: So like as a small example, kind of, you know, we’ve set up all of our Spinnaker pipelines, and it’s been amazing. And the devs have been able to iterate really fast. And now we’re starting to see that we want to maybe start introducing some stability to our pipelines. And so to achieve that you want to start setting them up as infrastructure code. And like just starting to build that whole process into the development cycle is obviously quite challenging and takes time. But it’s allowed us to sort of think forward and continually improve on what we’ve been doing.

    So you always feel like you have kind of found the best solution at a point in time. You’re like, wow, this is absolutely incredible, and it’s going so well. And then six months from then, you’re kind of like, wow, it could be so much better. There’s so much more we can still keep doing to improve again. So I guess that’s what’s been quite an awesome part about using Spinnaker. It’s like at every sort of milestones it’s kind of like you can keep iterating and making it better and using it more efficiently.

    Ashley: That’s excellent. When you use a tool like Spinnaker open source to a software like that, that you feel like you’re just getting started sometimes. You thought you knew what you could do with it, and you find out there’s eight or ten different things that it evolves to from where you started.

    Coles: A hundred percent.

    Ashley: Well good. I wish you the best of luck on your talk. You know, I think you may get some folks attracted just to come hear about the drones. I hope you’re gonna have some pictures of maybe some of the orchards and the maps and some of the output that you do, is that correct, just to give folks an idea of kind of how this works?

    Coles: Yeah, we’ll definitely, I think, do a very brief sort of couple minutes on the product and just kind of explaining what we’re trying to do, and for sure. ‘Cause I guess it’s always nice to sort of give context of what Spinnaker is actually producing at the end of the day.

    Ashley: Mm-hmm. Well very good. Well it’s been great talking with you, and I wish you the best with the talk. And hopefully you get a lot of folks coming that are excited to hear about what you’re doing. Is this the first talk that you’ve given at a conference like this, or you’ve done that before?

    Coles: No, this would be the first talk, so I’m quite excited for that.

    Ashley: Really? Excellent. Well you know it’s–there’s lots of ways to give back to an open source community, and of course, writing code is certainly one of ‘em. But most people don’t write code. And I think one of the great ways of giving back is doing exactly what you’re doing, is going and talking, not just in the community and online forums and Slack channels or whatever they are–that, too–but participating in talks like this.

    Because it–I’ll predict you’ll have a lot of folks come up to you after the talk and want to know more or talk about what they’re doing and tell you about their idea or ask your advice on something. So it’s a great way that advances the software and the use of it, because you’re doing things that probably other people never expected or predicted. So congrats. That’s awesome you’re doing it.

    Coles: Thank you. Yeah, I guess that’s one of the things I’m most looking forward to is just sort of meeting a whole bunch of new people and kind of, I guess, like we learn so much, like especially kind of being in South Africa, not kind of in the tech hub of the world, I would guess, is when you go there and you start speaking to some people who are kind of leading the way in terms of the industry, it’s so valuable just to kind of hear what they have to say and what they’re using and how they’re solving problems. Because it’s so relevant and it can translate so nicely back home. And so, yeah, it’s always so much value in just meeting the people and just chatting through things.

    Ashley: Well excellent. Nick, thank you for joining the podcast today.

    Coles: Thanks Mitch. Will you be at the summit in a month’s time?

    Ashley: I don’t think I’m gonna be able to be there. I’ve had several invitations [Laughs]. If I can I will be.

    Coles: [Laughs] Okay.

    Ashley: But I’m like back-to-back for the next month going to conferences. So we’ll see. Hopefully. If I make it I will definitely look you up.

    So I’d like to thank Nick Coles, head of software for Aerobotics, joining us today. Again, he’s speaking at the Spinnaker Summit 2019, the dates are November 15th to the 19th in San Diego. By the way, Nick, you’ll enjoy San Diego, too. It’s a great city like San Francisco is. His topic is the future of farming with Aerobotics, aka drones. And on Sunday November 17th at 12:30 p.m. So please check out his talk.

    And thanks to all you for joining us today on this podcast and listening to Nick and I chat about his talk. This is Mitch Ashley with staging-devopsy.kinsta.cloud, and you’ve listened to another DevOps Chat. Be careful out there.

    — Mitchell Ashley

  • DevOps Chats: Continuous Delivery at Airbnb

    DevOps Chats: Continuous Delivery at Airbnb

    Spinnaker Summit 2019 Preview: Airbnb is rapidly moving from a monolith Ruby on Rails application to a distributed SOA/Kubernetes architecture in Kubernetes. The new architecture uses self-service codified pipelines and easy webhook integrations, scale adoption and collaboration across the company. Even though continuous integration isn’t new to Airbnb, every team now needs to be able to scale CI across 100’s of containerized services in AWS EC2.

    Software Engineer Brian Wolfe co-led the decision to move to Spinnaker and build in more automation. At one year into the project, Airbnb has 40 services in production with many more to follow.

    This episode of DevOps Chats features a preview of Brian’s talk, “Scaling a Migration to Continuous Delivery (Airbnb)”. Brian’s talk is on Saturday, November 16 11:00 am, at Spinnaker Summit 2019 in San Diego.

    As usual, the streaming audio is immediately below, followed by the transcript of our conversation.

    Transcript

    Mitch Ashley: Hi everyone. This is Mitch Ashley with staging-devopsy.kinsta.cloud, and you’re listening to another DevOps Chat podcast. Today I’m joined by Brian Wolfe, who is a software engineer at Airbnb. Our topic is a talk that he’s gonna be delivering at Spinnaker Summit 2019 in San Diego. The topic of that talk is scaling of migration to continuous delivery. Continuous delivery is a hot topic right now. That’ll be happening on Saturday, November 16th at 11:00 a.m.

    Brian, welcome to DevOps Chat.

    Brian Wolfe: Thanks so much, Mitch. It’s really an honor to be talking with you today, so.

    Ashley: Well I’m more honored to have you on. I appreciate you taking the time. Tell us a little bit about you. Introduce yourself, tell us what you do and a little bit about–I think we know Airbnb–but tell us about what part in Airbnb that you work in.

    Wolfe: Cool. Yeah, so I assume most people who listen to the podcast know what Airbnb is, but we’re a worldwide community of hosts and guests. We bring people together to have really local experiences. I’m gonna skip the rest of that little spiel.

    Ashley: [Laughs]

    Wolfe: But I’ve been at Airbnb for about three-and-a-half years, and most of that time I’ve been kind of on the operational tooling side, and so a lot of that was observability works, so looking at your metrics and traces and logs, and then performance works, so like figure out how to make Airbnb–monitor Airbnb performance, make it faster and that sort of thing. And so but for the last year I’ve been the tech lead on our continuous delivery team, which is kind of an ambitiously named team. We’re not there yet. But the goal of this team was to really up level how we deliver software at Airbnb.

    And so my experience is mostly looking at how we’re making the transition from a monolith which was written in Ruby on Rails to an SOA over the last two years and kind of continuing on into 2020. And part of that is how do we scale how we deliver our software. It used to be a more core team who understood how to scale our monolith. And now every team kind of needs to know how to deliver that software, how to scale it and how to keep it working.

    Ashley: Now, is the move to an SOA, and so we mean service-oriented architecture, correct? Is that what you’re referring to?

    Wolfe: Yeah.

    Ashley: Is the move to that kind of architecture what prompted you to–or order Airbnb to–invest so heavily in this automated continuous delivery? Or was that kind of happening in parallel? Gives a little bit of idea of context of how this kind of came together.

    Wolfe: Yeah, so it definitely happened because of the move to SOA. And if you look at what our processes were before, it was very consolidated. Most people were applying the same code base. And so you kind of could have practices that worked for that one code base. But as we moved to SOA we now have hundreds of services that are being deployed, and you need to be able to scale those best practices across all of those services. And so that requires that you add these automation pieces so that humans can make fewer mistakes.

    Ashley: That makes a lot of sense. I imagine you have a kind of continuous integration happening at the front end of this, right, leading into your continuous delivery?

    Wolfe: Absolutely. So we’ve had continuous integration for a long time. And so all of our unit tests and that sort of thing and all of our builds are happening in continuous integration. That’s been the case for at least three or four years. And we have fairly mature tooling around that. And we actually did a migration to our new platform there fairly recently that containerizes all of our builds and all of our MCI so it kind of runs in a more uniform manner.

    Ashley: Mm-hmm. Is that moving to Kubernetes? Is that the change you’re referencing or something else?

    Wolfe: So even stuff that’s not running in Kubernetes, so as you can imagine Airbnb is about 10 years old, and so we have a lot of technical history. And so some of that stuff is gonna be running on bare EC2 instances, and some of that is running inside Kubernetes at this point. And so along with the move to SOA we’re also migrating to Kubernetes. And that migration is progressing rapidly. But there’s still gonna be a lot of stuff that has to be built on EC2 and running on EC2. But the builds themselves can happen inside containers.

    Ashley: I know we don’t think of Airbnb as something that’s built 10 years ago and evolved over time, right? We think of it as [Crosstalk] – [Laughs].

    Wolfe: Yeah, yeah. It has a surprising amount of history for such a ____ company, you know?

    Ashley: Well it’s an interesting, you know, it’s an interesting situation that you are involved in, because you’re not talking about a monolith application that was built 30 years ago or 20 years ago. It’s something fairly recent, Ruby on Rails, you know? So it’s still a contemporary programming language, something that everybody is familiar with. But even you have to go through your own kind of architectural and now continuous delivery evolution of that technology.

    Wolfe: Absolutely. I think it’s kind of surprising that even at a company that’s ten years old you have what we consider legacy parts of our stack. I think that happens quickly with a field that evolves as rapidly as ours. And if you look at how continuous delivery has evolved in the last ten years, it’s been a pretty big shift.

    One of the big things we’ve seen is, you know, we had a really good solution for deployment if you look in 2013-2014 time frame with our system and deploy board. But if you look at it now, it looks a lot like a really good CI system, like a really good continuous integration system. And the parts around actually having an organized pipeline that automates the deployment out, we don’t really have that in house yet. And so that’s really why we’ve been adopting Spinnaker as a solution.

    Ashley: Interesting. Well tell us a little bit about were you involved in the choice of bringing in Spinnaker? Was that happening while or before you joined Airbnb?

    Wolfe: Mm-hmm. So the decision for adopting Spinnaker happened last year. And so it was a discussion between myself, my manager and then Jing Jing and Jens, who are on the team. And it was largely about do we want to keep investing in deploy board, which was our existing solution. It was deeply integrated with all of Airbnb stack. Had a lot of stuff in it for how we deploy MonoRail, which is our big Ruby on Rails application.

    Now, do we want to build automation into that, or do we want to bring in something new? And so we did a big research project to look at what are all the available options out there, and what would those give us, and how could we make those work at Airbnb. And at the end of that we decided that we should not invest more in our internal solution and should instead bring in Spinnaker and customize that to kind of encode Airbnb opinions into that via extensions and then use that as our platform for deployment moving forward.

    So we’re still early days in that. We’re about one year in. We have about 40 services onboarded right now. These will be all production services, some of the really critical services out of Airbnb. But considering we have thousands of services at Airbnb, there’s some work to do to kind of make it the standard across.

    Ashley: Mm-hmm. Excellent. Well interesting. Tell us some more about I know in the description of your talk, you talk about that you’re gonna discuss how self-service codified pipelines and easy web hook integrations help you scale adoption and collaboration across the company. Tell us some more about what that is.

    Wolfe: Yeah, so one of the really harder parts about working at a company that’s growing as fast as Airbnb is its really hard to match the needs of everyone. And so you want to make everything self-service if you can so that, you know, I’m creating a standard pipeline that you know I deploy one thing, I do some AB comparison, I roll it out to my next environment, I do another maybe canary analysis rollout, and so on.

    So I can come up with standard components that people can use. But then you have stuff that people have built over time. And so our research team has regression detection that is really specific to them. And they think it provides really high signal if some functionality regression has occurred because you have things like search ranking and the actual layout of the results in the response package, and you need to make sure those things don’t change if you don’t intend them to change.

    And so they built a service to actually do that regression detection, and they want to integrate that with their deployed pipeline. And they want that to just be something that automatically happens every time that they run a deploy. You deploy to staging, you call this service with some arguments, you wait for that regression test to pass, and then if it fails you present some custom UI that says, “Okay, this is what failed. This is why. And this is how you can go learn about it some more.” And you move onto the next stage or you fail the pipeline and figure out what’s going on.

    What we wanted to do is make it really easy for teams to plug in services like this, ‘cause they do provide the most value. The way we’re doing that is by integrating, actually, with our interface definition language that we use internally. And so we actually extend Spinnaker to just provide a little web hook stage that people can reuse and call out to their services. But then they get a nice user interface on top of that. And so they can kind of provide just a base level of functionality within Spinnaker. And it’s kind of a minimal amount of work for us to onboard a new type of stage. Does that kind of make sense?

    Ashley: Yeah. I was just gonna ask you about that. Sometimes it can be a real challenge to introduce such a fundamental tool, like this is fundamental to your software development and delivery process and what kind of things that you can do to minimize the disruption, lower the bar in the area of entry for the teams. And a lot of it depends, too, on how much your organization does software process similarly, or is it very decentralized and everybody kind of does their own thing. It’s such an application if it’s been in a monolith kind of state. I imagine there’s a lot more similarities in how people work at Airbnb. But you’ve gotta still figure out how do you make this easy, right?

    Wolfe: Yeah. So I think we’re lucky in the sense that right now there’s really one deployment tool that people use, and that’s deploy board. And now we’ve introduced a second deployment tool. And our strategy here is actually we make the experience when you’re first onboarding actually look a lot like deploy board.

    And so Spinnaker lets you customize the user interface. And so we actually have a panel within Spinnaker that shows a view that actually looks really similar to the view that you get in our old tool. But then it provides all this power under the hood when you actually start your deploy. And so that’s been pretty valuable for us for onboarding because it doesn’t look that dissimilar, I would say a little rough around the edges, to be honest still.

    It’s, you know, we spent a lot of time optimizing that flow for deploy board. People get a little bit confused sometimes. But just having that initial experience look similar has made onboarding a lot easier. When we first proposed using Spinnaker, people were like, “I can’t figure out what’s going on in the UI. I don’t know what I’m doing here. What are you guys thinking?” But just by providing that similar viewpoint people are like, “Oh, I know what to do. I’ll just click on this button, which looks the same as we saw before, and that should start my deploy doing the right thing.”

    Ashley: You know, never underestimate the value of an easy-to-use user interface or process or whatever.

    Wolfe: Especially just having that familiarity, I think, it’s–Jens on my team has been a continual advocate for really pushing on, like, make this easy to transition from our old tool to our new tool. And I would say that there are some usability gaps that we’ve had to come across. But they’re getting better, I would say.

    Ashley: How far along in the rollout process of this are you? Are you in the still kind of beginning third of it, working with some of the early teams that are adopting? Are you now moving into more broader adoption of it? Where are you kind of in that process, that maturation process?

    Wolfe: Yeah, so we’re kind of in the–we’re calling it a beta stage. So in alpha stage we were really hitting some early adopters, people who were willing to take some risks. Now we’re onboarding kind of the big customers, the ones who are pretty demanding and making sure that this actually works for them. And so what we’re doing is we’re measuring how much do we prevent, and so how many rollbacks are we preventing, how many incidents are we preventing, and making sure that Spinnaker, as the platform we’re currently providing, is actually delivering business value.

    We’re limiting our onboarding throughout 2019 to about 50 teams. And at the end of 2019 we’ll have a really good idea of what works and what doesn’t. And in 2020 we’re going to open up the adoption to the full company. Kind of at the same time we’re going through a lot of scaling exercises to make sure that we will actually be able to scale to the full company.

    Ashley: And give us an idea of the 50 teams, what are usually the size of those teams? What do they range?

    Wolfe: Yeah, so the smallest teams will be probably about six or seven developers deploying a service. The largest team that we have onboarded has about 45 developers. And then we’re aiming for one of our front-end services, which is more monolithic, and so that’ll have several hundred developers working on it. But then there’s kind of an operational expert team that manages what that deployment looks like.

    Ashley: Mm-hmm. Well that’s a pretty good range in different sizes of teams. So I’m sure–and parts of the application–so I assume they bring different challenges with them. How do you measure success from your decision of both going down the path of implementing continuous delivery and Spinnaker? You talked about business value. You also talked about preventing how many rollbacks and incidents that also happened. But how do you demonstrate to the people that said we’re gonna spend this money having you go implement this, and tell them that it was worth it?

    Wolfe: Totally. That is a great question. And so one of the big motivators for us was actually for regression prevention. There are kind of two aspects for CD. One of them is increased productivity for developers. And then another one is having more vigorous processes for rollout. And we were definitely focused on the vigorous processes for rollout and having automated canary analysis by default for every service. And so that’s kind of the route that we’re pursuing, really, in 2019 in terms of proving the business value.

    Ashley: Mm-hmm.

    Wolfe: In 2020 we’re gonna have more metrics and more ability to say something around how productive we have made our engineers, or at least that we haven’t made the story harder. But right now it’s proven a little bit harder to quantify that. So we get people saying, “Yeah, we love it. We love that we can kind of ignore the deploy process now.” But aside from having a CSAT style score where you’re just saying, you know, rate how productive you are on a scale of one to ten, it’s been pretty hard to quantify the productivity gains we’ve seen with Spinnaker.

    Ashley: Well, that is often a challenge, right, for us in the software world in implementing technology, and as you’ve talked about some of the regression testing issues, maybe even incidents in the old process compared to the new, hopefully there are some metrics you can also use from the old methods to the new and to give you some places to start, at least, to benchmark those things.

    Wolfe: Certainly for regression detection we have. We have strong evidence that automating this pipeline–and so what I really think of Spinnaker as is a way to automate your run book. Just the fact that it is automated now has reduced the number of regressions that have gotten out to production. We have very strong evidence of that at this point.

    Ashley: Well and hopefully you can sing some of the praises, the good things of what the teams have learned so far with folks that are adopting it as you take on some of the bigger challenges like the larger teams do.

    Well now, I appreciate you coming on the podcast. It’s been fascinating learning about what you’re gonna talk about. And I wish you the best in your presentation at Spinnaker Summit.

    Wolfe: Thanks so much, Mitch. This has been a pleasure. And I hope to meet you at some point soon.

    Ashley: That would be fantastic. I would enjoy that very much. I’d like to thank my guest today, Brian Wolfe, software engineer at Airbnb for sharing with us his experience with Spinnaker and how they are migrating to a continuous delivery model and scaling that up in the organization. Brian’s gonna be talking at the Spinnaker Summit 2019, which is November 15th through the 19th in San Diego. His talk is on Saturday the 16th at 11:00 a.m. And I believe you’re gonna be doing that with Jens Vanderhaeghe? Is that correct?

    Wolfe: Yes, that is correct.

    Ashley: Awesome. Good to have you both up there. Well thank you everyone for joining us today. I’d like to thank our listeners for taking the time to check out the podcast and hear about this talk. This is Mitch Ashley with staging-devopsy.kinsta.cloud. And you have listened to another DevOps Chat with my guest. Be careful out there.

    — Mitchell Ashley

  • DevOps Chats: Autonomous Test Innovation Using AI/ML with Functionize

    DevOps Chats: Autonomous Test Innovation Using AI/ML with Functionize

    At the speed of DevOps, automated testing is essential for QA to maintain pace with that of software creation. Automated testing is ripe for innovation, too. How do we know we are performing the most relevant tests, that new functions in the software aren’t being missed or overlooked, or that highly dynamic applications aren’t outpacing static test cases and logic?

    Functionize aims to bring innovation to achieve autonomous testing, making testing, the creation and maintenance of tests more efficient. Founder and CEO Tamas Cser joins us on this episode of DevOps Chat to share several innovations his company brings to DevOps teams.

    Join us as we discuss how writing test cases in English help capture intent and result in less brittle tests over time, and about the ability to reach into Functionize during run time and programmatically change object models and internals. Learn how AI improves visual detection of web pages and changes in web page behavior, and how AI can root out test cases that may no longer be valid as software functionality changes.

    As usual, the streaming audio is immediately below, followed by the transcript of our conversation.

    Transcript

    Mitch Ashley: Hi, everyone. This is Mitch Ashley with staging-devopsy.kinsta.cloud, and you’re listening to another DevOps Chat podcast. Today, I’m joined by Tamas Cser, who is Founder and CEO of Functionize. Our topic today is a very interesting one—AI machine learning for automated DevOps testing. Tamas, welcome to the DevOps Chat podcast.

    Tamas Cser: Hey, Mitchell, it’s great to be here. Thanks for having me.

    Ashley: Well, absolutely. It’s my pleasure. Thank you for joining us. Would you start out by introducing yourself, tell us a little bit about you and also about Functionize?

    Cser: Absolutely. So, my name is Tamas Cser, and I’m the founding CEO of Functionize. I’ve been in technology for about 15 years, started as a consultant running my own company here in the Bay Area focused around development and we did, also, a lot of DevOps and building infrastructures kind of initially just saw posted and then cloud hybrid.

    And during those years, I bumped into testing a lot of times and really being at the front of the house and running the company and dealing with customers. It really become very apparent that testing is lacking in a lot of areas, and as DevOps got more and more initially mature and got traction, then we got into it more and, again, testing is a huge problem.

    So, really, that’s when I got into testing, tried to look at the problem, understand what opportunities we have now with cloud and big data that have not been explored and how that could be applied to solve some of these problems and this is really how Functionize was born about four years ago.

    Ashley: Mm-hmm.

    Cser: And ever since, it’s been an amazing ride. And now, you know, four years in, we have really acquired a great team and we have a very exciting product and customers and just, it’s been an incredible journey.

    Ashley: Great. Well, I definitely wanna hear more about the product and et cetera. Let’s start maybe with, in a software deployment world, especially in a DevOps world where we’re automating and we’re shifting things left, testing is only as good as the actual testing that happens. Bad testing automated is still bad testing.

    What are some of the hindrances that you see as we’ve changed to more of a DevOps style container, cloud? What kinds of things has that introduced into the testing criteria and how do you approach solving that differently than maybe a traditional, like with a Selenium style tool or something like that?

    Cser: I think that the largest change that comes with all of the incredible progress that we have made in software development, if you think about the entire DevOps toolchain and cloud and a lot of the automation that happens including containers is speed.

    And so, as with speed, software gets deployed much more frequently, and companies are able to iterate very, very fast around their product. And that provides a huge challenge for regression testing as well as getting testing on any of the new features that you’re releasing, and the maintenance especially around it is a huge challenge.

    So, I would say that, really, that speed that we have seen recently with some of these newer methodologies and then agile, multiple releases per day is really putting a huge pressure on the quality departments. And so, there’s huge opportunity to improve around that and really bring that up to speed, if you will, versus QA that has been kind of sitting stagnant for a long time, not seeing a lot of innovation if you look back the past 15 to 20 years.

    Ashley: Well, you know, it stands to reason, if you’re producing a lot more software much quicker, you need to test, obviously, much more quickly. So, automation, of course, is very important.

    What do you also think about the problem of how do you know to test the right things when you’ve got the right kind of coverage, and keeping in mind, you know, you may have developers, DevOps, engineers as well as QA people that are defining what those test criteria are?

    Cser: So, I think coverage is a very interesting topic and it’s an important topic that a lot of our customers have, how much coverage should I have, do I have the right coverage? And this is also an area where we can collect data and understand and analyze data from the application itself and even from the live user usage, in order to see where the patterns are, what’s being used, what potential impact it might have on the software if a particular feature or a functionality would break.

    And so, again, I think there’s a huge opportunity in bringing that—this is certainly one area that we’re focused on with our autonomous next generation capability where we’re trying to close the gap on that and have an impact on companies’ abilities to analyze and understand what, really, they need to test.

    Ashley: I believe, if I recall right, with your technology, your product, you actually write the tests in kind of an English format, right, rather than code? Is that correct?

    Cser: So, yes, we do have an English version of the desk creation capability. That’s not the only one, but it’s one of the key areas we’re innovating around is to understand user intent and what the test case needs to do and be able to take modeling around that, and that becomes really interesting and really important in how your test cases are maintained over time.

    So, these test cases become a lot less brittle for the very specific implementation and particular button or HTML element interaction. And so, that’s a really exciting area that we’re innovating in.

    Ashley: Excellent. So, I would imagine also, being part of a CI/CD process, you have to integrate with a lot of different tools or at least fit into a framework where, you know, you obviously don’t own the process end to end. Tell us a little bit about how you’ve approached that.

    Cser: So, the way we approach that—and again, this is a super important part of DevOps and building a piece of technology, these days especially, because the toolchain is incredibly important. And so, the way we approach it is two ways: One, we have generic APIs and give out of the box solutions to integrate with kind of your usual suspects—GitHub and Jira—

    Ashley: And Jenkins.

    Cser: And various different CI tools like Jenkins.

    Ashley: Mm-hmm, and Agira, yeah.

    Cser: The other piece that we’re doing that’s really interesting is, we’re—we have an ecosystem within Functionize that allows our users, basically, to program our run time environment that actually executes the tests.

    Ashley: Mm-hmm.

    Cser: And this opens up all of the object models and kind of hidden internals of the actual test case running. This is really, really interesting, because it gives you the capability to do a really, really deep level integration even in the middle of an execution with third-party applications or any other external functionality that you might have to perform that’s kind of a complex and advanced use case.

    Ashley: Now, am I correct in reading into what you’re saying in that while the tests are running, you can programmatically control Functionize and alter or adjust its behavior of what it’s testing or how it’s testing or something related to that?

    Cser: Precisely, precisely. So, this is kind of a first in the market capability and beyond just controlling sort of the outcome of the test case or the behavior, you also get access to the rich data sets that we’re collecting, including all of the visual screenshots or videos that we’re collecting doing execution.

    Ashley: Mm-hmm. And I understand, too, you use some AI machine learning in some pretty unique ways with screen captures, maybe other ways in the product. Talk a little bit about that.

    Cser: Absolutely, I’ll be happy to. So, as applications run again and again, obviously, how the application behaves and looks is really important. So, there’s one area that customers care about just around visual testing, which is, “Is my application displaying the way that I’m expecting it to and customers would expect it to?” So, it’s one key area that we’re working on and, with template recognition capabilities that can detect breakages on the page and this is based not on traditional kind of, you know, pixel by pixel computation, looking at the difference, but actual, deep learning that can understand how your page normally looks and where we normally see changes versus maybe change sets that looks like anomaly.

    Ashley: Mm-hmm.

    Cser: And then there’s another area—also, this is really interesting. Part of this that can be used in the root cause analysis, for example, of a page to understand what changed, what new elements were introduced and potentially have a test case that now is no longer valid could be automatically updated.

    Ashley: Oh, so, certain scenarios, maybe a page has changed sufficiently that the test no longer makes sense any more to run, that kind of a situation?

    Cser: Exactly. You mentioned that somebody added a couple of new form fields to a form and now those are required and you can’t complete the test. So, that’s something that we can do with analysis, visual analysis, as well as some DOM analysis can recognize an automatic usage and some solutions around it.

    Ashley: Wow, we’ve come a long way from the screen scraping vector graphics days of testing UIs, haven’t we? [Laughter]

    Cser: Oh, definitely. [Laughter] There’s a lot of exciting work that’s gone into it. And again, we are obviously drawing down on incredible open source capabilities as well and the advances that we’ve seen over the last year in an LP as well, so it’s really exciting progress that’s being made in the industry in general.

    Ashley: Excellent. So, talk a little bit about the cloud environment and, you know, we’re living in a world of, you may be a cloud-native application, but you may also be in a hybrid private cloud, maybe even not even a cloud data center, you know, private data centers and also multi-cloud. What are some of the challenges those environments bring and how might you tackle that?

    Cser: So, that’s a great question. Cloud environments in general are wonderful to work with for us and customers that are in the cloud and utilizing the cloud are good to work with for us. It does not represent any kind of a major challenge. It actually makes it much easier for us to work together.

    I would say that the cloud hybrid environments or on-prem environments certainly pose a challenge and there’s various different ways that we work with customers like that to still be able to bring the value of the cloud and the scale and kind of a lot of the modeling that we do and be able to still test applications that are hidden behind a firewall.

    Ashley: Mm-hmm. Is that also because of just the environments, maybe the testing tools, et cetera, that are in a traditional data center, private data center environment just are unique to that environment versus a cloud environment is more standardized if you’re in Azure or Google or AWS—is that why?

    Cser: I would say it’s more, it has to do with the access. It’s the way that we can access the application. It’s the client’s ability to provision environments that might be dedicated for testing, the way those environments can get the right resources, compute resources so it can handle the load that we will put on it as you go into agile testing. Versus in a traditional environment where the hardware is limited and potentially, you know, you have to deal with some secure tunneling into the system which is gonna slow it down. And so, overall, your speed decreases and the flexibility around those environments decrease. And the customer’s ability to provision new hardware, let’s say, because they would like to move faster for testing is greatly diminished.

    Ashley: And one of the things awesome about entrepreneurs is, oftentimes, they take a problem that they’ve had in their career, maybe their last job and say, “There’s a better way to do this. Let me go create a product, create a company, go do that,” since that’s part of your story, too.

    What was it that, in your experience with software and testing being a hindrance in and of itself? Were there specific problems or just, there’s gotta be a better mousetrap, or how did you—what did you bring with you from that experience to spin off and crate Functionize?

    Cser: That’s a great question. It really has to do with the pain, right? Getting the call from an angry customer when a particular bug regressed for the third time that week because your processes are not there—it’s really, really painful. And so, that’s really the initial starting point.

    The second piece that I would say that makes it a little bit unique is that—and we talk about this is, as well as shift left is really important so you can start testing early on. You’re taught it’s really important to shift right. It’s really important to test production environments because there’s a lot of things that happen these days when the application is literally compiled in real time in the browser with lots of real time dependencies, third parties and APIs and whatnot. So many things can go wrong in production that may not be actually code defects, they might have to do something with some other defect or some other dependency problem.

    So, that’s certainly one area that I am bringing Functionize in a heavy way that, and kind of the way that I see the future going, going both ways. We’re shifting left at the same time we’re also thinking about how do we test production.

    Ashley: Interesting. You know, you certainly hear a lot of shift left in test environments, not so much things happening in production environments. Are there certain, you mentioned applications kind of assembling, if you will, being brought together inside the browser at run time in a production environment.

    Are there other use cases like this for testing and production? I’m really curious about this.

    Cser: Oh, absolutely. There is many, and a few that I can think of that would be really interesting is a lot of the personalization that also is happening these days. So, we see customers struggle also with kind of these new style marketing capabilities that are operationalizing and these experiences that you would display to the user. It’s very difficult and challenging to test, difficult to see from the production environment. You have kind of the right data and the right experience, if you will, showing up.

    And so, that also provides opportunities for us to create specific products and features within Functionize to attack that problem.

    Ashley: Mm-hmm. It seems also like—this is, you know, sort of an edge case maybe today, we’ll see more and more of it, but serverless applications that are more event driven, you know, not traditional transaction driven might also be a good production test case for you.

    Cser: Oh, I agree. So, certainly, that would be an early, let’s call it corner case, rush case, but I think that we’re gonna see more and more of these coming as companies mature and these technologies become more mainstream.

    Ashley: Mm-hmm. Well, kinda going in a little bit of a different direction here, a little birdie told me that you all are up for some type of AI machine learning award—what’s happening with that?

    Cser: Well, we were just recently nominated for the AIconics Award here in San Francisco. So, we’re excited to be part of that. We’ll be attending and seeing how it goes, obviously. We’re honored to be there and being nominated. So, that’s very exciting to see a company being recognized for the work that we’re doing.

    Ashley: That’s awesome. Well, I wish you best with that and good luck. Also, I know there’s some things happening on the partnership for Functionize—what’s happening there?

    Cser: Yes, absolutely. We have a lot of great traction, and so, we are about to announce shortly, probably at the end of the quarter or early Q4 a very strategic partner, one of the largest global service innovators who are partnering with Functionize and we’re going to market together which is, again, I’m really proud of the team, and it shows the work that we’re doing and the work that marketing is doing to raise the awareness of the power of AI and machine learning and how that’s gonna change testing.

    Ashley: Mm-hmm. So, we’ll look forward to maybe some announcements coming up in the fourth quarter, then. Gotta keep an eye out for that. Great.

    Cser: Absolutely.

    Ashley: Yeah. So, I’ll kinda put you on the spot here, a little bit. If you had to say what some of the best practices that you’ve learned, not only at Functionize but also before when it comes to automated testing, you know, DevOps, cloud world—if you could boil that down to two or three best practices, where would you start? What would you tell people?

    Cser: I would say that finding the right tool and building the right toolchain is definitely incredibly important. The second I would say, like everywhere else, I would say you do need the right people and so, training and understanding of how do we apply these tools is absolutely critical.

    And then lastly, I would say that your strategy in that application, obviously, going back to the earlier conversation of what to test and how to test is gonna be critical for your success. And so, the tool, obviously, is primarily the area and certainly education is the second area that we’re very focused on, too.

    Ashley: I’m curious [Cross talk]—I’m curious about that, too. What are the kind of things that you look at or look for in people from testing skill or kinda the skills that you invest in current staff for training and skill development? What are some of those things that you look for?

    Cser: Can you clarify the question? In terms of look for in skill sets we’re hiring or skill sets—[Cross talk]?

    Ashley: So, your second item that you mentioned was the kind of capabilities and skills of your staff in testing and hiring those kinds of folks, also training for testing folks. What are some of those things that you’ve learned to look for that would be good advice for potential customers or current customers of Functionize?

    Cser: Yes, that’s a great question. So, I would say that, you know, obviously, like in any hiring, you want to hire bright people who are very motivated, so I’m assuming that’s the baseline.

    As far as the skill set goes, I think we definitely are looking at, obviously, professionals that have a lot of experience, but also, we can really enable and help people that are primarily manual or, let’s say, much less technical today and bring them into the world of automation and really drive a lot of value for our customer base. Because we see a lot of customers struggling with scaling and finding the right technical talent. So, I would say that that’s absolutely a critical point on being able to spread this to a wider audience, but at the same time I think that having technical users in part of the project, I think, is still very important.

    Ashley: Okay. Very good. Well, as always happens on these podcasts, we’ve run out of time. Tamas, thank you so much for being on the podcast today.

    Cser: Thanks for having me. It’s been a pleasure to be here.

    Ashley: My pleasure as well. Thank you to Tamas Cser, Founder and CEO of Functionize for joining us today and also, of course, to you, our listeners for joining us. You’ve listened to another DevOps Chat. This is Mitch Ashley with staging-devopsy.kinsta.cloud. Have a great day. Be careful out there.

    — Mitchell Ashley

  • DevOps Chat: The Kubernetes and Multi-Cloud App Journey with Rafay

    DevOps Chat: The Kubernetes and Multi-Cloud App Journey with Rafay

    After Akamai acquired his company, Haseeb Budhani decided to take on his next challenge and start Rafay.

    Rafay focuses on the complexity and repetitive aspects of deploying and managing Kubernetes applications. Rafay addresses complexity by providing application abstraction, cluster blueprinting and enterprise-ready integration, making Rafay a great candidate for multi-region, multi-cloud, hybrid and edge/MEC adoption.

    Join us on this episode of DevOps Chats where Haseeb shares some of the experiences and challenges that led him to found Rafay, and how to accelerate your path to Kubernetes and multi-cloud applications.

    As usual, the streaming audio is immediately below, followed by the transcript of our conversation.

    Transcript

    Mitch Ashley: Hi, everyone. This is Mitch Ashley with staging-devopsy.kinsta.cloud, and you’re listening to another DevOps Chat podcast. Today, I’m joined by Haseeb Budhani, CEO and co-founder of Rafay. Our topic today is lifecycle management for containerized apps. Haseeb, welcome to DevOps Chat.

    Haseeb Budhani: Mitch, thanks for having me. Great to talk to you.

    Ashley: Yeah, nice to talk with you. Always love talking with a fellow entrepreneur. Could you start by just introducing yourself, a little bit about your history, what you do and a little bit about Rafay?

    Budhani: Absolutely. So, Rafay the company is slightly under two years old at this point. We have a product available on the market that a number of customers are using which, essentially, helps them simplify the ongoing operations and life cycle management for their containerized applications, which they may be running in public cloud, hybrid environments, or any combination thereof.

    Fundamentally, there is a brunt of work that every company does as it relates to ongoing management of the containerized apps. You know, essentially, they build a platform internally on top of Kubernetes, and our pitch is, “Hey, why are you spending time rebuilding the—kind of reinventing the same build over and over again in every company? Let us help you with our solution which focuses on the repetitive effort that every company goes through, so that you, Mr. DevOps Engineer, can focus on the true value that your company is looking for out there. So, you focus on company value. We’ll help you solve for the repetitive work that you have to do, anyway.”

    Ashley: Very good. Now, you started this company after selling your prior company, Soha Systems, to Akamai, I think you took about a year before you jumped out and started Rafay. So, why did you pick this particular problem? What was it about it? Was it something you experienced in Akamai, in your prior company, or just as you were kinda taking that time to really think about what you wanted to do next? Why did you pick this space?

    Budhani: Yeah. You know, Akamai, it was, surprisingly, a lot of my friends who’ve kind of seen exits, you know, kind of, they talked about acquiring a company, it was tough, and I didn’t like the experience or whatever. I actually had a lot of fun at Akamai. It wasn’t as stressful as it used to be. I mean, you know how stressful startups are, so it was a pain-free experience, really. You know, an incredible sales organization. They really understood our product, took us to this whole new level.

    But, you know, sometimes an idea gets stuck in your head and it’s just impossible just to walk away from it. So, in our prior startup, so pre-acquisition by Akamai, we were practitioners of the problem that we are trying to address. So, essentially, we were the customer.

    Some of our colleagues at Soha were effectively building a platform like the one I just described, wherein we were kind of writing at this layer on top of container orchestration so that we could be running our application across many locations. So, at the time of exit, we ran—I don’t know, somewhere between 20 and 30 POPs globally across multiple cloud providers just so we could have our cloud based security platform at Soha running in a number of locations globally.

    And it was very, very painful. And I think at that point, we didn’t realize that when you’re building a platform, when you’re so focused on the core value which is the security, you know, this DevOps thing is important, but you don’t spend as much time strategically think about it. And that, after the acquisition by Akamai, Hemanth, my Co-founder at Soha and also here at Rafay, we spent a lot of time thinking about this. Like, why did we have the experience we did? Because it was not a good experience. And then we thought about, “Well, who else has this problem?”

    It turns out everybody has this problem where anybody building a SaaS platform or running applications in the cloud, particularly when it comes to containerized apps, because the community is so new, there is a whole lot of supporting cast, if you will—mythic. And it’s gonna come. It’s just, right now, it’s missing because it’s so new. And that got us thinking, “Well, somebody’s gotta solve this problem.” You know, and when you kinda look around and say, “Somebody’s gotta solve this problem” and nobody raises their hand, maybe you should do it.

    So, that got us thinking, “Well, we should—we need to go solve this problem.” So, we left. And I’m sure both of our wives were very unhappy when we said we were gonna leave Akamai. [Laughter]

    Ashley: [Laughter]

    Budhani: But—no, it was a worthy cause, and we’re very happy we left because we solved a really important problem for the community.

    Ashley: Well, congratulations for starting another startup. I think it’s one of those kind of in your blood type of things. So, talk a little bit about what’s unique about Rafay. I’m assuming you took some of the same problems that you experienced of what you’ve built at Soha and probably along the way said, “Hey, you know what? This might be a good idea for a product for containers and containerized apps.”

    And I think I would just add to the conversation by saying, as you start to build larger and larger containerized apps, you realize what some of the management challenges really are around it, and so that’s, you know, borne out of that is a new opportunity. I’m assuming you had a similar kind of experience.

    Budhani: Exactly right. You know, you read some blogs about, you know, like, Kubernetes. And it makes things look very easy. You can get something up on your laptop and you’re like, “Oh, look! I have a Kubernetes cluster running on my laptop! Isn’t this easy?”

    Ashley: [Laughter]

    Budhani: No. When you run it in production with multiple applications, it’s hard, and everybody knows it, this is not new information. But the challenge is that, because there is as yet no manual for this, we only learn by doing. So, we have to kind of jump into this and we start building stuff, and then we go, “Oh, this is hard,” and then by that time, we already have pretty large teams in most companies who are working on this. And that’s the problem and, of course, that is the opportunity, also.

    Here’s a better way of thinking about this. So, because of our experience at Soha, some things we recognized. Not everybody on our team we could train to be experts at the platform.

    Ashley: Mm-hmm.

    Budhani: Some people are experts and most people are not gonna be. And this is very typical in most companies, you have a DevOps team of five, seven, 10, 20 people who build a platform and then another hundreds or thousands of engineers, depending on the size of the company maybe consuming the platform. Because basically, they need to understand this—they need an abstraction layer. That was a key insight for us, having spent time talking to companies, you know, once we were out kind of looking at, you know, what do we do next?

    So, everybody seems to need some notion of an abstraction layer so that I, as a developer, I check in code and just magically, my application shows up somewhere. I don’t want to think about this underlying complexity. And the second thing that we saw, which VMware, surprisingly, validated two weeks ago when they did a bunch of announcements at VM World was that nobody’s running a single cluster.

    Ashley: Mm-hmm.

    Budhani: They may start with a single cluster, and because it’s easy to keep adding nodes to their cluster. But the practical approach around clusters is to run many small clusters. And VMware spent a lot of time talking about this idea of many small clusters.

    So, that implies solving for multi-cluster operations, multi-cluster federation is a critical thing. And that continues to be a gap in the community right now. There are a number of approaches that have been taken to solve for multi-cluster federation in the community, but—and, you know, a lot of companies have talked about this, there are tools that exist out there to solve this problem, but fundamentally, they make certain assumptions that are not borne out in practice. And that was a unique thing that we saw as a gap and we addressed it very elegantly in our platform.

    There’s a number of things we do, for example—I’ll give you another example. And anybody listening to this podcast, as you hear this examples, you may kinda look back to your experiences and go, “Yeah, I’ve faced this.” So, you bring up a cluster. Let’s say you are some e-commerce company. So, you need some clusters that need to be PCI ready, or PCI compliant. Some clusters, for marketing, because all they do is some experiments—I don’t need to be stringent for them. And then you have people who are just writing code, developers.

    So, you need a blueprint for a cluster, because each cluster may need different things. How do you do that? It’s not easy. It’s not easy. You know, there’s no kind of help chart for help charts. What do I do then? These are some of the problems, right?

    So, if you really spend time talking to the DevOps community, these are some of things that they struggle with and today, they all take a bunch of open source tools and try to build something for themselves. Some do it very elegantly. Some just never get to that level of perfection because they are so busy. And this is my pitch to every customer. My entire team, all of us—we do one thing. We do one thing. We help you run your containerized applications.

    Ashley: Mm-hmm.

    Budhani: You have 50 things to do. Let me help you with this, right? Let us be an extension of your team. Let us, you know, really take this one stress away from you—of course, pay money for that, [Laughter] but focus on the things that are more important to the company.

    Because getting the right script in place or whatever in place just so you can run your Kubernetes and run it better—this is not how you make money. If your competitors all ended up buying from a vendor to solve this problem versus build it themselves, they’re gonna probably sell a single more unit of widget that you’re selling, right? Focus on the value. Don’t focus on the very complex but effectively undifferentiated work.

    Ashley: Well, what you’re saying makes a lot of sense of why rebuild it all yourself? Why learn all the hard lessons that others might have learned? And I take it that you do things through your abstraction layer with your blueprinting of essentially kinda creating what we might think of in software as here’s patterns that we might use, here’s blueprints we might use for different clusters.

    Budhani: Yeah.

    Ashley: And I know a common mistake I’ve seen happen is, it’s really easy to throw everything into one cluster and then start to say, “Well, that doesn’t work. How do I start dividing it up into, besides just geographic location, what are the other techniques for doing that?” So, those are the kinda things that you’re bringing to the table with Rafay, is that correct?

    Budhani: Yeah, absolutely. Yeah. I mean, I’m sure many of us continue to have the Design Patterns book from college on our shelves, right?

    Ashley: Sure, yeah.

    Budhani:[Laughter] That’s the idea, right? Some things are best done in a specific way, and if there’s a way to do that, if it’s a best practice way to do it, you do it. Not every company may have their own quote-unquote blueprint, and that’s okay. But let’s help you manage those better.

    In effect—you know, and this is fundamentally, we talked about, practically, what Rafay is selling today. But if we fast forward three, five years, we sell something today, but we still have to have a raison d’être, right? Why are we doing this? I think, fundamentally, there is a system of record needed in our industry that maybe doesn’t exist as well as it could. Where is that one place in a company where I can go and say, “Hey, where is ____ right now? How is it doing? Three months ago, how was it? Who changed what last time? Where is that information today?”

    It’s in people’s heads, sure. It’s in a Git repository. We all talk about GitOps, right, which means somebody wrote a script and checked into a git repo. Yeah, but maybe the guy who wrote that script left six months ago—now, who understands that script? That is a very common occurrence—people write scripts and then they move onto the next job. Where is that system of record? Where is that system of governance for an application?

    Particularly with containers, there is an opportunity to very elegantly address that problem, right? Have a system that allows you to do that. One system where somebody who is non-technical can log in and say, “Hey, just from a compliance perspective, is this app running on all the right clusters right now, or not? One shot, can I know that?” That is where I think we will all go as a community.

    And as I kinda think through what are those bread crumbs to get us there, some of the things that we are working on as a product company today, my hope is that we get there. We get the opportunity at Rafay to build that system of record that people will really treat as a strategic system for their own companies. Anybody in the company should be able to come and look at, “Hey, where is my—what is the status of my app right now? How has it changed in the last month?” That, to me, is such an important and worthy goal to be focused on. Let’s see how—you know, what it takes for us as a group to get there. But that, man, that—I mean, somebody needs to go solve that problem.

    Ashley: Now, I haven’t used your product, but one of the things that I found really intriguing about your go to market and how you talk about it on your website is, you actually present use cases, more than just generic things, but something specific to, like, factory operations for IoT and customer experience app, modernizing in store retail experiences, 5G, edge deployments. Are those part of the kinda blueprints of what you talk about, or is that more, you know, on a broader scale, “Here’s how we might advise you about how to implement a containerized cluster environment?”

    Budhani: So, these are the business use cases that are—the last one we talked about was this retail use case. It’s a fascinating problem. Many retailers are out there talking to larger vendors and in some cases to their ____ providers who happen to be large customers. And they’ve asked the question, “Hey, I want to deliver a better experience inside my store. Is the point of sale system in the retail store up or not?”

    That’s a good question. How do you find out? We can’t run these tests from out of the store, because that’s a private network, basically. So, I wanna run an app inside my store. Oh, yeah? Where do I run it? Where is this infrastructure? This is not a—there’s no data center, here. They have probably one rack in the back somewhere, maybe with one or two machines. So, then maybe I should run containers here—but then, who’s gonna manage my containers, because I have 1,000 stores, it’s not one. And if I have to send IT people to every store to deploy an application and upgrade it regularly—well, I need an army now, right?

    It’s just a small question. I need to run a small, tiny, 200 ____ app. It becomes this big thing. What an interesting problem. But this is a real problem. Think about any retail location where they have a lot of refrigeration capability. They’re selling ice cream or they’re selling whatever drinks and whatnot. If any one of these refrigeration units is not working, you lose a lot of money.

    Ashley: A very big deal.

    Budhani: And, of course, it’s a bad—right? So, they run these sensors, okay? Who’s collecting this data? Where is the sensor running? Is it a tiny little sensor? But then the sensor data has to be collected maybe locally. Okay, what do I do now? Right?

    All of these small things, these are real business problems that kind of snowball into this bigger problem around application management, right? So, if I—if you consider, if we called a retail store or some head office for some retailer and said, “Hey, do you wanna buy a container management platform or life cycle management containerized app?” Like, “What is a container?” Right? [Laughter]

    Ashley: Yeah.

    Budhani: Their people know, but they may not know, right? But they understand this problem, right? It’s like, “I need to run apps here, because this is how I’m gonna make more money. I want to provide a better online and offline experience for my customers.” That opportunity, right?

    So, that’s why, as we hear use cases, as we partner with different companies, and we see very specific and repetitive use cases come up again and again, our hope is that we continue to list them on the website. So it’s more meaningful for even the business or more kind of—you know, non-technical people in the company to come to a website like this and understand why, you know, their engineers are talking to us, for example.

    Ashley: You described yourself as a SaaS based company, it’s a SaaS service. Do you also provide any operations support, kinda outsource of operations, too, or is it strictly a service that others then operate on top of your SaaS?

    Budhani: So, it’s a service, and we do provide kinda more of a solution architecture model ready to essentially assist in our customers thinking through how they’re gonna consume a platform like this, right? Because a platform like this is not—it’s not a box that you plug into your network, if you will. This is a strategic conversation where we’re thinking about how to embed it in some way into your processes or into your, you know, own organization.

    So, there’s some level of consultative conversations that are happening up front, definitely, and on an ongoing basis, also. But from an operations perspective, we don’t see that when Rafay is fundamentally a software company, you know, a lot of pre-sale support on the solution side. We are continuing to engage with a number of companies who excel at that, who look at Rafay as an enabler not just for themselves but for their customers to whom they’re delivering a broader service. So, think systems integrators, DevOps consulting organizations who, whether Rafay exists or not, they are solving this problem and they have other customers, right?

    Ashley: Mm-hmm.

    Budhani: So, those companies look at us as an accelerator for their own services and for their own offerings. So, that’s been how we’ve been approaching the market, so we can obtain a high margin business and be, effectively, a software company. That’s what we know how to do well. But there’s so many great organizations out there who have an incredible skill set as it relates to operations and integration. That’s where we step back and partner with somebody, and at that point, be our tool in the tool chain and they go deliver a great experience to their customers.

    Ashley: Very good. That helps a lot. So, talk about who your most ideal customer would be. Is it a large enterprise, medium to large? What kind of an organization would say, “This is exactly what I need—come help me”?

    Budhani: In theory, this is a horizontal problem. Anybody who’s doing, you know, running containerized applications has this issue. But fundamentally, what we, at different stages in the company, look for different kinds of companies in terms of just, you know, low hanging fruit, et cetera.

    What we find works well today is, you know, relatively mature organizations who are doing somewhere between a couple of hundred million to a billion dollars of revenue, they are modernizing their applications, moving to the cloud, becoming more and more cloud native. And they’re the ones who are very interested in going fast, because there’s some—usually, there’s some business reason why we’re doing this. They’re not just modernizing for the sake of modernizing, there’s competitive pressure, there’s cost pressure. They may need to be in different parts of the world, et cetera.

    There’s reasons why they need to be up and about quickly, and they will look for any shortcut to get them there. A shortcut in this case would be, you know, “Hey, I could build this myself. Should I?” Right? And what I just said, right? This is a question we, in some way, shape, or form always want to highlight to our customers, which is—hey, you have a really strong engineering batch. You can do anything. Should you do everything? Do you still build data centers? And that, of course—that’s a, you know, asking the question, “Are you still using, building data centers?” that’s a straw man, of course. Because, hey, I mean, maybe other than 100 companies in the world, nobody’s building data centers any more.

    Ashley: Right, right.

    Budhani: It makes a point. In our industry, we continue to go up the stack in terms of abstractions, right? I mean, all the way from assembly to op and now nobody’s building subnets any more. Private subnets, published subnets—you just make an API colony and you get one in Amazon, right? Why am I building this stuff?

    Similarly, we are providing them yet another level of abstraction which allows them to move faster. Because fundamentally, the name of the game is, “Hey, if I don’t move fast, if I don’t produce product fast, I’m gonna lose.”

    Ashley: Now, I know you worked with AWS as your et cetera VMware. Is there a particular environment that you’re best suited for? Like, if you’re a VMware customer, you’re gonna—we’re gonna connect right into your environment, or is there a technology stack that’s more easy for you to integrate than others?

    Budhani: So, this is a very agnostic platform. I mean, VMware environments can be anywhere. And I’ll highlight a specific thing that we’ve thought through and we implemented as part of the platform.

    So, let’s assume for a minute that it’s a mature organization, they do have some hybrid environments, and they maybe have some Tectonic set up and they’re running VMware there, and they also have Amazon, and that’s okay, right? So, in one environment, maybe they’re using PKS from VMware, PKS Essentials from VMware. But in Amazon, they’re using EKS, because that’s what Amazon sells for Kubernetes.

    Okay, now what do I do? Am I gonna build two sets of platforms to support two different environments? Not a lot of people do that. Or, with a platform like ours, what we say is, “Hey, just—look, if you have an opinion on Kubernetes, go with it, and essentially, we’ll add value on top of your existing Kubernetes throughout the operator.” It’s a Kubernetes concept called an operator.

    Ashley: Mm-hmm.

    Budhani: Or, if you don’t care and if you say to us, “Hey, you know what? Just bring me the, I don’t know, whatever the latest upstream Kubernetes is, 115 or whatever—just bring that.” No problem, we’ll work with that, also. But, for our SaaS platform, to help you manage your environments behind a firewall in a data center or behind a security group in Amazon, you don’t ever happen to open a single portal in the firewall for us to get in from outside. The entire system is designed to be secure such that there are sessions being launched from inside out, so your clusters run an agent, basically. They reach out and broker all the connectivity without making a security issue kind of pop up on the security team’s radar.

    So, this is another thing, by the way—anybody listening to this podcast thinking about some sort of a multi-cluster or multi-cloud solution? Please let’s make sure that you’re not signing up for a vendor who’s asking us to make maybe not the best decisions when it comes to security and asking us to open up ports, setting up IPsec links—you know, all of these are bad ideas that lead eventually to other problems that you haven’t thought about today.

    So, these are some of the things where understanding our enterprise security, understanding our enterprise problems are important up front, so we thought through these things and we designed a platform to be able to work not just in any environment, but to work in these environments in a very secure fashion.

    Ashley: Well, certainly, I think, with all your experience, you can help de-risk things and provide kind of a framework or a path for enterprises.

    Well, unfortunately, we’ve run out of time. See, I feel like we could talk about this for another hour and a half and share a lot of good stories about it, too. I’d like to thank you—thanks for being on the webcast, the podcast, here.

    Budhani: Yeah, thanks for having me. It was a really fun conversation. I look forward to talking again soon.

    Ashley: Yeah, I do, too. Keep us informed about any new news. So, I’d like to thank our guest today, Haseeb Budhani, CEO and co-founder of Rafay. I’d also like to thank, of course, you, our listeners, for joining us today. This is Mitch Ashley with DevOps, and you’ve listened to another DevOps Chat podcast. Be careful out there.

    — Mitchell Ashley

  • DevOps Chat: Value Stream Management in the Enterprise with Plutora

    DevOps Chat: Value Stream Management in the Enterprise with Plutora

    Doing anything at enterprise scale brings with it a whole new set of requirements. With potentially thousands of requirements across hundreds of applications and releases, it can be a daunting challenge to pull together all the DevOps activities across a large enterprise. This has given rise to the new term Value Stream Management (VSM.)

    Like Salesforce for managing the sales pipeline, VSM is about IT seeing across all its initiatives, teams and projects. Are high-value requirements making it into production systems? Is DevOps making the expected impact at enterprise scale? Are VSM tools well integrated into the DevOps toolchain? Or are teams slowed by extra or unnecessary work? You must keep a handle on it all but we don’t want the medicine to be worse than the disease.

    In this episode of DevOps Chat, we explore VSM with our guest Bob Davis, CMO at Plutora. Plutora tackles the VSM challenge at some of the largest banks, financial institutions, telecom and mobile carriers, insurance, hotels and internet technology companies.

    Transcript

    Mitch Ashley: Hi, everyone. This is Mitch Ashley with staging-devopsy.kinsta.cloud and you’re listening to another DevOps Chat podcast. Today I’m joined by Bob Davis, chief marketing officer at Plutora. Our topic is an interesting one. We all talk about DevOps in organizations, but how do you achieve high performance DevOps. Bob, welcome to DevOps Chat.

    Bob Davis: Hi, Mitch. It is great to be here. Thanks for including me in your agenda for the DevOps talks.

    Ashley: The honor is mine. Appreciate you being here, Bob, definitely. Would you start–have you introduce yourself. Tell us a little bit about what you do as chief marketing officer and just a little bit about Plutora.

    Davis: Absolutely. I’m a two and a half year veteran of Plutora, having come in shortly after Plutora got its first round of funding, which was very exciting for all of us and we really were launched on the path of providing a better outcome for folks transforming into DevOps and that’s what Plutora does. Plutora is a company that provides a SaaS platform, large enterprises, your typical brick and mortar of the past, banks, the insurance companies, et cetera.

    These folks come to us and we work with them to provide the ability to visualize their entire pipeline to become better software developers ultimately and to corral the morass of tools that they’ve assembled in their transformation and really make sense out of them, make them efficient and make the organizations successful, and that’s what we do.

    Ashley: I’m especially interested to talk with you on this topic because when you look at your customer list you have some very well-known names in Fortune 500, Fortune 1,000. It isn’t just a bunch of other startups. It’s people that other–

    Davis: That’s correct.

    Ashley: –companies and other people would recognize. So handling financial transaction, a lot of that, telecommunications, et cetera. So you come at this with a unique perspective and when we think about DevOps we’re thinking about that central core of team of bringing developers and operations together, but to really just scale it across an organization at an enterprise level, that’s a much different problem. What do you see as the challenge as you work with customers? Why do they come to you saying, “We can get it this far, but we need to get further and here’s what we need to solve”?

    Davis: It is no question that that particular request from prospects and customers alike is the number one thing we hear, that is they’ve embarked on the transformation to DevOps. They’ve started maybe with automating testing. They’ve started by providing maybe a CI server and that automated deploying capability somewhere. That’s where they’ve begun their journey and they started with that automation and they’ve immediately found that while they’ve got gains there on that part of the developmental pipeline, the overall performance did not increase as was expected. They’re basically standing there going, “What am I doing wrong?”

    The big problem people have really surround this word visibility. The lack of visibility results in tremendous issues all around the development process, not just for developers, not just for ops, but for every one from the program office to the security guys to the governance and audit people and so forth as we deal with people where those kinds of things are really important. Traceability, auditability, these are–software is running very sensitive businesses and those things are very important.

    So without visibility, in order to build those things in messes up my DevOps process. I’m going along the path. I’m ready as a developer to move into a test cycle and all of a sudden the security guy says, “Hey, I need to check that for security issues,” and he’s coming to the party late and all of a sudden there’s a two week delay. People get frustrated and CIO’s are saying, “We’ve invested millions in tools and we’ve invested millions in organizational changes to adapt to agile and DevOps. We’re getting worse, not better, why?” And that’s the problem.

    Ashley: When you say visibility, to me if I peal that word apart, I think what you’re really saying is it’s not necessarily visibility by the DevOps team itself always, but it’s the intersection with all those other parts of the organization, the policy parts of the company, the audit part of the company, the security part of the company. It’s all the other people that are involved in that business tool chain, if you will, of getting product out the door. You can’t just do it alone. Maybe yes, if you’re a three person in your basement start-up, sure you can. But when you’re talking about scaling this in an enterprise level, they have processes for stopping and checking things. Is it ready to go in production? Has it been checked for security? All those things. Is that what you mean by visibility?

    Davis: It is and part of understanding the scope is to understand the scale of what these enterprises do. We talked to one group in a large insurance company in the Midwest that had 35 applications, that’s it. That doesn’t sound like a lot.

    Ashley: Yeah. It doesn’t.

    Davis: But that week, the week I was there they had 1,600 releases in process to be released that week. When you peal that down and what are 1600 releases in the context of 35 apps? Thirty-five apps, an application might have a whole portfolio of releases that need to be managed. So this visibility of the understanding of where my application is in the development cycle from the time that it starts to be an idea to the time it releases.

    So that process to know what features are being released, where they’re being tested, what versions are they being tested with, when are they ready for a user test, when are they ready for preproduction final test, all those steps along the way have different tools, different teams often, required shared resources like environment, and they depend on other teams in the group and other people’s efforts within those teams to meet at the end point at the same time because, like I said, a particular release is going to have multiple different teams, different project teams.

    So the visibility is across an entire pipeline first of all from the beginning of requirements to the time it gets released, but the visibility is also between pipelines to understand where dependencies are, to understand when other things that might affect my ability to release, like security, like audit, how those things come into play. I’ve got to have visibility across all that so that I can participate and predict when something is going to come out. That’s really important. If I’m building–for example, if I’m building a check cashing app and I’m a major financial institution, part of my development team is an iPhone app and I’ve got a team that’s building that and they’re following the rules of Apple and they’re following the rules of UI and they’re looking at how I can do forms and they’re doing all that.

    On the backend there’s a data base that’s probably sitting on a main frame which has incredible security around it and it has to be upgraded to be able to provide the ability to process that automated mobile cash check. And that’s probably being done, like I said, on a main frame with a different development cadence. Maybe it’s got a release schedule that’s once every six months and this is app is supposed to be out in two weeks. How do I coordinate that and how do I bring the security people in to do that right?

    So in order for that to be the case I’ve got to have a single system of record. I’ve got to have automated disability and dashboards to alert people as to when things are coming. I’ve got to be able to take information that’s being generated from multiple different places, multiple different applications, multiple different application methodologies, and correlate them in a way that gives me concise single page dashboard views of where things are so I can answer the question, “When will this release go out?” So those are all the kinds of things we do.

    Ashley: Well, let’s take that word visibility and turn it another 90 degrees and look at it this way, which by the way, I think most enterprises would kill to have, what’d you say, 36 apps? That’s a side point. Anyway, if you took visibility–

    Davis: Well, that was one–just to be clear, that was one group. That was one group.

    Ashley: Oh, one group. Okay. I thought you meant the whole enterprise. I’m like wow. I know a lot of companies that ____ “How do you do that? Where’s the course for that? Sign me up.” So anyway, take that visibility word. I think a big part of what you’re talking about is it’s not just a tool chain for CICD going from development all the way through deployment. It’s also the planning, the scheduling, the coordination of releases. It’s coordination of getting information to people who maybe it’s not just the information, but the exceptions ’cause you’re doing that many releases that frequently that often.

    You don’t wanna read 20 reports every day on the security compliance of every release that’s going out the door. You wanna see what’s the five things I need to be aware of that’s holding something up. So there’s a lot of coordination planning and scheduling that kind of activity that you guys focus on that’s I think more unique in what you do.

    Davis: I’m really glad you brought that up, Mitch, because that’s something that’s coming on the horizon quickly, but has typically not been considered. When people talk about DevOps, even Agile and DevOps, when they talk about software delivery pipeline, very often what they’re actually talking about is check-in to production. The clock starts when I start to check my code into my source code repository. All development tools, whether it’s automated testing CI servers, CICD pipeline, management systems, et cetera, they’re all built around that second half and backend of the process.

    To me the real issue goes back to the initial requirements planning. There’s some really nice tools about – on the enterprise agile planning side that have come up that are pretty mature. The PPM market itself has been around and good long while and there’s some really interesting and great products that allow me to establish the priority of my development work, establish the priority with my development teams, and to do that based on the strategy of the company, the budget allocations of the company, the expected value returned by the applications and projects in question whether it’s revenue or whether it’s cost savings.

    And I set that up and I say, “Here’s my portfolio. This is the priority. I’ve got one ____ applications that I’m going to be driving. That’s my priority. Go.” And the teams all drive and they’re confident they’re building according to the priorities of the company. They enter into the pipeline and I lose visibility. Boom. It’s gone. While that product is being developed things happen. Scope has to change. I might have to pull features out that are particularly thorny, from a development point of view, in order to meet my schedule. Decisions have to be made.

    Do I pull that out or do I keep it in ’cause it’s so critical and I accept a delay? How do I answer those questions if I can’t tie that back to the epic and the strategy, the theme, if you will, on the safe vernacular? How do I answer that question unless I can unequivocally understand how that connects back to the original idea of the application? The input is clear. I prioritized it right. Am I really ____? Am I getting better? Are people using the product? How are the NPS scores and how do those impact?

    There’s all kinds of questions that you can ask that really define whether I’m doing better ’cause fast for fast sake is not that interesting. Fast to get to the value faster, big deal, and understanding that value. That’s why I year ago Forrester started to talk about value street management, the measure of success is ultimately is the business benefiting from the effort. Without the kind of visibility we’re talking about that whole effort is compromised.

    Ashley: In different capacities I work with a group of CIO’s. I work with developer community. And a common theme across that is now that I think DevOps is taken the barriers of the organization, something that we all get rid of and start working together more closely in a true value chain or true pipeline chain is all the communities are struggling or want to know more – how do they more effectively communicate with the business side and with the C suite.

    And not to make this about me, but I was on a podcast here recently and that question came up and I said, “Well, if you’re a developer you want your CEO’s attention, say ‘I have an idea that will increase revenue, bring us new revenue streams, cut cost, get us to market faster,’ any one of those business terms and they’ll immediately put you on the calendar.” I think that’s where you’re going and where the analysts are going with this value stream management.

    Davis: Totally. That’s exactly where they’re going and that’s the key question that people wanna answer. It involves velocity. It involves quality. It involves all the things that people wanna talk about with respect to DevOps, but it has to be done in the context of an awareness of the impact on the business. It’s – in many respects when I first came to Plutora and I was talking to the founder and CEO about what Plutora did, my first reaction was, “This sounds an awful lot like Sales Force for IT.” The problem is very much similar to that problem.

    System of record, like what Sales Force provides, inside this platform, our platform, for example, that connects all the dots and automatically serves up the kind of metrics that the C suite wants to find out about what’s going on with the developer not having to change the tool they’re working on. If they’re using Jira, if they’re using Jenkins, if they’re using Selenium, whatever tool they’re using they can continue to use that tool, but the platform, the management platform will take all that information and correlate it and provide visibility to management on exactly what the value is that’s being delivered. and the metrics of what’s the cycle time of getting released.

    When I have a bug, how fast to fix it? If it’s a said one, what’s my meantime to repair? Am I getting better at quality, et cetera? These are all the kinds of things that guys wanna know and they wanna compare the scope of release that went to production with the scope of that release that went into planning and see how well they did on all kinds of levels. To be able to have the information to the C suite and everybody in between flow up to them naturally, without having the practitioners do anything out of ordinary to their normal day to day jobs, you have this perfect scenario.

    You have this catwalk, if you will, across the software manufacturing process that gives the management the ability to see it without having to change the way the developers go and the information that can be served up, I can establish stakeholders. I can have a security stakeholder, the audit stakeholder, whatever my organization unique needs are. Let’s be clear, software is the competitive differentiator of organizations these days.

    Ashley: It rules the world, right?

    Davis:It rules the world and it’s–yeah. It’s for better or worse and as it becomes more of the thing, you can’t just be a bank that happens to use software. You have to be a software company that is choosing to be in the banking business.

    Ashley: Yep. Technology company.

    Davis: And if you–yeah. You’ll be out of business.

    Ashley: Well, Bob, I appreciate you spending some time on the podcast. I know we could go a lot longer talking about this. Maybe sometime at Half Moon Bay at Barb’s, there we can grab something sometime.

    Davis: That sounds like a great plan, man. Any time. Any time you can always hit me up there.

    Ashley: Food and something healthy to drink, yummy. Yeah. Of course.

    Davis: There you go.

    Ashley: Well, I’d like to thank you, Bob Davis, CMO from Plutora, for joining us today. You’ve listened to another DevOps Chat podcast, I’d like to also thank you, our listeners, for joining us. This is Mitch Ashley with staging-devopsy.kinsta.cloud. Be careful out there.

    — Mitchell Ashley

  • DevOps Chat: From 0 to 1,000 Deploys Per Month with Redbox

    DevOps Chat: From 0 to 1,000 Deploys Per Month with Redbox

    Many organizations want to implement the latest DevOps/SRE practices, but many struggle with transforming the existing processes over to new streamlined processes. How does an organization transform from zero automated deploys to 100s/1000s a month? Most importantly, how do you show the value of Spinnaker to your business as a whole?

    November’s Spinnaker Summit 2019 in San Diego allows us an opportunity to hear and learn from DevOps engineers, developers and practitioners using Spinnaker open source software. Joel Vasallo, manager of Cloud DevOps at Redbox, shares a preview of his talk where he shares his experience about using Spinnaker open source to increase the number of deploys into production. Joel found Spinnaker so useful because it is very extensible, covers AWS well, strong support for Kubernetes, different deployment types and has the automation to create the software delivery platform that fits the need.

    Joel’s talk, Transforming Software Delivery using Spinnaker, is on November 17, 12:30 pm PT at Spinnaker Summit 2019 in San Diego.

    As usual, the streaming audio is immediately below, followed by the transcript of our conversation.

    Transcript

    Mitch Ashley: Hi, everyone. This is Mitch Ashley with staging-devopsy.kinsta.cloud, and you’re listening to another DevOps Chat podcast. Today, I’m joined by Joel Vasallo who is manager of Cloud DevOps at Redbox. We know Redbox is the movies, the kiosk, the online service—great.

    Our topic today is going from zero to thousands of deploys a month. Boy, if you’re not doing a thousand a month already, that sounds like a great place to be. This is actually a preview of a talk Joel is doing at the Spinnaker Summit on “Transforming Software Delivery Using Spinnaker.” It’s 12:30 p.m. Sunday, November 17th in San Diego. Joe, welcome to DevOps Chat.

    Joel Vasallo: Hey, thanks so much for having me, Mitch. I’m glad to be here.

    Ashley: Glad to have you here. Well, tell us a little bit about yourself, what you do at Redbox and also, hopefully people know Redbox, but just a bit about Redbox.

    Vasallo: Yeah, no, my name is Joel Vasallo, I’m the Manager of Cloud DevOps at Redbox involved with software delivery to the cloud. In my spare time, I’m a GDG organizer, so a Google Developers Group here in Chicago. But about Redbox, I mean, you said a lot of people know what Redbox is. We’re the kiosk that’s usually found in your grocery store. So, we do DVDs and Blu-rays, but we also, a little known fact is, we do have an on demand platform, so video on demand, and that’s kind of our new thing that we’re doing right now. And there’s actually a lot of motivation around building fast deployments, and that’s kind of gonna be my drive for the talk.

    Ashley: Awesome. Now, just a question for you for cloud DevOps—does cloud actually include software that’s being delivered or information being delivered out to those kiosks, too, or is this really for the on demand platform in the cloud that you—

    Vasallo: Yeah, there’s some aspects of it as well. I mean, there’s a lot of APIs in use across the company, so there are some that do kind of traverse the cloud in that regard—yeah.

    Ashley: The only thing better than APIs are common APIs, right?

    Vasallo: [Laughter] Exactly.

    Ashley: Reuse.

    Vasallo: Yeah, especially if you have a solid delivery platform—of course.

    Ashley: Absolutely delivery platforms. Okay, so, let’s jump into it. We’re giving a preview of your talk, so there’ll be a lot more content, of course, in your actual talk. Tell us a little bit about, if you’re already achieved this thousands of deploys, you’ve been at this for a while. Tell us a little bit about how you got started on this journey.

    Vasallo: Yeah. You know, to preface a bit, I mean, I’ve been in the Spinnaker community since, like, late 2015, when it was still kind of an early release out of the open—in the open source community. And at a previous role at Gogo, I used to work at Gogo and we implemented Spinnaker basically from day one. And prior to that, we were using some other open source—Netflix open source tooling as well.

    But in that journey, we were going from a traditional monolithic, you know, to a microservices architecture, and the goal was to essentially build a way to deploy to the cloud as fast as possible, right? And just taking and building on those practices, you know, applying all the things, the same—the 12 factor concepts, immutable infrastructure, routing, all that stuff.

    It kind of transformed into some open source tooling that we’ve written in my history, and there’s also some practices when using Spinnaker to make such a big change in an organization, right? I mean, the whole topic of DevOps in an organization, there’s so many great talks out there, right?

    Ashley: Mm-hmm.

    Vasallo: And that’s definitely one of the—I guess, arguably, the hardest bit is, “Well, what does the DevOps mean for you?” Right? That’s the typical phrase you hear when people say, “Well, we’re DevOps.” “Well, what does that mean?” Right?

    Ashley: Well, now, you’re talking, obviously, at the Spinnaker Summit, and it sounds like you’re pretty committed to the product. This isn’t a podcast about that specifically, but you sound pretty dedicated to it. Why Spinnaker?

    Vasallo: You know, it’s—the core fact is, it kind of extends a little bit past just deployments. It has—it’s really, the biggest thing is, it’s so extensible. Right out of the box, you’re just coming out swinging. The fact that Netflix, Google, Microsoft, a lot of big names—Amazon as well, how could I forget?—are contributing and helping build this product up, you have some pretty big names out there, right? And why not lean on the shoulders of giants. They’ve seen, I would argue, probably one of the bigger challenges in terms of software delivery, especially at an in demand space such as, you know, for example, Netflix, right?

    Ashley: Mm-hmm.

    Vasallo: And building on that success, you know, you basically get those batteries included, right? You get some battle tested software, and I think that’s one of the biggest things.

    But in addition to that, I mean, it’s extensible. It covers Amazon very well, it covers Kubernetes, it has support for various deployment types and various customizations as well, some of which are customizations you need to build, and some customizations are actually built into the platform. And the best part, it’s open source, so that’s my favorite thing.

    Ashley: One of the key words I didn’t mention that I know your talk’s gonna be about is automated deploys, right? You’re not doing hundreds of thousands with a lot of manual process. Talk a little bit about the key of automation and you kind of assume a development team is gonna be all about—yeah, let’s automate this and make our lives easier, but maybe it’s not that easy to get from point A to B where you do have that automated. Talk about that part of the journey.

    Vasallo: Yeah, no, I agree with you. And I mean, it’s an unfortunate thing in the landscape right now where a lot of companies are rushing into the DevOps landscape, right, and they just say Jenkins, right—Travis CI, GitLab, whatever. And, you know, you’re essentially flooding tools on the market, right? And it’s very tough.

    But having a lot of tools doesn’t mean you’re DevOps, right? There’s some fundamental business changes that have to occur in order to achieve that delivery that we’re talking about, right?

    Ashley: Mm-hmm, yeah.

    Vasallo: So, with that, it’s kind of taking—how do I call it?—the little infinity loop of DevOps, right? Where that’s great and, you know, the plan, code, build, release—you know, that little cycle, it’s great, but you can’t use that as a guiding principle. There’s not enough detail, right? And then you can have the extreme on the other side, where it’s like the Candyland, I like to call it, where it’s like the Candyland of ITIL process where you’re just essentially passing Go and then passing down a ladder, if conditions. That’s too complex, and that’s hard to explain. So, there has to be something kind of in between.

    And that’s kind of what I’m gonna be focusing on in my talk and really asking the questions of how do you lay out an effective delivery process? What are some steps—what are the steps at your organization to get changes out the door? Because it’s a different answer for every company, right? There’s no one truth out there, unfortunate as it sounds.

    Ashley: Mm-hmm. Well, and everybody’s environment is unique. Maybe some of the problems are the same, but there’s also, you know, what is unique about your organization, whether it’s the business technology or just people, the team you’re working with has strengths in certain areas.

    Vasallo: Yeah, of course. And that’s something that I kind of build on in the talk is about building minimum viable products, right? And the argument is, many minimum viable pipelines, right? Really laying out, end to end, what needs to be done to get software delivery to occur. And the best way to start is really to look at what’s kinda currently happening, right? Lay out, you know, change has to come in, goes through phase two, phase three, phase four, and then ultimately phase five—profit, right?

    Ashley: Mm-hmm.

    Vasallo: Well, what are those phases and how much investment should be done in each spot, right? You can spend your time automating, I don’t know, some aspects of your pipeline, but what’s the return on investment, right? Would it have been better served in, for example, automating quality versus automating security scans, right? What’s gonna be the most immediate value in shaving time and getting you back from, you know, these long, drawn out pipelines to potentially fast and efficient ones, right?

    Ashley: It’s interesting. What was sort of the starting problem that took you down this path? Was the business asking you, “Hey, we need stuff faster, faster, faster”?

    Vasallo: Yeah.

    Ashley: Or was it just your own, “We don’t have resources, so let’s figure out how to automate it”? I mean, what kinda got you going on this?

    Vasallo: Yeah. So, a little bit about myself, I joined Redbox back in 2018 with the mission of building a software delivery pipeline to get software—me and my team to get software into the cloud. And with that, there’s also the transformation aspect of, “Well, what is the cloud? How do we integrate with the cloud?” And a lot of questions kinda came up.

    Fortunately, we took a big investment in doing a cloud native approach, you know, from day one. Which means no lift and shift, which is awesome. It’s arguably the hardest path, [Laughter] as we probably all know, but this means essentially from the ground up, building a software delivery platform, right, that’s integrated from the, all the way down to the development layer and all the way up to the automation layer. So, we’ve been very fortunate to work on something like that.

    That’s kind of where we started, and that was the motivation of building a truly cloud native pipeline, so, yeah, working on that stuff.

    Ashley: Now, you mentioned you’ve developed some tools, too, and with Spinnaker being open source, have you contributed a lot back to the community yourself, or are you more of a consumer of it? You know, everybody plays a different role, and not everybody’s, you know, having an open source project.

    Vasallo: I’ll selfishly say that I’m more of a consumer of it.

    Ashley: That’s okay.

    Vasallo: I would love to contribute more to it. I’ve done a few contributions in my past, but in recent times, I haven’t.

    With that being said, though, my team is empowered to make change and drive changes, so we do have one or two changes in flight, and one of them actually got merged this past week. So, we’re looking to help grow and not just be consumers. We’re hoping to also build and solve the same problems we’re seeing, maybe, for other people as well.

    Ashley: Well, good. That’s great if you get a chance to do that. Keep in mind, you know, exercising it, especially at the volume you’re talking about, that is super valuable, too. That’s what helps improve software, whether it’s open source or paid for. So, that’s another—

    Vasallo: Well, yeah, and I mean, it’s kinda like what you’re saying. If you’re not paying for the software, maybe you can pay for it within support in the sense that Spinnaker’s community is an open community. There’s a Slack channel out there and you can join—anyone’s more than welcome to join, ask questions, and get answers from potentially some high up people. And Netflix, even people in the community in general, right?

    We’re all on the same journey, and I mean, that’s arguably what kinda kept me involved in the Spinnaker community is just seeing such an open and collaborative group of individuals just willing to help each other—special interest groups every week or so meeting and just saying, “Hey, what’s the direction of the project? What else can we do? What are some challenges we’re seeing?” It’s actually been pretty fun.

    Ashley: What are some of the objections you typically see or maybe that you even ran into about automating and doing deploys at that volume?

    Vasallo: Yeah. So, it’s not like we’re doing 1,000 more changes a month, right? It’s arguably, running fast can initially be seen as a reckless approach, right? You know, if you said, “Hey, we’re averaging 10 deployments a month”—maybe that’s a standard organization—“to 1,000,” you’re gonna definitely open up some eyeballs, right? People will be like, “Wait, wait, wait. What do we need that for?”

    Ashley: “What about stability? What about up time? What about”—

    Vasallo: Exactly, exactly. But the argument is, it could be said that smaller changes typically result in less down time, right? If you can break down a purchase path change in some company, right, and maybe saying, “Hey, we’re gonna make a UI change and then we’re gonna do a back end change that’s backwards compatible and we’re gonna do, then, an API tier change to support that change such that everything hooks together,” you then broke up one large change of three major components into three separate small, individual pieces that can be rolled out safely, right? And tested and embedded as well. That’s the most important part. Because you’re only as good as your automation and testing as well.

    Ashley: Now, did you do much shift left in terms of doing more security testing or vulnerability testing earlier in the Dev cycle, too, as part of this?

    Vasallo: Well, so, that’s the beauty of it. I mean, essentially, the pipeline doesn’t, at least for us, doesn’t start at Jenkins or the build time. It starts all the way at the commit. So, we have a lot of tooling, at least at Redbox, that’s built to empower developers. The thing to also note about those 1,000 deployments—those deployments aren’t, you know, some Ops team clicking an OK button. They are developers promoting their own code, and they are empowered to write and deploy code to production when they feel it’s appropriate to release, right?

    But in doing that, we don’t just give them the keys to the city, right? We build a nice, safe, and reliable pipeline—essentially, a paved road, if you will. That’s the kind of term I’ve been hearing it described where, you know, you can ride the paved road and have fun. You can get off road and it’ll be bumpy, but the majority of the time, the paved road will get you to production as fast as possible.

    Ashley: Right. Yeah, that paved road concept comes from the Netflix paved road talk that they gave. I think that was OSCON or something.

    Vasallo: Yeah, there are definitely a lot of great things from the engineers there, a lot of inspirational things kinda came out of it. I mean, again, it’s building on the shoulders of these giants, right? They’ve seen, arguable, large percents of Internet traffic, right? And at that level of success, you know that some of those concepts definitely can be applied elsewhere, right?

    Ashley: Mm-hmm. Well, you know, one of the things that’s maybe a little bit misleading in saying automated deploys is, it isn’t just the deployment process that’s automated, is it? Talk about what all the parts that are automated to be able to deploy at that volume.

    Vasallo: Yeah, no—definitely. I mean, it’s essentially part of the—part of the talk is also gonna, the little secret of it is, it’s not like you can roll in Spinnaker and have an automated deploy and you’ll be good, right? There’s a lot of aspects of a pipeline, and as I was alluding to from a commit layer, right, whether you do your static analysis for security early on, whether you do your unit tests at build time, whether you have an automated way of storing in archiving artifacts, right? You need to know what is making it out to your environment stage prod.

    In addition to that, you need to build that infrastructure automatically, right? So, we have things—you know, you can use things like Terraform and CloudFormation as well, right?

    Ashley: Mm-hmm.

    Vasallo: But your load balancers, your security groups—again, we’re talking within the context of AWS—how do you reliably and repeatedly build your infrastructure in such a way that it can support that rapid deployment, right? You can’t just have an automated pipeline where it’s like, “Okay, hold the lights, let’s go call up somebody and say, ‘Hey, we need a load balancer. Can we get that done?’ ‘Sure, give me five minutes.’” Even five minutes is too slow, right? We need to have an automated way of creating and also getting fast feedback to our developers when these changes work and when they don’t work, right, so they can help triage and figure out what’s going on.

    Ashley: It’s almost like I’ve said before—DevOps is an overnight success 35 years in the making. [Laughter]

    Vasallo: [Laughter] That’s a good one, that’s really good.

    Ashley: Where you kinda get there one step at a time, but it is, and I think that’ll be something interesting to anticipate about your talk is really thinking about it holistically and what that entire process—not that you’re gonna go into every bit of detail in that amount of time, but it is about all of that and connecting the toolchain, the work flow, and the process, people process part of it into something that can be automated and happen at high scale.

    Vasallo: Yeah, that’s definitely something to be said there. I mean, a lot of—it’s generally a soft topic, but it’s arguably one of the more important ones is, there has to be a cultural transformation and understanding if you truly wanna be as successful as some of these giants, right?

    Building the concepts of open culture, right, working together, cross functional teams—there’s no magical fix out there in the world to fix a broken process in a broken culture, right? And that’s the toughest thing. And then that’s something, definitely, we touch on as well. And I alluded to it earlier with the concepts of, for us, that’s empowering our developers. So, building compelling tools for our developers to use every day they come to work.

    Ashley: Well, you mentioned keys to the kingdom for the developers. I can’t imagine you did all this without some form of senior leadership backing. How did you get that and where did you get it from to go down this path?

    Vasallo: Yeah. So, it definitely started very grassroots, right? Say the challenge to get to the cloud, right? That’s an easy challenge and I think a lot of—sorry, that’s an easy statement to say, but what does that mean? And for every organization, it’s different, right? For some, lift and shift is the most appropriate solution if you’re not willing to invest in the cultural changes, right?

    But for me, at least in my experience, it’s always been finding the hero in the regards of not necessarily a hero mentality, but finding that one champion who can help drive that, you know? That team who is under a tight deadline and you’re like, “Hey, you know, you’re a relatively new project. You’ve never really seen—you’ve never really gotten deployed. How about we just take an experiment and see if we can automate your app end to end, right?”

    Ashley: Mm-hmm, mm-hmm.

    Vasallo: The good news is, it’s a pretty low risk, it’s a new feature. The timeline, obviously, is always a concern, right, in terms of meeting business deliverables. But in doing that, working with those individuals, you not only can get fast feedback as to the proposal of the pipeline, you can also build a relationship that essentially is like DevOps and Devs working together to kind of create this delivery solution, if you will.

    Ashley: Hmm. Well, I feel like we just scratched the surface, which is probably a good thing, because it’s the preview of your talk. I appreciate you coming on the podcast today, Joel.

    Vasallo: No—no worries. Hey, thanks so much for having me. Again, I hope to see everyone at the Spinnaker Summit. It’s gonna be a great time, so I hope to see everyone there.

    Ashley: It is. Well, definitely, I’d like to thank Joel Vasallo, Manager of Cloud DevOps at Redbox for joining us today. Again, Joel’s talk at Spinnaker Summit, which is at 12:30 p.m. Sunday, November 17th in San Francisco, it’s titled “Transforming Software Delivery using Spinnaker,” so be sure and check that out if you’re making it. Now you have a reason to go to San Diego in the fall coming on the winter.

    Vasallo: Other than the great tacos, yeah. And everyone, join us on the Spinnaker Slack, we’re happy to help out and answer any questions as well. We’re more than a welcoming community, so hope to see more folks there.

    Ashley: That’s a great point. It’s a very vibrant community. So, we wish you all the best in your talk, and everyone, you’ve listened to another DevOps Chat podcast. I’d like to thank you, our listeners, for joining us today. This is Mitch Ashley with staging-devopsy.kinsta.cloud and you’ve listened to another DevOps Chat. Be careful out there.

    — Mitchell Ashley

  • Cloud Native Tracing and Observability: Why You Care

    Cloud Native Tracing and Observability: Why You Care

    Is observability the new monitoring? Or is observability, and tracing, fundamentally different? Like any IT industry trend, it can be difficult to discern as many jump on the trend bandwagon, appropriately or not. The Splunk .conf 2019 event presented Splunk with the opportunity to bring clarity to the terms and the reasoning behind the acquisitions of SignalFx and Omnition.

    Monitoring is mostly about taking events, telemetry data and established data points and thresholds into monitoring, alerting and problem management processes and tools. Its tried and true, and has progressed with the evolution of typical monolith, distributed apps and systems. Enter cloud-native applications, composed containers, microservices and service meshes. With great flexibility comes complexity, and cloud native isn’t immune to such an adage.

    What would have been a monolithic application is now shattered into hundreds or even thousands of smaller pieces of application and app technology software. The complexity of externally monitoring cloud native quickly surpasses the capabilities of most monitoring approaches.

    Observability, as recently acquired SignalFx CTO Arijit Mukherji shared this week, is built upon three pillars: metrics, tracing and logs. Metrics show when you have a problem, tracing points you to the problem and logs help find the root cause—a reasonable way to define and segment observability. It also, not surprisingly, fits well with the logic of Splunk plus SignalFx plus Omnition.

    To move beyond monitoring requires instrumentation, built into the software and APIs as part of the software creation process, a DevOps process. SignalFx in part brought Splunk auto-instrumentation, which during run-time, identifies frameworks and libraries in use within applications and can capture tracing instrumentation. Omnition brings even deeper tracing chops to perform tracing across large service meshes of microservices. Add Splunk’s capability to correlate data across the business, applications and operations data and you complete the picture with the ultimate goal of making observability easy for developers.

    The above might explain why developers, ops and DevOps care about observability and tracing, but should the business care? SignalFx CEO Karthik Rau connected the dots nicely during a conversation this week. Digital transformation strategies require speed and agility, but also demand more risk-taking. Confidence in risk-taking comes when accompanied by the ability to respond rapidly to changes and failures. A software deployed 10 minutes ago may need to be backed out or corrected rapidly when the users’ experience goes negative. That requires a rapid determination of what the problems is, and the ability to take immediate corrective action, including automated action.

    Cloud native and DevOps not only enable disruptive, digital transformation strategies but must be accompanied by rapid and automated responses when negative business and customer impacting conditions occur. Customers don’t care when CPUs are taxed, but they do care when a mobile app’s responses fail or are slow. The move to cloud native is served well when backed up by such capabilities to respond in near real-time when problems occur.

    — Mitchell Ashley

  • DevOps Chat: CI/CD Velocity for Large Monolithic Services with Pinterest

    DevOps Chat: CI/CD Velocity for Large Monolithic Services with Pinterest

    Spinnaker Summit 2019 Preview: Software Engineer Rainie Li played an essential role in implementing Spinnaker as part of Pinterest’s CI/CD pipeline. The results moved Pinterest from two scheduled deployments per day to continuous deployments, greater than 15 during business hours.

    This episode of DevOps Chats features a preview of Rainie’s talk, “How we introduced CI/CD for Pinterest’s largest monolith services (API and Web) to improve developer velocity, quality & reliability (Pinterest).” Topics including how Spinnaker was selected, important metrics, Pinterest’s future CI/CD platform Hermez, Canary analysis and lessons learned from the journey are shared.

    Rainie’s talk is on Sunday, November 17th, 1:30 pm PT, at Spinnaker Summit 2019 in San Diego. Joining Rainie on the talk is Jasmine Qin, software engineer with Pinterest.

    As usual, the streaming audio is immediately below, followed by the transcript of our conversation.

    Transcript

    Mitch Ashley: Hi, everyone. This is Mitch Ashley with staging-devopsy.kinsta.cloud and you’re listening to another DevOps Chat podcast. Today, I’m joined by Rainie Li, who’s a Software Engineer at Pinterest.

    Now, Rainie’s gonna be talking at the Spinnaker Summit 2019 coming up in San Diego, November 15th through the 19th, and her topic she’s presenting on is how Pinterest implemented CI/CD pipeline for large, monolithic, or monolith services, APIs and web, focused around helping them improve developer velocity, quality, or liability, et cetera. So, some real hands on experience that she’s planning on sharing. And she’ll also have a co-presenter, Jasmine Qin, and she’s gonna be joining her, too, at that presentation.

    Well, Rainie, welcome to DevOps Chat.

    Rainie Li: Thank you, Mitch.

    Ashley: Awesome to have you on. Thanks for joining us. Would you just introduce yourself, tell folks a little bit about you, the kind of development work that you do at Pinterest and also maybe a little bit of what you’ve done in your past to become a software engineer?

    Li: Yeah, sure. So, hi, everyone, I’m Rainie Li, I’m working at Pinterest infraorganization and our team is focused on doing continuous delivery platform for all the internal engineers. Before I joined Pinterest, I worked at both Microsoft and Amazon, also as a software engineer, and I found I have a very strong interest in infrasite, and it turns out, continuous delivery is a very important feature for software engineers. So, that’s why I ended up here.

    Ashley: Very interesting. Well, you’ve worked at some very large online services. I know that Pinterest has, what, about 250 million users, and of course, Microsoft and AWS are very large organizations, so you’re kinda used to working on some pretty big infrastructure projects, it sounds like.

    Li: Yeah, I am. [Laughter]

    Ashley: Good, good! Well, great experience. So, tell us a little bit about your talk. I guess maybe start with, you know, you came to Pinterest, maybe about where they were in this process of implementing continuous delivery and where you kinda picked up in the process. Were you at the beginning or they already started?

    Li: Yeah. I kinda joined the one that already started, but I still drive the whole production readiness for Spinnaker at Pinterest. I also drive the design review for Spinnaker as well. So, I kinda joined the one that already picked Spinnaker for Pinterest, but I’m the main person to continue to deliver this project.

    Ashley: Oh, so you kind of picked the project already in progress and then take it forward from there as one of the leads on it?

    Li: Yeah.

    Ashley: Very nice. So, when you said—so, you do a review of Spinnaker. Is that your internal usage and changes in implementation to it, or do you also contribute things back to the open source? Any of that kind of thing?

    Li: Currently, we only reveal how we use it internally and we do some customization for Pinterest only. We haven’t contributed to the upstream yet, but I think long term plan, we would like to contribute to upstream.

    Ashley: Yeah. Well, you know, just using it at the volume that you do, I’m sure, is a way of, a form of contributing back, too. It’s gotta certainly exercise Spinnaker and a lot of other tools about that, too.

    So, tell us about—so, you picked up when they had already gone through the implementation and you’re kinda furthering it, or were they still in the process of implementing CI/CD continuous delivery?

    Li: I picked up when they were still in the process of implementing a CI/CD platform.

    Ashley: Mm-hmm.

    Li: But they kind of already decided to use Spinnaker, but we still have to do Pinterest specific customization and how to make it production ready at Pinterest, that kind of work.

    Ashley: Ah, excellent. So, sounds like you had quite a bit to do with Spinnaker even when  you joined of getting it ready for production. Why don’t you talk about some of the things—I know you’re gonna talk about this in your Spinnaker Summit talk. What are some of the things that you need to do to Spinnaker to get it production ready for that kind of an environment that you’re in?

    Li: So, I think there are several key pieces. The first one is, we have to make sure authentication and authorization are working as expected for Spinnaker, because that’s the most important thing—like, we have to make sure users can go over our OS process before they can—

    Ashley: Your OS, mm-hmm.

    Li: Yeah. The second key piece is monitoring, which includes metrics and alerts. So, basically, we exposed all the Spinnaker service metrics and created a dashboard for those alerts, which we can get paged when there’s some issue that happened if Spinnaker is down or something like that. These are the two most important things for production readiness, I think, for us.

    Ashley: That’s extremely important when you’re automating something like a continuous delivery, continuous integration. It’s automated, but you have to know is it working right?

    Li: Yeah.

    Ashley: Are things happening and are there problems or are you meeting the kind of metrics that you were expecting?

    Li: Yeah.

    Ashley: I know something you mentioned in the description of your talk is what kind of metrics are important to measure. Can you say a little bit about what are some of the metrics that you’ve learned, both while you’re implementing it and now that you’ve had it in production that are really important to watch?

    Li: Yeah, definitely, those metrics can—like, latency and ____ P90 and also, like, those SLA related metrics from all Spinnaker components are very important. I think especially the gate service, which is Spinnaker API Gateway, it’s like the entry point for the rest of Spinnaker components, the metrics for Gateway is extremely important.

    Ashley: That’s kinda the integration component of Spinnaker, right?

    Li: Yeah.

    Ashley: Everything talks to it. It’s sort of that—I don’t know if it’s a central hub, but it’s certainly where all the APIs tie in together.

    Li: Mm-hmm.

    Ashley: I’m sure if you’re having a problem with that, then you’re gonna know [Laughter] there are probably lots of failures happening in the system.

    Li: Yeah.

    Ashley: Do you look at, do you also measure how many builds you’re doing per hour, per day, or how many deliveries into production you’re doing? Is that metrics that you capture through Spinnaker or you do that elsewhere?

    Li: That’s not the metrics that we are capturing in Spinnaker, but we do measure from the pipeline execution history.

    Ashley: Mm-hmm.

    Li: So, currently, we are using Spinnaker to do the continuous delivery for two major services in Pinterest. We finished around 15 deploys per day for these two major systems and each deploy contains roughly 10 commits, and we think this pipeline, it increased our productivity a lot, because previously, we have to do manual deploy.

    Ashley: Mm-hmm.

    Li: Yeah.

    Ashley: Do you also have, then—I know this is outside of Spinnaker, probably, but do you have most or all of your testing automated as well, or are there still some manual steps for testing before it goes into a final deployment?

    Li: You mean testing for the service itself, or testing for Spinnaker service?

    Ashley: Testing more of the software application part of it.

    Li: I see. Yeah, so, currently, we are using a separate stage to run integration test jobs in Jenkins. That’s how we do the testing step. It’s kind of automated already, like, we don’t have to interact with humans to trigger the build or humans to run the test. We don’t need that. It’s just a stage in the Spinnaker pipeline. We run the interpretation test job in Jenkins, yeah.

    Ashley: Mm-hmm, excellent. Okay, good. We have kind of a feel for your workflow, your tool pipelines. It sounds like it’s very automated.

    Li: Mm-hmm.

    Ashley: So, going back to the topic of implementing something at that kind of a scale, both, you’re a very large organization and you’re doing a number of deployments per day. I know there’s some things in Spinnaker—well, you have a platform called Hermes. And is that your own central workflow engine, or is that part of Spinnaker?

    Li: Oh, so, Hermes is our future CI/CD platform. So, we are in the, like, developing stage. Spinnaker will integrate with Hermes and now we will use Spinnaker as a back end workflow engine.

    Ashley: Hmm, okay.

    Li: Yeah.

    Ashley: So, that’s really your—that’s Pinterest’s own central workflow engine and you’re integrating Spinnaker into it as part of that feature?

    Li: Yeah, yeah.

    Ashley: Pipeline—tool pipeline, if you will. Okay.

    Li: Yeah, it’s like a CI/CD platform for Pinterest, and we are going to open source Hermes as well in the future.

    Ashley: Oh! Well, that’s good news. I’ll look forward to that.

    Li: Yeah. [Laughter]

    Ashley: Do you know kinda time frame when that’s gonna be happening, or is that a little farther in the future, don’t know yet?

    Li: I’m not sure the exact timeline, but somewhere early next year or at the end of this year.

    Ashley: Okay. Well, that’s not too far away. That’s great. We’ll look forward and thank you for contributing that as open source. So, talk some more about what are some of the lessons that you learned about going from where you joined Pinterest and where things were with Spinnaker’s implementation, maybe some things you had to do, some of the things that you learned? Here’s a good way to do it, here’s a mistake I made and I fixed it by taking a different approach—what are some of those lessons learned?

    Li: Hmm, I think the lessons for me is, because we deploy Spinnaker services to Kubernetes platform and I never used a Kubernetes platform before I joined Pinterest, so that’s the biggest lesson I learned here, like, how to set up all the Spinnaker components in Kubernetes platform, yeah.

    Ashley: Uh huh, good. Are there—so, learning how to do that was important. Were there any specific lessons that you learned about doing that, or is it just really just kinda learning the how of doing it?

    Li: A specific lesson I think is some special situations where the Kubernetes platform decides to rotate cluster and do some platform testing. On the Spinnaker service side, we have to handle this kind of scenario well instead of bringing down the sides or have some happy customer user experience. That’s a lesson I learned from there.

    Ashley: Okay. I think you also mentioned in your show notes or your description in your talk that you did some canary analysis as part of that as well?

    Li: Mm-hmm.

    Ashley: I assume that was a new thing for you, or have you done that before?

    Li: Yeah, that’s also a new thing, but we have a separate team which is the configuration team which, they are mainly implementing this feature. So, for Pinterest, we support OpenTSDB as a metric, but I think Spinnaker canary analysis can only support Prometheus or Datadog, those metrics. So, I think the ____ team integrated this OpenTSDB into Spinnaker canary analysis component so that we can provide Pinterest ____ on Spinnaker UI.

    Ashley: Ah, interesting.

    Li: Yeah.

    Ashley: Okay, great. What are some other things that you’re planning on talking about at your talk during Spinnaker Summit?

    Li: I think how we use Spinnaker and how we customize Spinnaker at Pinterest and the lessons we have learned when we adopted Spinnaker here. I think these are the main topics I’m going to talk about, yeah.

    Ashley: Mm-hmm. Great. Well, I know you weren’t at Pinterest when they decided to use Spinnaker for this project or for part of the toolset. But based on what you’ve seen and what you’ve learned, do you think Spinnaker accomplished the goals that they had for why they chose it?

    Li: Definitely. Yeah, we like Spinnaker a lot. [Laughter]

    Ashley: What were some of those goals, do you recall? Were there certain, we wanted to get to a certain number of deployments per day, or were there other things that they were trying to achieve that were looking to see how you’re doing?

    Li: I see. So, we don’t have a specific number for the goals, but Spinnaker definitely helped us to use pipeline for deployment. Without Spinnaker, we have to go to each stage, do manual collect deployment one by one, which is not very efficient. And currently, we can use Spinnaker to manage deployments in a consistent and a repeatable way. We really like pipeline, which can provide a sequence of deployment stages. We don’t need too much manual interaction, and—yeah, that’s, I think, the goal we adopted Spinnaker is, we need the pipeline feature and I think it works well for us. Also, the canary analysis report is very useful for us as well.

    Ashley: Yeah, I know a lot of people are very excited about using that.

    Li: Yeah.

    Ashley: And you mentioned that you integrated OpenTSDB as part of your observ—I can’t say it [Laughter]—observability stack, there we go, got that out. Now, I think you also have a deployment system there that Pinterest uses to deploy into AWS, correct?

    Li: Yeah. We have a deployment system which is also open source called Teletron, which we used for four years to deploy to AWS VM. But this deployment system does not provide a good pipeline, so that’s why we integrated with Spinnaker to have the nice pipeline feature, yeah.

    Ashley: Excellent. Well, you’ve sure done a lot of work. What’s kinda next? What’s the next set of projects that you’re thinking about or you’re currently working on now that you’re at this point with deploying Spinnaker into your CI/CD pipeline?

    Li: I think the next project is, our team will implement Hermes at the Pinterest main CI/CD platform. Spinnaker was working at the back end workflow engine, and we were also working on migration, like most of the services from AWS and VM to Kubernetes. Yeah, that’s two major projects that we are going to be working on.

    Ashley: Ah. Well, certainly, the move to Kubernetes, I’m sure that’s a pretty substantial project, a pretty big one.

    Li: Yeah, yeah. [Laughter]

    Ashley: Well, good. Well, thank you so much for joining us on the podcast today. This is very interesting and I think you’ve got a really compelling and interesting talk. I’m looking forward to hearing more about it and have others to get a chance to attend.

    Li: Yeah, thanks for inviting me, Mitch. I’m glad to introduce Pinterest using Spinnaker at the Summit and I’m very happy to talk about more there.

    Ashley: Absolutely. Well, thanks for contributing with your talk. So, you’ve listened to another DevOps Chat podcast. I’d like to thank my guest today, Rainie Li, who is a Software Engineer at Pinterest. Again, her talk during Spinnaker Summit is on Sunday, November 16th at 1:30 p.m. This is according to the agenda currently on the website, you can look there for updates for the Spinnaker Summit site, and her talk, again, is on how we introduced CI/CD for Pinterest’s large monolith services, API and web, to improve developer velocity, quality, and reliability. And as you’ve just heard, there’s a lot of information that Rainie’s gonna be sharing with you.

    So, thanks for joining us, Rainie, and thank you, of course, listeners, for joining us here today. This is Mitch Ashley with staging-devopsy.kinsta.cloud, and you’ve listened to another DevOps Chat podcast. Be careful out there.

    — Mitchell Ashley

  • DevOps Chat: Monitoring Spinnaker on GKE with Miles Matthias

    DevOps Chat: Monitoring Spinnaker on GKE with Miles Matthias

    On a project to move one of Google’s Fortune 500 customers to Google Cloud, Google Kubernetes Engine (GKE) and Spinnaker open source, our DevOps Chat guest ran into the message “hang tight.” No scripts or documentation. Not to be delayed, Miles Matthias, Google Cloud Consultant with Container Heroes, filled the gap and contributed his work back to the Spinnaker community.

    This episode of DevOps Chats features a preview of Mile’s talk, “Monitoring Spinnaker with Prometheus Operator on GKE” that he will be giving on Saturday, November 16th, 3:45 pm PT, at Spinnaker Summit 2019. Miles also talks quite a bit about Canary testing in this episode.

    As usual, the streaming audio is immediately below, followed by the transcript of our conversation.

    Transcript

    Mitch Ashley: Hi, everyone. This is Mitch Ashley with staging-devopsy.kinsta.cloud, and you’re listening to another DevOps Chat podcast. Today, I’m joined by Miles Matthias, he’s a Google Cloud consultant and he’s with Container Heroes—great company name. Our topic today is actually a preview of his talk at the Spinnaker Summit 2019. His topic is gonna be Monitoring Spinnaker with Prometheus Operator in a GKE Environment, Google Kubernetes Engine. His talk is on Saturday, November 16th at 3:45. Miles, welcome to DevOps Chat.

    Matthias: Hey, Mitch. Great to be here. I’m really excited about the Spinnaker Summit.

    Ashley: Well, tell us a little bit about you, what you do and a little bit about Container Heroes.

    Matthias: Yeah, sure. So, like you said, at Crowd Consultant, I’ve done a whole bunch of software development and infrastructure setup and architecture design throughout my career and now I’m helping other companies set up their cloud infrastructure, make migrations, modernize their application development and all sorts of things like that. And I’m a partner at ContainerHeroes.com, and we’re a group of kind of people like me that have been in previous startups, CTOs, who’ve done the whole PC route, we’ve built a bunch of custom applications for clients throughout out career and now we enjoy consulting people that are using the cloud.

    For the past year or so, especially I guess the last nine months or so, one of our big clients, we’ve been helping Google with one of their new customers that was complete on prem to cloud, to Kubernetes containerization migration and they utilized Spinnaker. So, I kinda was brought in as the Spinnaker expert to help them get it up and running, to help customize what they needed in their environment and contributed as much as they could back to the open source, including some stuff that I built for using Prometheus Operator, which we can touch on a bit. And so much so that then the team invited me to give a talk at the conference, so I’m excited about it.

    Ashley: Sounds like a very—very relevant talk, especially sharing your experience from working with this Google customer.

    Matthias: Yeah.

    Ashley: So, tell us a little bit about when you came to this project and first started working with the customer. Had you worked with Spinnaker before, was that a new technology, or were any of these new technologies, or you’ve kinda done all this before?

    Matthias: So, I’ve done a bunch of CI/CD before. Spinnaker is still in its early days, so I have very limited experience with Spinnaker as an application. It’s still—yeah, it’s still very nascent on the development scene. The entire kinda concept of continuous delivery is still very new in the industry, even though it’s been around for a little while.

    And so, CI has been very well developed, everybody knows Jenkins, everybody has ____, a million CI tools, content of continuous delivery and having more thought out, sophisticated delivery strategies that can help you do more automated continuous delivery like canary analysis and different deployment types and automated rollbacks and all these sorts of things are kind of brand new and Spinnaker obviously really helps with a lot of that stuff.

    Yeah, I had some experience, but was really excited to jump in and kinda get into something that’s really taking off in the industry.

    Ashley: Tell us a little bit about the app that you were challenged to work on. I understand this was already, had been developed with Jenkins and Spinnaker before and then you were moving it to GKE, or what was it?

    Matthias: No. So, the client had 50 different job applications, a bunch of different things, and it was all running on prem. So, part of the big effort was to obviously move them onto Google Cloud, but also to help them develop a CI/CD process that allowed them to make a commit in their repo, have artifacts be built in CI, have tests run and then have those passed to Spinnaker and CD and then deployed to GCP.

    They had no experience with Spinnaker and didn’t have any kind of CD solution setup, really, especially for the cloud. Like I said, they were completely new to the cloud, so this is all new stuff for them, so I helped them set it up, install it, configure the options all we wanted, manage it and then introduce more and more advanced features as the project got more and more migrated over to GCP.

    Ashley: Mm-hmm. So, it sounds like maybe a more traditional Java application environment maybe is not even continuous integration yet. Were they down the path with that?

    Matthias: Yeah, they had Jenkins on prem.

    Ashley: They did?

    Matthias: And so we—you know, yeah. So, the architecture was moving everything to Kubernetes and so, you know, one of the GCP projects has GKE cluster just for CI/CD. So, running Jenkins and running Spinnaker on it. So, they had Jenkins on prem, and moved that over to the cloud, hosted it on Kubernetes, and then we installed Spinnaker right alongside it on the same GKE cluster, right?

    Ashley: Mm-hmm.

    Matthias: And then used those tools as the basis of their CI/CD pipelines to deploy to other clusters within the cloud for them.

    Ashley: How about the size of the application? How would you quantify how kind of big or extensive it was?

    Matthias: Like I said, they had a—this was a true microservices organization. So, you know, they have 30 to 50 different applications as microservices, each individually being deployed. As far as traffic, I mean, so the applications, like I said, are microservices, so they’re pretty small in and of themselves individually.

    As far as traffic and resource usage, it’s a very, very large, large client. Like, pushing the boundaries of some of the largest things we’ve seen.

    Ashley: Hmm. Okay.

    Matthias: So, that was really exciting—really cool.

    Ashley: This was a Fortune 500 company, I understand.

    Matthias: Yeah.

    Ashley: We’re not talking about the specific company, but—

    Matthias: Sure, yeah, yeah, yeah.

    Ashley: Well, good. Well, what are the sort of things that you’re planning on talking about, then? I mean, I know this is about the monitoring aspect of it, so you know, are you touching on Prometheus Operator and probably how to configure it or how you decide it and what kinds of groupings to create, what kind of rules and alerts and those kind of things? Are you gonna be talking more fundamental architecture? What are your thoughts?

    Matthias: Kind of a bit on everything, because this kind of—this is kind of some required setup in order to do monitoring and in order to do canary analysis, even, if you’re gonna use Prometheus for canary analysis metrics. And it’s also just fundamental to how you wanna have this setup if you’re running Spinnaker on Kubernetes in general.

    So, a little background on Spinnaker as a project, which I’m sure other people that are familiar with Spinnaker definitely know, you probably know yourself. Spinnaker was originally created by Netflix. Netflix is still very much a VM shop, right? So, they don’t use Kubernetes, and everything runs on individual instances. And so, Spinnaker can be deployed to run on Kubernetes and that support is there, it has support to deploy applications to other Kubernetes clusters.

    But there are still a few things like monitoring, like canary analysis where, in the Kubernetes world, we do things a little differently, and Prometheus is one of those examples, right? In the VM world, you spin up Spinnaker, you spin up one Prometheus instance—and Prometheus obviously is a metric store to collect metrics and then usually is paired with Grafana and ____ and things like that in order to—all CNCF open source projects in order to give you graphs on these metrics.

    And so, there was previously, in the VM world of Spinnaker, running it as Netflix does, a microservice of Spinnaker that does monitoring, that connects to all of the different components of Spinnaker, listens for their metrics and then reports it to Prometheus. Cool. Like, it looks great. They even have some dashboards.

    So, Grafana dashboards, if you’ve ever worked with those you can just upload some JSON files for your dashboard and then you click there and then you’re like, “Hey, look at these dashboards!” You can get real time dashboards based on the Prometheus metrics of Spinnaker. And Spinnaker, all the different components that it’s made up of and the different metric based on the things that it’s doing so that you can see in there when it’s processing your deployment, here are some of the metrics that are coming off, right?

    Ashley: Mm-hmm.

    Matthias: So, all of that support was there. When you want to run Spinnaker on Kubernetes, though, that support was a little…not so much. [Laughter]

    Ashley: Mm-hmm.

    Matthias: And that’s the trouble that we kinda ran into. So, the thing that was there was the concept of, “Hey, we have this monitoring component in Spinnaker,” and if you enable monitoring, that monitoring component will be enabled as a sidecar container to every single pod, every single microservice that Spinnaker is composed of, and it knows how to talk to that microservice and get the metrics from it and then report it somewhere.

    Cool—good. Okay, I can get that. What about the monitoring service then having those metrics that are just collected from all the different components and putting it somewhere? Ideally Prometheus, you can also do Stackdriver or, I believe there might be Datadog support now, but mainly, the two that we see used a lot, Stackdriver and Prometheus.

    Ashley: Mm-hmm.

    Matthias: However, in the Kubernetes world again, especially when you have a bunch of different Kubernetes clusters, one of the tools that you end up using is called Prometheus Operator. And that is essentially a Kubernetes operator that, when you apply to your cluster, will automatically install for you Prometheus Grafana, an alert manager as pods, deployments running on your cluster. It’s a way to provision your infrastructure so that when you deploy a bunch of applications in these clusters, you already have some monitoring infrastructure set up, right? I can deploy and I can basically go to my organization and say, “Give me a new cluster and then I can deploy my application and start admitting Prometheus metrics and there’s already a Prometheus instance on the cluster, so like, I’m good to go.”

    In order to get the Prometheus, the version of Prometheus that the Prometheus operator creates to then go and fetch the metrics from the Spinnaker monitoring component that already existed and already knows how to get those metrics from the Spinnaker different components, you have to kinda connect those two ends, right? And in the Prometheus Operator, the Kubernetes world, basically, you just apply Kubernetes manifest, the CRD called the Service Monitor, and that basically tells your instance of Prometheus on the cluster, “Hey, go fetch the metrics from these places,” right? And so, it’s pretty easy, but there was no setup way in Spinnaker. So, you would read the documentation in Spinnaker and they would have very detailed, very nice—like, if you’re running Spinnaker on VMs and you want to do monitoring and you wanna use Prometheus, just use this fancy and nice, easy setup script, right?

    Ashley: Mm-hmm.

    Matthias: Super nice. Great. And then literally in the documentation, it said, “If you’re running on Kubernetes. Hang tight, support is coming.” [Laughter]

    Ashley: [Laughter] The “hang tight” documentation feature. Wonderful!

    Matthias: Hang tight—yeah, right. So, none of this is extremely complicated. There are a lot of different components and you have to be careful about how you set them up, right? But the very nice setup script that was available for just running plain old Prometheus on VMs was very nice and like I said, it installed a bunch of Grafana dashboards for you and did a bunch of nice stuff for you. There was no such thing for Kubernetes, right?

    Ashley: Mm-hmm, mm-hmm. Kinda starting from scratch. Not completely, but at least to take that next step.

    Matthias: Exactly, exactly, exactly. And so, what I did and what I contributed to the project was creating a setup script for this exactly, saying, “Hey, you’ve got a cluster, you’ve got Spinnaker already installed on there, you’ve got Prometheus Operator installed on there and you’ve got a Prometheus Operator installed on probably a bunch of other clusters, too. Here’s a setup script that connects Prometheus instance to Spinnaker so that Prometheus can go and fetch those metrics about how Spinnaker is running and how Spinnaker is handling your deployments.”

    Ashley: Mm-hmm.

    Matthias: And installs a bunch of these Grafana dashboards that the open source project has already created for us to be able to say, “Here’s a Grafana dashboard about each microservice—Clouddriver, Gate, Orca, all the others, they each have their own dashboard in Grafana that you can go look at and they’re preconfigured and it’s really nice, it’s really neat.” Again, those just had to be kind of converted into the way you apply these in the Prometheus Operator world which, again, is applying your Kubernetes manifest to say, “Hey, it’s actually just a config map with a certain label on it.” “Hey, here’s a dashboard,” and when you apply that, the Prometheus Operator knows how to hook into the Kubernetes master and goes, “Hey, there’s a new config map that is a Grafana dashboard. Let me go fetch it and add it to Grafana for you,” right? That’s the whole purpose of Prometheus Operator.

    So, that’s kind of the thing that I set up was saying, was giving you—so now, and then I PR it to the documentation. So, now, when you go look at the documentation, it says, “Hey, if you’re running Spinnaker and Prometheus Operator on Kubernetes, there’s also a fancy setup script for you, too!” [Laughter]

    Ashley: There you go. Now people will have to start at the same place you did.

    Matthias: Yeah, really.

    Ashley: So, are your planning, on your talk, actually walking through this, the process of the script you built and how to set all this up, get the dashboards to come up, et cetera, or you’ve kinda taken a little different approach?

    Matthias: I don’t want to anger the demo gods, so I’m not sure if I’ll actually be, you know, typing it, running the actual scripts, but I can certainly have something set up that’s like, “Hey, this is the end result. Let’s step through what the script is actually doing to help you understand what components need to be matched up where, and then here’s the end result,” right?

    The other kinda couple things that I’ll probably touch on in the talk are—which is why it’s important. So, all of that is monitoring Spinnaker as an application. So, like, your DevOps or your Release Engineering team, people that are involved in these things who are gonna monitor Spinnaker as an application and are gonna see, “Oh, Orca is overwhelmed and it has a bunch of tasks in its queue that it can’t pop off, so let’s add some more replicas of Orca” or something like that.

    That’s what this script helps you set up is this dashboard to see the performance, to see those metrics, right?

    Ashley: Mm-hmm, mm-hmm.

    Matthias: And other reasons that it’s important are, and the other things I’ll touch on—well, first of all, the one thing I’ll probably touch on is that there’s kind of an effort currently in the works and I hope by the time of the talk, maybe there’s a little more effort built towards it, but there isn’t really a whole lot of guidance in the Spinnaker community about, “Okay, here are the metrics we emit. Here’s what they mean, and if they go outside normal ranges, here’s what you should do to remedy that situation.”

    Ashley: Mm-hmm.

    Matthias: There are no kind of run books, as you might call them, for how to respond to a Spinnaker instance that is under a really extreme load or kinda gone configure sideways or something like that.

    Ashley: Interesting.

    Matthias: Rob Zyner kinda leads the Netflix effort. He’s done a great blog post about some of the metrics and things like that and what each of them mean. But that’s kind of the only resource out there. So, there is kind of this effort in the community to say, “Hey, let’s put together some actual run book and match some of these dashboards that we have in Grafana and some of these metrics and be able to tell people, ‘Hey, if you run into this situation, here are the step 1, 2, 3 to check, and if these are the case, then here are the three options for alleviate that problem or how to address the problem,” right?

    Ashley: Mm-hmm.

    Matthias: So, hopefully, by the talk in November, we might have some progress on that in the community.

    The other thing that I’ll probably talk about on the topic and why the concept of the Prometheus Operator is so important in the Kubernetes world is for canarying.

    Ashley: Say a little bit about what canarying is with your deployment processes.

    Matthias: Yeah, sure. Sure, so, canary testing is—it’s a form of testing that evaluates one of your release candidates. It compares it to—what it does is, it actually runs your candidate, and it runs the version that is running in production, starts a new copy of it, and it runs it side by side. So, then it is—and then it diverts usually a small amount of traffic or just, you know, will have 100 pods behind a service and you’ll add 2 pods so they’ll get a proportional small amount of traffic. And you’ll run it and you’ll see how it performs. You’ll listen to the metrics and you’ll see, “Oh, this is how your candidate that you want to push performed.” And in canarying, in the configurations in Spinnaker, you can say, “These are the metrics I care about when evaluating, when deciding if it performed well and if it performed badly.”

    For instance, let’s say you have a new version of the application that just, like, 500—500 errors out on every request, right? You run a bunch of unit tests, you run a bunch of integration tests and ideally, at that point, at one of those steps or earlier, you would’ve caught that, right? But tests are written by people and sometimes people don’t write them or sometimes some situation that depends on live traffic and live data actually exposes some types of problems, right?

    So, let’s hypothetically pretend you have a new version that comes out and it’s just 500s all over the place. What canarying says, the kind of theory behind it is—let’s actually run it on a small amount of traffic, let’s see what it does. And then canarying would go, “Oh, hey, your release version that you wanna push out? It’s just doing a bunch of 500s, and you told me in your configuration, if you see so many, a spike in 500s, that that’s bad, that we don’t wanna actually then deploy that to all 100 instances running in production,” right?

    That canarying says, “Hey, that’s not a good thing to promote,” right? Like, “I gave it a chance, I ran it on some things, I compared it to what the version that is currently running or our new instance of the version that’s currently running, and it didn’t do well. So, you know what? I’m gonna error out, I’m gonna show you this is how it behaved, you guys go fix it, and then when I come back, I’ll run those things again. And if it performs well, then great, we can release the new version.”

    And that’s kind of one of the main components that you need for continuous delivery. Because you can imagine a world where developers are just committing, committing, committing, committing.

    Ashley: Well, yeah.

    Matthias: They do a bunch of tests, they run some tests, to them it all looks good. And then you just have this other automatic component over here that’s doing canary analysis and that can just test it, it can actually put it into production, give it a little bit of traffic, see how it performs, and if it performs well, then let it go, you know? Promote it, right?

    Ashley: Yeah, that’s extremely useful.

    Matthias: Deploy on a Friday, deploy on a Saturday—who cares, right? It’s actually seeing how it’s running, all the other tests before it’s even gotten to that point have passed, obviously, and if it runs well, then run. And this is what enables organizations like Google, Netflix, and others to deploy thousands of times a day, because it has this automated system to go in and say, “Give it some traffic, see how it runs, and if it runs fine because of the way we configured how we told the system what does it mean to run fine, right? Great—then let it go.”

    Ashley: We’re kinda running up against our time, here. I do have one last question.

    Matthias: Sure.

    Ashley: What kind of folks should come to your talk? Obviously, it seems like developers, people who are working on the CI/CD pipeline, maybe DevOps engineers. Are there other folks you’d recommend, or are those the right folks?

    Matthias: Those are definitely all the right folks. I think if you’re interested in Spinnaker or you’re currently using Spinnaker where you are, want to know about monitoring it as an application and managing it as an application when it experiences a bunch of load and how you want to connect all the dots if you’re running it on Kubernetes and you’re running multiple other Kubernetes clusters that also are using Prometheus Operator, this is a great talk to go listen to.

    Ashley: Again, Miles is a Google Cloud Consultant with ContainerHeroes.com. He’s speaking at the Spinnaker Summit 2019 in San Diego. Now, that conference is the 15th through the 19th of November. His talk, again, is on monitoring Spinnaker with Prometheus Operator on GKE on Saturday the 16th at 3:45. Thank you, everyone, for joining us. You’ve listened to another DevOps Chat. This is Mitch Ashley with staging-devopsy.kinsta.cloud. Be careful out there.

    — Mitchell Ashley

  • DevOps Chat: Service Mesh Tracing, from Envoy, Omnition to Splunk

    DevOps Chat: Service Mesh Tracing, from Envoy, Omnition to Splunk

    As application functions get smaller, containerized, become microservices and combine into service meshes, a new set of challenges crop up. What functions does each service perform? What state constitutes services in trouble? What are the dependencies between services across a complex service mesh? How can we instrument observability and tracing between the service interactions?

    Constance Caramanolis, software engineer at Omnition, joined us on DevOps Chats, recorded just before Splunk’s acquisition announcement. After working at Microsoft, Constance joined Lyft, in part to work with Envoy–an open source project created by Lyft that brings upstream and downstream tracing across a service mesh. Constance recently joined Omnition, while still in stealth, to help “flip tracing on its head.”

    It’s genuinely a fascinating conversation and gives us a window into the challenges and solutions to managing a service mesh at scale. Listeners should also check out our DevOps Chats episode talking about the Splunk’s acquisition of Omnition with Rick Fitz, SVP and GM of Splunk’s IT Markets Group.

    Transcript

    Mitch Ashley: Hi, everybody. This is Mitch Ashley with staging-devopsy.kinsta.cloud and you’re listening to another DevOps chat podcast. Today I’m joined by Constance Caramanolis, software engineer with Omnition. Constance, welcome to DevOps Chat. Great to have you on the podcast.

    Constance Caramanolis: Thank you so much for having me. I’m very excited.

    Ashley: Oh, I’m excited too. I love having, pardon the term, but rock stars like you, developers that are doing some cool work on our podcast. The topic is service mesh at scale, but before we get into that, would you first just introduce yourself, maybe tell us a little bit of your background as a software engineer, what you do currently at Omnition, and if you could tell us a little bit about Omnition?

    Caramanolis: Yeah. So I’ve been in the industry for several years. I first started off at my Microsoft. So I got a lot of good experience in terms of building large components that need to run reliably. So I worked on Windows and Windows Film there and then I moved to Lyft three and a half years ago where I purposely joined to work on Envoy with Matt Kline and Jose Nino and others within Lyft. At Lyft I worked on all aspects of Envoy from configuration management to adding either integral features for open source community or just rolling it out within Lyft.

    The last few months I was on Lyft I worked within our data platforms team just building a variation of a work flow tool. So now I actually just joined Omnition not too long ago. And I’m gonna be focusing a lot on the open telemetry component. We’re using open telemetry through so much valuable data that tracing provides that it’s actually, instead of just looking at individual traces, you’re able to get a whole understanding of what an application does and what impacts _____ have and how it propagates using that tracing data.

    Ashley: Easy to see why you made the transition from your work on Envoy at Lyft and now at Omnition. Well, let’s start out with talking about service mesh, microservices, of course containers, et cetera, et cetera. That’s certainly how we’re developing applications these days. Service mesh brings some other things to it, I think some more complexity because you’re talking about a configurable low latency infrastructure layer that’s really designed to handle network based and process communications at high volumes. So that right there tells you that there’s complexity involved and any time you’re building software, creating applications, how you run that in a service mesh architecture has got to bring with it some challenges. In your experience, what are some of those challenges that you’ve seen?

    Caramanolis: I think the big one, especially how it ties back to Envoy’s main goal in provide observability is that as you break components apart and put different parts of the internet, doing what is still working versus isn’t working and having a consistent definition of that is very challenging. Good example that we love to use when we’re talking about Envoy, especially before people have adapted Envoy is that there is no standard definition of what a failure is. Some people may say that a 503 is not an error. Please don’t do that.

    Ashley: I’ll avoid that on this podcast.

    Caramanolis: Thank you. But if you don’t have–not having standard definition of error across different languages and tools makes it actually hard for people from different teams or just say higher up VP’s or directors to know if things are working or not. Observability just across any service mesh is pretty complicated. Another one thing is definitely coordination in terms of topology especially if we’re talking about coming from monoliths. One of the benefits of a monolith is that you know where all the code is, you can figure out where the ____ ’cause it’s usually within one repo and you can just read it.

    But then the other–that’s definitely great ’cause you need to see where everything is, but when it gets really large it can take a long time to build, deploying it. You have some pieces of code that run maybe once every ten minutes. So why does that need to be running alongside something that’s high priority. It needs to be really efficient. When you’re bringing things apart it’s gonna be harder to actually correlate how things interact. As everyone usually tries to keep documentation up-to-date, sometimes you’ll forget like, “Oh, I’m actually calling an API that just moved to service B and ___ service A.” So keeping that mental model of the service topology is very challenging and is so dynamic.

    Ashley: Especially with microservices architecture. We’re creating so many more of them too.

    Caramanolis: Oh, my goodness. Yeah. Especially it could be say one person’s start of day where like, “Hey, just create a new service and deploy it and let it invoke a very low priority API call.” That’s–you immediately depend on how many people you hired, ____ example blowing up the number of services very quickly. I would say probably third challenge of microservices. One is in actually operation. Is that from an operation point of view how do you remediate things when things go wrong? ‘Cause unfortunately, things will go wrong and usually the goal of any application is stay up as long as possible and to serve customers, whatever your definition of customers are.

    How do you build tools to handle things when things do go wrong and low do you maybe standardize that so that way you don’t have to train a company’s engineers in ten different tools to say, “Handle a certain error case or three different error cases and standardize in that”?

    Ashley: Very interesting. It’s compelling to me why you might pursue a path of using a tool like Envoy. Tell us a little bit about how did you end up using Envoy.

    Caramanolis: I joined Lyft to work on Envoy because I had actually met Matt almost a year before I ended up joining Lyft, but he was talking about how–when we were talking about it it’s just there is a need for Lyft to scale ’cause especially one of the big problems that was experienced is that we were not getting really good observability on our ELB’s. So sometimes we could see things–we know that say I was–for example, I was trying to ____ ride, but when we’d look at our metrics we wouldn’t be able to see things failing.

    So that definitely clearly led to an initiative you saw that our observability along, just our regular request path need to be improved. So Matt and I–and that’s just one example. Matt definitely talks more about the motivation of why Lyft built Envoy, I think in one of his earlier talks. It was definitely–and also just it was an area–I’ve always loved infrastructure and backend. And so being able to build something that the entire company is reliant on, Lyft being able to scale is a road blocker to us being successful.

    So Envoy allowed us to scale because we could add 10 or 100 services instantly and not have to worry about service discovery. So just being a part of something so critical and the opportunity to just learn so much and I’m gonna say I definitely–I was intimidated by the project. So I loved the idea of joining a project where I’m gonna get–mentally get my ass kicked in terms of I have no idea what this is. It’s swimming in the deep end with sharks and I can’t wait to see if I can learn how to swim and see how great–how this shapes me as a developer.

    Ashley: You are a courageous person. Not everybody is willing to expose themselves that way to that much risk. So congratulations. That’s great.

    Caramanolis: Thank you.

    Ashley: Fantastic. More of us should do that. Well, you gave a talk just recently, I think it was December of ’18, KubeCon talking about reducing in the meantime to detection specifically with Envoy for service mesh. Say a little bit about that.

    Caramanolis: Yeah. So a lot of the talks–so I’m gonna say we, as a community. There’s a wonderful community of either contributors and maintainers and also just people who bring blogs around Envoy, a large of the community has talked about either Envoy’s features and how it’s helped them from more of a technological point of view of how it’s helped them to scale and become reliable and just debug issues. It was definitely one thing that we’ve experienced internally is how do we translate that so the rest of the developers, and when I say the rest of the developers, like application developers. Not the DevOps or infrastructure engineers who are either enabling ongoing ____. How do the rest of the application developers use Envoy to their benefit day to day?

    So I was trying to make a talk that would show how I use Envoy at Lyft to identify where issues are coming from and build that up to be more digestible to everyone else. Also tried to highlight that I know with the amount of work I’ve done on Envoy I have built my own blind spots about things that I’ve forgotten or were critical. So I was using an example that I had a very lose definition of HP status codes before I joined Lyft. I’m sure there’s so many application developers who don’t have that–who don’t have a solid idea of what those mean.

    So I would have to give presentations on what these meant or what these ____ meant and how they’re related to Envoy just so that way everyone would have the same set of tools going forward to better debug issues when they happen. So yeah. I was trying to make that digestible as a presentation.

    Ashley: Interesting. Yeah. I heard the talk went over very well. I’m sure you had a lot of people come up to you after with some good questions or wanting more information. It’s a great way to share. When you and others contribute back to the community, not just code or other things, but also talks, and people get to see real applications on Envoy and the technologies. Anything you’ve walked away with in terms of learning since giving that talk about either reducing meantime to detection or implementing Envoy for service meshes?

    Caramanolis: I think some of the bigger learnings is that–or maybe the learnings that resonated more with me is that these types of talks need to happen more especially within Envoy and giving say an intro debugging talk or even making interactive. So say if we’re able to set up a test environment where we had Envoy’s, like a mini microservice environment. Then you can do test requests and see how they failed. And after give people a hands-on safe way to learn how Envoy worked, that’d be really great because you can see everything on a slide. You can listen to it, it can resonate, but applying that to day today is a little hard without there being someone to bounce ideas off of.

    Envoy as a project itself has been really, really successful and there’s been amazing contributions and hearing ____. I was very lucky to work with a lot of really smart people from Google and all these different companies. I’m just mentioning Google ’cause those are the people I worked with most closely. All these really, really smart people, I got to learn from them indirectly or directly, I should say. But also at conferences people would come up to us and say, “Oh, I was trying to do this with Envoy,” and you’re like, “Oh, I never thought of that.”

    And Envoy Con, there was just like 30 minutes talk of people would come up with, “Oh, we’re trying to do this one case here. So we built our own filter and we did that thing that way.” You’re like, “Woah, that is really cool.” Just seeing how other people think a problem differently really could highlight either gaps in our own knowledge or like, “Oh, maybe we should try their approach.” So I think maybe that’s the most valuable thing about conferences. Maybe I could share with people, but hearing what other people have thought about this topic really teaches me a lot about different ways to solve a problem.

    Ashley: It really is. A talk really is almost like initiating the two way conversation gets to happen. That’s what kind of–seize it with lots of good information for people to come up and talk with you after. So that’s awesome.

    Caramanolis: Oh, yeah. Actually that’s a really good way of saying it. Yeah.

    Ashley: Well, what advice would you have for anyone who is maybe not experienced in using Envoy, if they’re just approaching it. How would you suggest that they get started and maybe are there some early lessons learned that you could share with them about how to more effectively use it?

    Caramanolis: Usually when these types of questions come up I always ask people what problem are they trying to solve. At least one common question that was at KubeCon, and since this is the context of KubeCon, is “Should I migrate to Kubernetes or Envoy?” I always reply back with, “What is your most pressing problem? Are you having issues with network observability or standardization of error cases, or is it that your current deployment pipeline isn’t working as expected?”

    So say they wanna focus more on the Envoy part, then I wanna–then it was like finding out more about their topology. So the Envoy configs. Anyone who has definitely worked with Envoy and has listened to this will probably giggle at this. The Envoy configs are definitely overwhelming, especially for those who are initiated into it. And so starting off really simple. Either Matt–I think Matt and myself, one of us had talked about how we had rolled that Envoy at Lyft. So we started off with either do it at edge or do at ingress with no service mesh or egress and slowly built that up there is like building up at one part of the interactions instead of doing it all at once because all at once, there’s so many components that can go wrong.

    One misconfigured value, it could be technically correct value, but say you put the wrong port, but you hear it everywhere, then finding out where that wrong is very hard. So definitely my advice would be starting off small, set a really clear scope, either ingress or egress out of one service or set of services or just at your edge and then building that trust within your developer community. So once you have that definitely start educating the rest of the developers who are using it saying, “This is how the errors look like and this is what that means,” and helping them see that value, ’cause once all the developers see the value it definitely–at least for my simplify their lives ’cause they no longer have to spend time like “I know this one error is failing, but either which service is causing it or do I know if it’s a network issue or is it bad application code?”

    With Lyft, Envoy was able to very much isolate it to service B is having a bad day and we can see if it’s having any other impact anywhere else, but at least we know where to focus on service B. If there are any questions, the Slack community, the flag channel for Envoy is very responsive if you do run into issues, like always ask questions there or post an issue and get hub and people do make a really concerted effort to reply as quickly as possible. I would say the community–

    Ashley: That’s fantastic.

    Caramanolis: I really respect them and love them and I do miss working with them on a more daily basis.

    Ashley: Can you maybe translate or transition into a little bit of what you’ve been doing at Omnition? What kind of work are you doing now?

    Caramanolis: Yeah. I’m actually gonna relate a little bit back to my KubeCon talk. Part of my KubeCon talk was talking about how I had this one issue. I know it’s hitting–I get, like say, if I’m gonna use an example, I think with the example I used and the talk was I can’t see photos. So I know it hits on the edge and it goes from service A to service B to service C. What allowed me to do that with Envoy is that Envoy has very clear metrics of your upstream and downstream colors. And to avoid defining that within here ’cause some people have different definitions, pretty much Envoy tells you what services your dependent on.

    So without having those metrics of knowing where your service is dependent on, sometimes it is hard to track down what piece of code is causing an issue. So one way actually people do that is actually with traces. If you trace–say I know that this request is failing. I could look for this trace and then see that it’s going from service A to service B to service C, which is really valuable. So what Omnition is trying to do is actually foot tracing on its head because usually the normal paradigm is to look at an individual trace and then try to correlate other data around it. Either it’s like an input value to a request.

    It’s either, “Oh, we know it’s always service B that’s having–after ____ is having a bad time.” And so it’s usually start from a really granular data point and build out the information. We take all the trace data and you’re able to see that everything went from service A, some things go to B, some things go to C, D. And then so if something’s going wrong it’ll be a red line and you’ll just say, “Oh, I know that service B is caught in this error. Let me look more into it.”

    Ashley: That seems extremely useful.

    Caramanolis: Yeah. And it’s like–it’s actually–it almost would make my talk from KubeCon obsolete ’cause I’m trying to teach the ____ with using Envoy as like I follow the service graph, but I follow symmetric that Envoy produces.

    Ashley: I wish you the best. It’s tools like that that really are essentially to be able to grow the infrastructure the way that we’re building software now. Congratulations on the move and I’m excited for you, Constance. I wish you all the best. Thanks for being on the podcast.

    Caramanolis: Mitch, thank you so much. It’s such an honor to be on the podcast. I had a lot of fun.

    Ashley: Well, I did too and I’m honored to have you on here too. You’ve listened to another DevOps Chat podcast. I’d like to thank my guest, Constance Caramanolis, software engineer at Omnition and thank you too, of course, our listeners for joining us. This is Mitch Ashley with staging-devopsy.kinsta.cloud. You’ve listened to another DevOps Chat podcast. Be careful out there.

    — Mitchell Ashley