Author: Mitch Ashley

  • DevOps Chats: Automic + CA + Broadcom, a Release Automation Journey

    DevOps Chats: Automic + CA + Broadcom, a Release Automation Journey

    The life of an acquired company can be an exciting one full of unexpected twists and turns, and that’s been true of Automic software. Scott Willson, a repeat podcast guest and longstanding member of the DevOps community, joins us again on DevOps Chats. He describes his acquisition journey as being akin to a “barracuda, eaten by a great white shark, eaten by a whale.” Scott is product marketing director, Release Automation at CA Technologies, a Broadcom company (the new proper name for Automic + CA + Broadcom).

    Scott shares what it’s like being a software DevOps company inside Broadcom, whose strong roots come from the chip manufacturing business. While software has its many differences from chips, of course, Scott shares how many of the DevOps principles and heritage from lean manufacturing have made the transition easier. There are valuable lessons here for all of us.

    Amid all the change, Automic shifted its products to the freemium model. Automic Continuous Delivery Director is now free to use for up to 10 active releases. We talk about the upcoming releases, which includes machine learning. Join us as we bob and weave through the acquisition story into the world of continuous automated delivery.

    Transcript

    Mitch Ashley: Hi, everyone. This is Mitch Ashley with staging-devopsy.kinsta.cloud and you’re listening to another DevOps Chats podcast. Today, I’m joined by Scott Willson, he’s product marketing manager, Release Automation at CA Technologies at Broadcom—that’s a mouthful. I think we’re gonna explore a little bit of that.

    So, our topic today is really catching up with what’s been happening with the Automic product and CA’s, this part of the acquisition back into Broadcom and just kinda catch up with Scott. So, he’s been on the podcast before. Scott, welcome to DevOps Chats.

    Scott Willson: Thank you, Mitch. Glad to be here.

    Ashley: Appreciate you being here. Would you start just by introducing yourself for any of our audience that doesn’t know you, a little bit about what you do, and tell us a little bit about CA Technologies at Broadcom.

    Willson: Yes, yes, so my name is Scott Willson—two Ls in Willson. I think many listeners may know me. I have been, in the past, an author with some of Gene Kim’s DevOps Forum papers, and I’ve been at the DevOps Enterprise Summit several times. I’ve been fairly active on blogs and all that stuff.

    And where my kinda claim to fame was, I came from a company called Automic. That’s spelled Auto-mic for the phonetic spelling of it. My background is, just to throw this in here, too, my journey, as I always tell people, is I started from a development standpoint writing the code, migrated my career kinda doing the DevOps thing before anyone called it DevOps, like many others.

    Ashley: Mm-hmm.

    Willson: And yadda yadda yadda, here I am now with a software company speaking all things DevOps and trying to help people get better at releasing their software.

    Ashley: You’ve certainly been a part of the community for some time now, so we appreciate all your contributions. Well, you’ve been through quite a journey. I know being acquired by CA and then CA being acquired by Broadcom—why don’t you kinda catch us up? Maybe tell us a little bit about that journey. I think you described it as an 18 month history going through that.

    Willson: [Laughter]

    Ashley: You know, we don’t have too long on the podcast, [Laughter] but if you can kinda give us a little bit of that path that you’ve been on.

    Willson: Yeah, yeah. You know, I kinda looked at it like a, it was almost like a barracuda eating a great white who was eaten by a whale, right?

    Ashley: [Laughter]

    Willson: It was a really interesting, very quick journey. It was a couple years ago when I was with Automic Software. We were acquired by CA Technologies, and about a year or 18 months after that acquisition or so, Broadcom bought CA Technologies to the—not just our surprise, but of course, to the surprise of the entire market.

    Ashley: Mm-hmm, yeah, it was.

    Willson: There was a lot of press about it.

    Ashley: It was a surprise.

    Willson: And so, yeah, it was a very, very quick and interesting turnaround. We are now known as CA Technologies, a Broadcom company. As far as Automic goes, we are retaining the Automic brand, so you’ll see a lot of our old software and all that will have the Automic brand. And not just Automic, but our current leadership has really felt to really keep and embrace the brands that we’ve had. So, you know, Blaze—CT is still gonna be Blaze, and Rally is gonna be Rally, those type of things.

    Ashley: Mm-hmm.

    Willson: So, the brands are remaining, which I think is useful, because even though CA had remained Rally—I know I, anyways, always refer to it as Rally, because—

    Ashley: I do, too, yep. [Laughter]

    Willson: – right, because that’s what I always knew it as. [Laughter]

    Ashley: Especially knowing them from the beginning, it’s hard not to call them Rally, but yes.

    Willson: Right, exactly. So, Automic is still going to be Automic from a product brand positioning. And then from a company standpoint—yeah, we’re still CA, just a Broadcom company. Within Broadcom, our group is known as the Enterprise Software Division, so we are a distinct division within Broadcom.

    Ashley: Mm-hmm.

    Willson: And then, to add to that, to the little journey, what has been interesting to me is—I remember it was a year ago, the last Forum meeting I was going to or what many of us in the inside called those meetings, the Gene Kim’s Great Slumber Party.

    Ashley: [Laughter]

    Willson: Right? We kinda talked about being lean, and some of the people, I remember having discussions with and running papers, we just really got to talking about what it means to run lean and the lean manufacturing and the whole thing, right, that DevOps is based on and that The DevOps Handbook espouses and all these kinda things.

    And after that, right, three months after that meeting it was announced Broadcom was gonna buy us and then they actually did, and here I am all this time later. And I bring that up because what was interesting to me, Mitch, is that Broadcom embodies this. They run lean.

    Ashley: Mm-hmm.

    Willson: And it’s been an interesting journey to go from the short time at CA, which ran big, to now being with Broadcom, which runs lean. And they run even leaner than what Automic did, which was in ISV, right? You know, Independent Software Venture.

    Ashley: I was curious about that, because I know Broadcom, you know, I had worked with them as a chip supplier with companies like Broadcom ad Qualcomm and, you know, manufacturers of chips in telecommunications think about the world, they’re much more, you know, operational efficient. It’s all about fractions of pennies and the cost of their goods to maintain profitability, but that’s super high volume. So, it’s a different mentality than a software company.

    Willson: It is, and there are some significant differences, I suppose, from the business or sales aspects. But from a production aspect, there really wasn’t a whole lot of difference. When you think about taking a product to market, right? Now, we were doing this with software, they’ve done it with chips, but Broadcom is an engineering excellence company. They really pride themselves on having the best of fill in the blank of engineering, whether it’s engineering processes, engineering whatever. In fact, what has been interesting post acquisition is that our engineering staff for the bulk of our products have actually increased. Broadcom is very keen on ensuring that the products we deliver are best in class.

    Ashley: Mm-hmm.

    Willson: And so, they realized a way to do that in the market is to make sure they’ve got the top minds developing and producing and shipping these products, right?

    Ashley: That has to be refreshing to you—you and the team.

    Willson: Oh, yeah. It was very interesting to see that that was the case. It’s been a unique acquisition as far as that goes, and to see that—no, they wanted to increase the head count of R&D. [Laughter] It’s important for them.

    Ashley: Mm-hmm. What? [Laughter]

    Willson: Exactly. Like—wait, what? They want the products to improve, to get better, and to continue to be viable and market leading. And so, they embody the lean processes to do all of this stuff. And so, it’s been—there’s a little bit of a learning curve, I think, on both sides to adapt to this new way, but here we are. You name the things that are automated—I mean, Broadcom automates just about anything and everything. The very stuff you read about, Gene Kim talks about in his book, saying—and even in Nicole Forsgren’s book, Accelerate, you know, automation is at the key of a lot of these things, and Broadcom embodies that.

    Ashley: It is. It’s a much different kind of company. If you’re not a hardware person or haven’t worked at a chip company, what I said, it’s—fractions of a penny make a difference in cost, and that cost can be in materials, but it also can be in process. And when you multiply that by millions upon millions of chips that they may ship a particular line, that’s a lot of money. So, it makes a huge difference in the profitability of the company—

    Willson: That’s right.

    Ashley: – to be able to reinvest and give you more engineers to be part of your team.

    Willson: That’s right. And so, software is a little different, right? It’s not the pennies like that. Obviously, with software, the margins are different, because you’re not dealing with a physical thing, a commodity, right? But the principles to be efficient like that are absolutely espoused by DevOps. And one of the things that this transformation internally has occurred—because we’re transforming ourselves from a software company, right? I think we’ve made some press releases talking about our BizOps platform, a holistic platform—

    Ashley: Right.

    Willson: – that we’re offering the market. No longer should you be thinking of us like the CA of old where you had a bunch of different siloed tools. Now, they’re all coming together in a holistic way to provide real business value and create a true, interesting platform for everybody.

    Ashley: Mm-hmm.

    Willson: But just this way of doing things lean like that and embracing the agile way of turning things around has been a cool thing to behold.

    Ashley: That has to be really good, because I haven’t talked with somebody that’s had that experience—not that others haven’t, but to have that common lexicon frame of reference of thinking about producing software like a software factory just like a hardware based factory. You have a lot of synergies and commonalities there. Yeah, there’s differences, too, but instead of software being this mystery to hardware guys or vice versa, right, you can talk and work together and understand the efficiencies, improvements, speed, impact, quality—all the things that drives the manufacturing company also drives software.

    Willson: That’s right. And what’s been exciting, too, with this opportunity to change—really, the Broadcom offer has provided us an opportunity to change things and to really force these integrations and synergies between what used to be separate divisions of CA but are now different teams and the interoperability that’s happening between them to deliver this BizOps platform is really exciting. It’s cool to see that now it’s like—look, we have all these things, but what’s the point, unless we’re focusing on business outcomes, right?

    Ashley: Right.

    Willson: We can actually drive a business outcome if there’s harmony and cohesiveness between all these things, and that’s very exciting. In fact, one of the other things along these same lines that was announced is the free tier program—I don’t know if you see the PR a couple weeks ago, Mitch—

    Ashley: I did.

    Willson: – but now, a lot of our products are offered for free.

    Ashley: Yeah.

    Willson: So, for example, a continuous delivery director, or Automic Continuous Delivery Director, if you go to CDDirector.io, you can just use it for free, up to 10 active releases. And Rally has some free tier as well. So, basically, we’re now offering the freemium versions of what we had, and—

    Ashley: Have the products changed substantially to be able to do that? Are they less functional or are they time based or just volume, kinda getting folks to up to date?

    Willson: It’s really just volume, right? Yeah. They haven’t really—well, a lot of them didn’t really need the change. A lot of the products that are offering this are the ones that are already SaaS based already. There will be more that are coming, of course, as things are being migrated to be SaaS based, and that effort is under way.

    But an example of Automic Continuous Delivery Director, there’s not really a time limit. It’s 10—10 active releases. That’s a lot of releases.

    Ashley: Mm-hmm.

    Willson: And you can just use it, and use it at will for that. You wanna get more than that or grow it across the enterprise or make it the standard or something, same with the Rally and Blaze CT. Oh, well, yeah, then you can buy the licenses that let you use it more broadly, right?

    Ashley: Mm-hmm. That’s a great model. I mean, I think we’ve sort of disproven the crippleware approach. Nobody likes that.

    Willson: [Cross talk], yeah.

    Ashley: Exactly, and being able to use it on enough of software production; you talked about 10. You’ve gotta be able to get a pretty good feel for what the product does that’s gonna do what you want it to do and start to integrate it into your tool chain workflow platform.

    Willson: Right, and these are 10 active releases. In other words, 10 releases in parallel, right?

    Ashley: Mm-hmm.

    Willson: Well, that’s—like you said, you should really be able to get a feel for it.

    Ashley: Yeah.

    Willson: Being able to run that many releases in parallel to each other.

    Ashley: Well, that’s awesome. That’s great to hear. I’m excited to hear about more of that coming from you, too.

    Do you have a new release coming up, here?

    Willson: Yeah. So, speaking along those lines, we have a new release for Automic Continuous Delivery Director. We’re pretty excited about this release. It is our first release that includes some ML capabilities and heuristics that are involved. Basically, what we were looking to do with this is, amongst other improvements and advances we made in the product, right, to help with continuous delivery, one of the things we looked at is that there’s significant bottleneck in continuous delivery, the methodology, which is QA or testing.

    Ashley: Mm-hmm.

    Willson: Now, as you probably know, Mitch, continuous testing is technically a subset or a component of continuous delivery.

    Ashley: Right.

    Willson: And I might argue that unless you’re doing continuous delivery, you are not doing continuous delivery.

    Ashley: Right.

    Willson: You need to be doing—

    Ashley: You need to think of continuous delivery is plan, dev, test, deploy, operate, right?

    Willson: That’s right, and you’re testing continuously, right? But the problem is, QA wasn’t really built to be agile themselves. So, they kinda became the bottleneck, not just from the way that they run, but you, as a QA person, right, testers, you build out all these tests. And so, you release this small change per DevOps per agile per continuous delivery, right? And it goes into QA and you run, you know, 1,000, 10,000 tests—and I’m not making these numbers up. Literally, a lot of our large enterprises have thousands and 10,000 tests.

    Ashley: Oh, yeah. Totally.

    Willson: Alright, so then, we’re in the waiting game, right? And maybe it takes however long for the test results to come back and then you get feedback to development. So, we found with our, a lot of our large customers, you know, the Global 2000, that this was a significant bottleneck. So, now what we do, what occurs, is that when you do that release and you’re a pipeline, right, you’re a release pipeline and it gets pushed to QA, one of the things that we do is, we look at several heuristics.

    So, we come around and say, “Alright, well, what are all the tests? What is the historical responses or outputs of the tests?” Right? How many of them failed? Which tests are flaky? How many tests are new tests? We want to actually run these tests first. So, a flaky test, for example, it ought to run at the front, and if it fails—pfft, we’ll send that feedback immediately to development or to QA so that we’re identifying these things earlier, quicker, faster, rather than having to wait on the back end for this, right?

    Ashley: Try to get to that 80/20, right? Get the things—

    Willson: That’s right.

    Ashley: – upfront that are gonna cause the most issues, find the most bugs.

    Willson: Right. So, we’re looking at all the data and intelligently figuring out the tests for you. The other thing that we’ll do is that we’ll look at the code changes you’ve made and we will map those code changes. At least in the case of Java apps, we’ll be able to map the code change to specific test suites.

    Ashley: Mm-hmm.

    Willson: So, now, when you submit that one component change, you are actually getting identified just the test runs that affect or exercise that component. So, altogether, when you put it all together, Mitch, what occurs then—now, when a change gets pushed into your continuous delivery pipeline, when it hits QA, you’re in a position to not have to run all 10,000 tests. You’re in a position now to run the high risk tests, the tests that you have marked are critical and important, right? So, those are included, that’s taken into account, too. And the test that we know actually exercised the change that you put in.

    So, you’re able to run what you have to run to ensure you have a high quality and also order them in such a way where there’s high likelihood of failure that Development is fed back that feedback earlier and faster. And so, we’re very excited about this release.

    Ashley: That’s awesome.

    Willson: Yeah.

    Ashley: Boy, gone are the days, I can remember when I would judge the quality of the release by how severe the bugs are and if we haven’t found really big problems yet, we’re not through. There’s more to come. [Laughter] You know, that was the heuristic. [Laughter]

    Willson: That’s right. Exactly. That was the heuristic.

    Ashley: That’s right. We’re far from those days, aren’t we? [Laughter]

    Willson: Well, you know, because one of the other constraints, if you think about agile and continuous delivery, right—small change and you make more frequent small changes and the “State of DevOps” report just came out, right, a week ago or this week, and it again is showing how this way, this approach is actually reducing the bugs and reducing ________ quality.

    All those things are great. But that now means one of your biggest resource constraints is time. And what we’ve found with our customers, it was time in QA. So, by intelligently assigning the tests on a run with a particular change, we’re able to shorten that down. You’re not having to wait for all 10,000 tests, you’re just waiting for the five that matter to run against that little small change and—pfft, get it shipped out the door.

    Ashley: Exactly, right. Be smart about testing, yeah.

    Willson: That’s right—be smart about testing. Exactly right.

    Ashley: Now, do you have a date when this is going public? Are you in beta now? Where are you at?

    Willson: Yep. So, if you were an on prem user of the product, it is actually GA now. It actually was up on the download site yesterday and—

    Ashley: Oh, nice.

    Willson: – it can be downloaded and you can go through the upgrade process. If you’re on our SaaS platform, you know, October 2nd, it’ll just be online working and available for you in the beauty of SaaS, right?

    Ashley: Nice. Well, great. That’s fantastic. Awesome. Well, congrats on the great release. I hope it works out very well and also the free tier brings some new customers, new people onto your platform and expand that as well.

    We have a few minutes left. I know you’re gonna be at the DevOps Enterprise Summit, correct, in Las Vegas?

    Willson: Yeah, exactly. So, we’ll have a big booth there at the DevOps Enterprise Summit, so certainly looking forward to seeing everybody wearing our Broadcom wear. We’re gonna—not to give away too much, but I think we’re gonna have some pretty cool giveaways at our booth at various times, so, you know, definitely come check us out.

    Ashley: Come early, come often for the swag.

    Willson: What—get the swag? [Laughter] That’s half of what you go for. Especially if you have kids, you know? I see some of these people, they’ve got bags full of this stuff, right?

    Ashley: Mm-hmm.

    Willson: And I’m like—yeah, you’ve got kids at home. [Laughter]

    Ashley: You don’t wear a Small—you don’t wear a Small. What are you—what are you doing, there?

    Willson: [Laughter] Exactly right. Exactly right. But we’ll have some things, a couple of the giveaways we’ll have I don’t think you wanna give to the kids, let’s just say that.

    Ashley: Okay.

    Willson: And if you’re going to the Gartner Symposium, we’re gonna be at the three upcoming Gartner Symposiums in Orlando, São Paulo, and Barcelona as well.

    Ashley: Fantastic.

    Willson: So, yeah.

    Ashley: Well, good. We’re gonna have to check in every six months or so just so we know what swag wear, logo wear to look for you in.

    Willson: [Laughter]

    Ashley: This time it’s in Broadcom—I’m kidding you, of course. But almost—it’s almost true.

    Willson: It’s almost true. Feels true. [Laughter]

    Ashley: Scott, it’s been great having you on the podcast. Appreciate you being here.

    Willson: Thank you so much, Mitch. Appreciate it and look forward to seeing you and Alan at the DevOps Enterprise Summit.

    Ashley: Absolutely. Look forward to it as well, and great to catch up with you about everything happening. Well, I’d like to thank you. I’d like to thank Scott Willson, Product Marketing Manager—or sorry, Product Marketing Director of Release Automation at CA Technologies, a Broadcom company—I think I got that right—

    Willson: You got it right.

    Ashley: – for joining us on the podcast. [Laughter] And thank you—you, our listeners, for being here as well. This is Mitch Ashley with staging-devopsy.kinsta.cloud. You’ve listened to another DevOps Chats podcast. Be careful out there.

    — Mitchell Ashley

  • DevOps Chats: Enterprise Continuous Testing With a Shift-Right Mindset

    DevOps Chats: Enterprise Continuous Testing With a Shift-Right Mindset

    Test automation and CI/CD are evolving very rapidly to achieve great speed and impactful results by DevOps teams and Agile organizations. Our DevOps Chats guest, Tricentis Chief Product Officer Wolfgang Platz, contributes his experiences to the state of the art with his newly released book, “Enterprise Continuous Testing.”

    Automating software testing just for the sake of automation is “doing the mess for less.” What is the business value of each software element you are testing, and what is the right strategy for testing these software implemented functions? Automation certainly brings with it speed, but are we ultimately making incremental improvements to testing which doesn’t tap into the full power and benefits of automated software testing.

    Our discussion with Wolfgang focuses on the need for expanding dev-testing with higher-level integration, system, and end-to-end user experience testing, which are often performed well by shared testing and operations teams. Now, think Shift Right. While Agile and dev teams often focus on the progression of creating new and exciting capabilities for the business and our customers, we still require that regression mindset that consistently validates unit tests, valuable business functionality, system integrity, performance, and operational needs.

    Transcript

    Mitch Ashley: Hi, everyone. This is Mitch Ashley with staging-devopsy.kinsta.cloud and you’re listening to another DevOps Chats podcast. Today, I’m joined by Wolfgang Platz, who is Chief Product Officer and Founder at Tricentis. Our topic today is a new book that he has just released, “Enterprise Continuous Testing.” It’s a great topic, talking about agile and DevOps and how we transform testing within that environment. So, Wolfgang—welcome to DevOps Chats.

    Wolfgang Platz: Thanks, Mitch. Appreciate it. Thanks, everybody for listening to this podcast. I’m actually on it to be with you and have the chance to introduce a bit of what Tricentis is about and what this book is about, and thanks for having me.

    Ashley: Absolutely. Well, let’s jump right—and we’re honored to have you. Thanks for joining us. Why don’t you jump into, just tell us a little bit about yourself, maybe why you founded Tricentis, and then we can get into the book.

    Platz: [Laughter] Actually, that’s a great question. If you look us up on the Internet, you’re gonna see our official founding date, which is 2008, but the journey actually and the story, the history starts a bit earlier.

    I founded Tricentis as a services company first in 1997, and bringing some software QA services to the world, but I had to acknowledge, frankly, that all the tooling out there just wasn’t sufficient when it came to test automation. When you do test automation, I always call it, you run into kind of a honeymoon phenomenon. Meaning, you record your first test cases and everything looks bright, but when you try to replay them and when you want to have them up and running even when changes happen, then you fail and things break.

    So, the only way out was to come up with my own tool there, and that was the birth date of our software product offering. Then, in 2008, it was mature enough to switch into a product company, and from there on, it has been a fantastic journey upwards. Now, we are, as you may know, placed as a leader in Gartner’s magic quadrants now for the fourth time in a row, so.

    Ashley: Mm-hmm. Congratulations, yeah.

    Platz: We are about to take it all over, that’s great.

    Ashley: Well, fantastic. You know, it’s a familiar model that’s kinda transitioning, pivoting from that services into a product company—and congrats on that successful transition. It’s not always a successful one. [Laughter]

    Platz: Yeah, that’s true.

    Ashley: So, if you started in 2008, that was probably around the area of agile. DevOps wasn’t really kinda taking off quite yet. I think that’s about when it kinda started. So, you were already, I’m sure, thinking about how teams are starting to work differently in agile methodology, maybe kinda leaning, starting towards a lean kind of an approach.

    Platz: Absolutely.

    Ashley: Yeah, interesting. Well, tell us about why did you write this book. I know continuous testing, CI/CD—all of that stuff is certainly at the heart and center of DevOps and everything kinda starts there. Automated testing is, of course, a key part of that. So, certainly there is a need for this topic. Why did you decide to write a book?

    Platz: So, what I had to acknowledge in talking to customers and coming up with this product was, software testing is treated in a lot of organizations like a stepchild. And I mean, everybody is aware of the relevance of development, don’t get me wrong, and everybody is aware of the good things that IT can bring to use.

    But software testing is kind of living in the shadows, and that was one experience. And when I then talked to customers, I found out that their ask and their intuitive need is always kinda the same. They see that software testing is here, they to some degree accept that there is a need for software testing, but all of a sudden, they say, “Well, I mean, it’s a lot of effort, so we wanna have this automated.”

    Ashley: Mm-hmm.

    Platz: And the first reaction to dealing with software testing in a more professional way is, “Let’s automate that thing.” And automation is a super element aspect of an improved software testing, don’t get me wrong. However, bringing large enterprise customers up to speed on software testing is actually more like a transformation journey. It is a change agenda you need to go through.

    Because it’s not just about automation. A friend of mine said automation is just doing the mess for less, right?

    Ashley: Mm-hmm.

    Platz: But there is so much more potential in optimizing software testing when you have a more comprehensive look at it, when you have a clear understanding what is the business value of each and every functionality you wanna cover with tests. What is the concrete need for test cases in order to address the business risk that may be within certain functionality, and then what is the right strategy of executing these tests? Are you gonna run them through user interfaces? Are you gonna run them through APIs? Are you gonna run them with the use of decoupling, which service virtualization enables you to?

    Only if you have that comprehensive perspective, then you can make continuous testing really work. Otherwise, it’s just gonna be little things, little improvements that certainly help, but don’t untap the full power of a better software testing world, and that is why I’ve been coming up with this book. I think it is overdue to make people aware—“Wait a minute, it’s not just jumping about onto the automation train. This is not gonna be the full story.” And you’re gonna leave a lot of stuff behind if you just think that way. Does that make sense, Mitch?

    Ashley: It totally does, and I think if I kinda pull a couple things out of what you described, one is—automation is part of it, but continuous is also another aspect of changing the kinda whole paradigm of how you think about testing. Because I remember it wasn’t that long ago where you had a QA team and a development team.

    Platz: Correct.

    Ashley: And of course, that created automatic tension because the QA team tried to find every problem and what has to be fixed before it gets released and change control meetings and all that kinda stuff. It was highly manual, even if some of the testing was scripted. But now, in this world of continuous integration, continuous testing that’s automated you know, it’s less an argument between teams and people and that kinda friction and more about how do we make sure we’re automating the right things? How do we make sure that we escalate when problems need to be fixed, when builds break and tests fail and make sure that we’re testing the right kind of functionality?

    So, it’s engaged the developers—and tell me if you agree with this—it’s really engaged the developers much more heavily in the creation and embracing that testing rather than being a separate function someone else does. True?

    Platz: Yeah, absolutely. I mean, you pointed it out already, Mitch. In former days, it was like Dev, Test, and Ops providing different teams to software development life cycle. And what has happened, though, is that these big, we call them test centers of excellence where you would have a large number of manual testers, usually concentrated with some management capability, a little bit of automation maybe, even.

    This piece, these large TCOEs, these test centers of excellence actually, they have either vanished or they are—or everybody’s questioning themselves if they should keep going with this model. And fully right so. Because what we now see with the need for speed in development is that the cooperation between Dev and Test needs to be so much closer, needs to be so much more fluent.

    So, what we see is that these test centers of excellence are torn apart. And what we call out there as a recommendation is that you, if you have a very slim, so to say, team, a best practice team or we call it digital TCOE, that on one hand has this tight connection between Dev and Test, but on the other hand, provides some guidance on how to do testing in the best way.

    Ashley: Mm-hmm.

    Platz: And we see a lot of companies switching into Dev Testers where you have this hybrid form of developers and testers established directly within the agile teams.

    Ashley: Mm-hmm.

    Platz: And I think that’s certainly the right way to do. However, on the other hand, we wanna be very much aware that complex system landscapes require higher levels of integration testing and some flavor of an E2E user acceptance test, which tends now to be forgotten when people go into an agile and DevOps practice.

    So, I think—yes, we see the TCOEs erode. They got into Dev testers, which makes a lot of sense, but please, guys, be aware that there is a need for higher levels of testing which has not vanished overnight, and someone needs to take care of that. What we see is that these needs tend to shift more towards the operations team towards a kind of shared services team now in large enterprises, but it is a, what I would say contra streaming, like the shift left pushing testers into development. We now see kind of a shift right movement also, which makes sure that some higher levels of test are still covered in these shared services organizations.

    Ashley: Yeah, it’s interesting you pointed that out that way, because you know, you do hear a little bit about shift right and you think about the world that we’re operating and it’s much more complex, because we’re operating in cloud environments, maybe doing cloud native, maybe not, but we’re working maybe in multiclouds, private data centers, private clouds, et cetera.

    All these combinations of environments—and, you know, some apps might be in one place, some might be in multiple. There’s a lot of permutations of that environment that an Operations team is going to say, “How do I know this stuff is gonna work, and when it breaks, how do I know you’re gonna be able to fix it quickly? Diagnose it, test it, and apply the fix quickly.

    Platz: Absolutely.

    Ashley: And that automation is also the automation of—yes, unit and functional tests, but also getting into being able to do some performance regression testing, those kinda, that may be specific to that environment, maybe even specific to an iteration release or kinda backing up and reapplying changes or backing out changes.

    So, it’s in this complex environment that relying on manual processes seems like almost an impossible task. You have to have automation, you have to have a philosophy about doing testing. I hate to wax on here. Maybe I’m singing your song here, [Laughter] but that seems to be the environment we live in and that we have to figure out the best ways to do it. Am I on track here?

    Platz: No, no—that’s cool. And I’m glad that you mentioned regression, because what we have to acknowledge is that the entire mindset of agile development, per definition, is progression rather than regression. And that is good, because agile development is about pushing out new releases, pushing out new capabilities, new features, new functionality faster than ever.

    But what that mindset means is that all the agile Dev teams and the Dev testers that are within the agile teams, they also have this progressive mindset. It means that yes, if you work with them, they will accept the need for unit tests, and some of them will create great unit tests. We have seen this within our customer base. But at a certain point of time, when this is shifting towards regression, then you might find out that having unit tests set up once is one thing, but keeping unit tests up and running is a different game. And what we see is that as soon as you shift from progression into regression, you need to have people with a different mindset that really accept regression as the key aspect of their software testing behavior. If you don’t do that, you, time over time, are gonna run into issues in complex system landscapes.

    And we see that, actually, I call it the evaporation of unit tests over time, which we have done a great survey with some Swiss banks, and we found out that having unit tests set up is something you rather easily can do. But over time, with more releases, you’re gonna see that their coverage, their potential, their power is eroding. And people are then referring more to the higher levels of integration to also catch these kinds of errors going forward, which makes a lot of sense.

    But I just want to make you aware—don’t see agile development and Dev testers as the sole instance of your software testing. If you do that, you’re gonna run into issues in a complex, multi-system landscape with a high extent of regression with a heavy backbone with a true enterprise system landscape.

    Ashley: So, Wolfgang are you saying that—you know, we used to say that you had to have separate QA teams, so it was people who didn’t create the software that could more fully test it and not kinda, you know, testing your own software is hard because you assume things.

    Are you saying that, even in this automated world, continuous testing world, we still need folks who have that sort of systemic thoroughness QA kind of approach to testing, even though we’re automating it, don’t strictly rely on developers to create a self-created unit and system kinda test. Is that what you’re saying?

    Platz: Exactly. That’s what I wanna get to.

    Ashley: Good. Well, tell us a little bit about your book. You know, there are a lot of ways to approach a technical book about a technical subject. Is this a prescriptive how to? Is this kind of a strategy guide to lay out how you put together a full program of continuous enterprise testing in agile and DevOps? Tell us a little bit about the approach of how you went about this book.

    Platz: The intent of the book is to provide a comprehensive overview of what you need to think about when you go for continuous testing. As I pointed out at the very beginning of our conversation is that just going for an automation of your software tests is not gonna be the solution. You need to have a way more comprehensive look onto the subject.

    And so, what the book really is gonna do, it’s gonna introduce the perspective of a value-slash-risk based testing to you, which is the foundation. Before we wanna do a software test, we wanna make sure that it’s worthwhile doing the test, right? If this is just a functionality with very, very minor relevance to the business, you might be good with just one test case, you know what I mean?

    Ashley: Mm-hmm.

    Platz: However, if it’s an absolutely mission critical capability of your software, then you want to make sure that all the different flavors of use cases and all the different procedures that may be within this specific use case are really covered, otherwise it’s, from a risk perspective, not bearable.

    So, having a clear understanding about the business risks-slash-the business value that is associated with each and every functionality that your software provides is the foundation. If you understand that clearly, you will be able to present kind of a map, your map of where you wanna put your dollars into. And guess what? This map is not gonna be just for your budget allocation, but it’s also gonna be relevant for your reporting, because your reporting should go towards business risk covered and not just for counting test cases.

    And guess what? Your management is gonna love it. All of a sudden, they see that somebody has a plan on how budget is allocated in the software testing space. So, it’s real cool stuff.

    As soon as you know that, you wanna make sure you have the right test cases at hand. We’re gonna give you an overview about different techniques of how to create the most meaningful test cases—meaningful in terms of what is the extent of additional business risks that I can cover with the specific test case, right? And when we have done that, we’re gonna make you aware that there’s different approaches to go towards automation. You can go through a user interface, you can go through an API. When you go through this, make sure you decouple systems so that you always know exactly what is the cause of a failure.

    Ashley: Mm-hmm.

    Platz: And when you have come up with that, you wanna make sure that you keep your test data stable, meaning that the system always has a reliable set on basic data at hand so that you’re not gonna be surprised by failures of tests which are just a matter of inconsistent or incorrect test data.

    So, all these things are gonna be introduced in the book. Of course, not into the very, very level of detail, but I think at least in a way that you know what it is about and you have a good basis from where you can start digging further into stuff. But ideally, you take the book and you walk home with it and say, “Uh huh. Now I understand what it is all about.” And I’m not gonna just go out there and jump on the next open source framework for test automation and say, “This is gonna be the holy grail of my whole journey,” you know? It depends on what you wanna achieve.

    Ashley: It’s about kinda creating a systemic, prescriptive framework for how to do this kind of work, because testing is, it sounds like such an understandable word, but it’s also so overloaded, because, as you talked about, there’s lots of ways to test, whether it’s through the API, the interface, there’s security testing, there’s black box testing, and you brought up the whole topic if test data. It’s hard to acquire in the first place and you have to main that, you know, evolve it with the functionality and the application—

    Platz: Oh, yeah.

    Ashley: – part of the regression, also. It’s really a very complex topic, and you get into performance testing, load testing, sort of chaos monkey-style resilience testing, all kinds of things that can happen. And then behind the scenes, now that we’re automating so much, we have this data that’s super valuable that we can communicate to the business. We plugged in code in here and it made it into production, and here’s how we know it was thoroughly tested and made sure that it’s gonna be reliable when we get it out in production.

    Platz: Yes, and this is actually where the journey is gonna go looking forward. What we’re gonna see happening is that more and more tracing information, logging information from production will be fed back into the testing loop, right?

    Ashley: Mm-hmm.

    Platz: I’ve pointed out the business risk courage. I mean, how do you get to a clear understanding of business risk? You get through that by knowing how often systems are used and what is the potential damage. These things—and nowadays, if you have a proper logging out there—can be obtained from production logs. And these production logs can be fed into, back into the loop by now influencing your test awareness on specific functionality from the very beginning, from the development perspective already, into the higher levels of testing.

    So, what we’re gonna see is that data, the use of data, even with some artificial intelligence involved, is gonna be a tremendous source for driving the further efficiency of the testing cycle. So, we’re gonna see, actually, the ideas of the value based testing in conjunction with having the right test cases at hand being even more optimized and speeded up by this data loop that we’re gonna get into the future.

    Ashley: Mm-hmm. So, your book is available, I assume, through normal channels like, you know, an Amazon and other places like that?

    Platz: Yes, correct. Just type it in, “continuous testing,” and you may wanna put my name in there also, and then you’re gonna be good to go.

    Ashley: Awesome. Excellent. Well, congratulations on the launch of the book. That’s exciting. I appreciate all the work. I haven’t written a book, but I’ve talked to many people that have, and I know it is a painstaking love of passion for the topic that kind of sees you through to the end and getting the book published. And I know you had some help with Cynthia Dunlop, so shout out to her for helping you on the book.

    Platz: Yes, yes, yes. It actually, that was a dramatic understatement, Mitch. Cynthia has not been of just some help. I would say she is the true author of the book. I was the one laying out the concepts, and it was stunningly how little time I had to provide to give all the basic information for her to come up with it. All the kudos actually need to go to Cynthia. I’m super excited to have my name on the book, but I feel a bit ashamed because it should be her name more prominent than mine, frankly.

    Ashley: Well, it sounds like both are well deserved, and your appreciation is certainly transparent to Cynthia.

    Platz:[Laughter] I hope so.

    Ashley: And, you know, we never go alone, right? We never get there alone, so it’s through working with other that always—it gets all of us there together, so. Well, congratulations on the book and thanks for being on DevOps Chat with us today.

    Platz: It was my pleasure. Thanks, Mitch.

    Ashley: It was my honor to have you. I’d like to thank Wolfgang Platz, who is Chief Product Officer and Founder at Tricentis, and of course, wishing him the best of luck with the book. Thanks to you, all of our listeners, for joining us today. This is Mitch Ashley with staging-devopsy.kinsta.cloud. Have a great day and be careful out there.

    — Mitchell Ashley

  • DevOps Chats: InfluxDB Cloud 2.0 Managed Service for Time Series Data

    DevOps Chats: InfluxDB Cloud 2.0 Managed Service for Time Series Data

    We can’t seem to generate enough data. And as they say, you haven’t seen anything yet! Cloud-based managed services are a natural solution to ingest, store and access large amounts of data from a variety of sources across private and cloud locations. An enterprise solution hosted in the cloud is good, but a developer-friendly, API-based solution with non-enterprise usage-based pricing reaches an even broader audience.

    Enter InfluxDB Cloud 2.0. The name just about says it all. InfluxData VP Products Tim Hall joins DevOps Chats to discuss the company’s new cloud-based time-series database offering.

    InfluxDB Cloud 2.0 represents that shift to a cloud-based, developer-friendly offering with much greater accessibility. We discuss the use cases; the free and usage-based pricing; the new FLUX language for querying, analytics and data processing; common APIs; and that TICK stack. TICK stands for Telegraf, InfluxDB, Choronograf and Kapacitor. It is a feature-rich conversation, so join us.

    Also, check out InfluxData’s Oct. 16th webinar; Optimizing Time Series Performance in the Real World.

    Transcript

    Mitch Ashley: Hi, everyone. This is Mitch Ashley with staging-devopsy.kinsta.cloud, and you’re listening to another DevOps Chat podcast. Today, I’m joined by Tim Hall, Vice President, Products, at InfluxData, and our topic is their new offering, InfluxDB Cloud 2.0. Tim, welcome to DevOps Chat.

    Tim Hall: Hey, Mitch. Thanks for having me.

    Ashley: Excellent. Great to have you on the podcast. Would you start by telling us a little bit about yourself, what you do at InfluxData and maybe for somebody who doesn’t know InfluxData does, give us a little bit about Influx.

    Hall: Sure. So, I’m the Head of Products here at Influx. I joined almost three years ago. I’ve got a background in big data open source technologies. I spent some time at Hortonworks running the Product Management Team there. They’re focused on the Hadoop space. Prior to that, I spent time at Product Management at Oracle addition HP, mostly in the monitoring and data integration spaces.

    So, Influx is a great opportunity to join a small company at the time. We’ve grown to almost 160 people now over the past 3 years, and it was sort of at the convergence of all the things I’d done in my career, from monitoring technologies to integration to big data. And so, it was great to be invited by Evan Kaplan, our CEO, and Paul Dix, our Founder, to join the team and see what I could do to help build a great product.

    Ashley: Fantastic. And tell us just briefly—InfluxData, what do you do?

    Hall: So, InfluxData is focused on providing a platform for time series data, which is about anything with a time stamp. Time series can be broken into two categories—regular series, which are things that you sample. Think of it like sticking a coffee cup in the river and pouring it into a bucket, and we’re the bucket, and you do that on a frequent basis—every 30 seconds, every minute, every 5 minutes, but it’s regular. It’s a regular time interval. And irregular series are also supported by Influx, which are things like events or logs, things that happen at any point in time that could be somebody doing key card access to a building, putting their ATM card into an ATM machine—those sorts of things. Those are all events.

    And so, we handle both of those two things and we have lots of folks using us, both in the open source community, on prem, or wherever they decide to play the software and now our cloud offering as well across a wide range of use cases, everything from DevOps monitoring use cases and building and assembling platforms for observability across to IoT use cases, both for consumer and industrial products.

    Ashley: Well, in a world where we just can’t seem to generate enough data, it sounds like you all are well positioned in the market. [Laughter]

    Hall: There’s more coming every day, right? So—

    Ashley: Yeah, it’s not getting less.

    Hall: Yeah, what do you do with it is really the question. And I think, again, we’re focused on providing a platform for developers. And our goal is really to drive developer happiness. And our motivation is trying to allow developers to solve problems quickly from the time of install to problem being solved. We call that the time to awesome. We want that to be as short as possible, and that’s not always the case with a lot of technologies out there. You’re fiddling with dependencies and other things and Influx comes out of the box with no dependencies and install it, get it running very, very quickly and get going and get your problems solved.

    Ashley: I do really appreciate your focus on the developer community. Actually, we—I just did a webinar a few weeks back with Tom Crow, your Community Manager, and he did an excellent job, by the way, of talking about how he worked with the developer community.

    So, let’s jump to InfluxDB Cloud 2.0. Is this your first SaaS online offering, is that correct?

    Hall: It actually isn’t, Mitch, but I’m glad you asked. So, previously, we’ve had InfluxDB Cloud 1.0, which was really taking our enterprise edition and providing a managed service on top.

    And the challenge we had with that is, we’ve seen good adoption and nice traction. We’ve had a large number of customers purchase and buy it, but the challenge is that it was a little on the expensive side, and it existed as a dedicated instance of our enterprise software that we ran and managed for them, but it was constrained on specific resources. And so, you would buy a plan, pay for it, and then if you continued to grow and use it, then you would have to upgrade, typically by contacting support, and we could do that in a matter of seconds, but it really required a lot more high level interactivity than maybe you might get from the vast majority of cloud services that you could think of, more consuming like a utility.

    And so, our focus with InfluxDB Cloud 2.0 was to deliver the first completely serverless time series platforms, deliver that elasticity for our customers to just adopt and go.

    Ashley: So, truly a cloud service, not just an application in the cloud, if you will.

    Hall: That’s right, that’s right. And all the capabilities we’re exposing in a multi-tenanted fashion, which means we’re obviously testing and validating across a large group of folks all at the same time, and keeping that platform up and available for everybody is our goal.

    Ashley: And I know you’re doing more than just storage in the cloud. You have visualization, UI—other capabilities? You wanna talk a little bit about that?

    Hall: Yeah, that’s right. So, you know, there’s sort of three parts to this story, right? Part number one is, you need to be able to feed the data in, and you need to be able to do that with velocity, and we need to store and land that data and we need to do that quickly and make it available for query and access and all the way through to visualization.

    And so, we have all those parts of the story together. The most recent addition also includes the ability to create monitoring checks that allow you to create queries against the data that you’ve stored and decide if they, let’s say, exceed a certain threshold or create a change of state in some way, shape, or form. And then from there, you can create a notification that can be sent to you through a variety of channels including HTTP POST, Slack—which is available in our free tier—and then, obviously, PagerDuty, which is a super popular mechanism for alerting your DevOps teams.

    Ashley: Mm-hmm. And I think, in addition to kind of integration level type functionality, you’ve also updated the API to InfluxDB in the cloud, correct?

    Hall: That’s right. Actually, one of the goals is, you know, through the 1.0 line of the TICK stack, T, I, C, and K, which many DevOps people might be familiar with, there were four parts—Telegraf, InfluxData (the database itself), Chronograf (which is our visualization tool) and Kapacitor (which was our anomaly detection).

    And we’ve really taken the I, C, and K portions of that stack and integrated them all together in our InfluxDB 2.0 offering. This is both open source as well as cloud. And what we wanted to do was bring together the language, first and foremost, that you use to interact with data, both from a query and a task perspective, and we’ve done that through a new language you’ve introduced called Flux—not surprising, the name of the company is Influx. So, Flux is the new language we’ve introduced. And then we’ve obviously kept Telegraf, which is our data collection agent that can be distributed and used to send data in.

    Now, one of the other things about T, I, C, and K is, each one of those elements had its own API. And so, if you worked with one, it didn’t necessarily translate to the other. So, we’ve unified and have created a common API as part of InfluxDB 2.0 that allows you to do everything from drive the creation of dashboard sales and dashboards themselves all the way through to submitting queries, checks, notification rules, et cetera, all behind a common API.

    And that API is consistent across the open source and the cloud editions. That allows [Cross talk]—yeah, sorry, that allows developers to essentially write an application, either against the cloud or against open source. And if they need to swap it around for any reason—for example, if your journey starts with the open source and you’re like, “Okay, now I need the scale and elasticity offered by InfluxDB Cloud 2.0,” they don’t have to make any changes to their application code. They can just change the end point and point out our cloud.

    Ashley: Very interesting. You know, we’re building applications that are API centric much more today versus let’s expose the application through an API. Did that change how you redesigned the API the way you’ve integrated all these components and does just kind of InfluxDB in the cloud more or less use its own APIs because it works together across these functions?

    Hall: It does. We’re actually consumers of our own APIs. And frankly, that’s been our philosophy for a while. The challenge had been previously that it was fragmented across the four components of the TICK stack and we’ve really tried to unify that and think about, again, delighting developers and giving a consistent API experience across them all.

    Ashley: One of the things you mentioned, too, about—and thanks for that information about TICK stack. I’ve got another question about that in a moment. But back to your 1.0 offering versus 2.0—have you changed kinda the pricing model, done anything to make that more accessible to either developers or enterprises

    Hall: Yeah, completely. In InfluxDB Cloud 2.0, it’s a completely usage based pricing model. We’re also starting developers off with a free tier, and so that free tier is rate limited, but you have access to most of the features and capabilities. We’ve obviously held off some of those things behind the paywall, but generally speaking, it’s now all pay as you go and we’ve got that pricing model published and up on the website as well as in the documentation that describes exactly how much you pay for reads, rights, storages, and so on, and you only pay for what you use.

    And that’s really the goal of providing that serverless offering. People wanna use it as a utility and, you know, compare and contrast with what we did with the managed service in the 1.0 line, you know, they were paying for things they weren’t using, right? It was a fixed price offering and so, you know, that’s why it looked a little more expensive. So, I think now, we’re, I think, price optimized for developers to grow as they’re successful.

    Ashley: And one of the things that probably in any developers mind listening to this podcast is thinking about developers definitely don’t like crippleware. So, you’re focused around data retention limitations, rate, that kind of thing. I’m sure you’ve heard that from your developer community about how to properly construct a free tier offering that’s still very usable to them.

    Hall: Yep, yep. I think we’ve got a pretty good mindset around that. On the data retention side, we’re currently allowing the storage of data for at least 3 days at 72 hours. That can be a lot of data in the world of time series. And the ingest rates are also, I would say, fairly reasonable exposed. Just to give you an insight, I actually am using the free tier myself. I haven’t arm wrestled by Product Manager to give me access to all of the features and capabilities, although I suppose that’s possible, but I wanted to see what I could do.

    Ashley: Mm-hmm.

    Hall: And so, one of the things I happened to do with my son recently is, we built a gaming PC for him. And, you know, one of the most expensive components, the two most expensive components in the machine are the graphics card and the CPU.

    Ashley: Mm-hmm, yep. I would’ve said the same thing. [Laughter] Do the same project with my sons, but yes.

    Hall: That’s right. And so, one of the challenges, of course, is that I wanted to make sure that we installed the CPU correctly and I was worried about it overheating. And so, we’ve got a nice, big, you know, heat sink on there. But I enabled the, what’s called the sensor’s plug in within Telegraf to basically give me access to the CPU temperature data. And I’m taking that and I’m actually sending it to Influx Cloud. And then what I’m doing is, I’ve created a check to check on the temperature and I’ve created warning and critical threshold levels, and then I send the notifications to myself in Slack.

    And so, generally speaking, it’s worked quite well, Mitch, but there’s been an interesting side use case and benefit that I’ve found, which is, my son comes home and decides not to do his homework.

    Ashley: [Laughter] Yep. I was just wondering.

    Hall: [Laughter] So, now I have a very uniquely tuned child monitoring service that allows me to recognize when he is gaming when he should be doing his homework. And he still hasn’t figured out how I’m doing it, so I’m definitely not sending him this podcast.

    Ashley: You’re too sneaky, dad. Parental controls and he doesn’t even know it. [Laughter] Too funny.

    Well, I’d like to touch back on the TICK stack, because I know you’ve been very active with providing open source software through the Telegraf and InfluxDB, Chronograf, Kapacitor. What’s the plan, then? You said you’ve kinda done a lot to the open source or to the code and built that into this cloud offering. What’s the plan with the open source going forward?

    Hall: Yeah, thanks for asking. I mean, one of the core values of our company is that we are strong believers in open source. And, as part of that, we’ve been working on the 2.0 offering for more than a year, and it’s gone through a series of alpha editions, and I think we’re currently up to alpha 18, which matches the features and capabilities that we have available as a GA instance for InfluxDB Cloud 2.0. And we may have some community members out there scratching their head, saying, “Um, so, what gives? How come this hasn’t gone GA?”

    And I wanna be really, really clear, because we have such a large community that we value their feedback, but I also value not giving them something that’s broken and something that’s not ready for prime time. And so, we’ve been very careful about labeling the 2.0 code line in our GitHub repository as alpha code as we work to complete the features and capabilities that we want to have within the 2.0 line.

    And so, the key things that are coming up for us next, there are sort of three things. Number one is, we’re moving towards sort of the feature completeness for Flux, our new language. We’re learning a lot of things about running Flux within the cloud environment and getting feedback from community members on that and there’s probably a handful of things that we’d like to finish there before we’re gonna call that sort of release one of Flux.

    Second is, we really need to create migration tooling, and there’s two things that we’re doing in that regard. Obviously, folks that want to land and store their data for longer periods of time will want to move their data from InfluxDB 1 to InfluxDB 2. And we’ve created sort of a new model inside of InfluxDB 2, which includes the notion of a bucket, which is where you place your data, and how to secure that bucket. And so, there are some migration tools that are required to bulk move that data across.

    And then last is, I mentioned, we’ve introduced a new primary language for working with data inside of InfluxDB 2, and it unifies the ability to both query that data for dashboard creation and report generation and these other use cases that people might build on top, but it also is the same language that’s used for tasks and sort of creating batch operations.

    And that was certainly not the case between InfluxDB and Kapacitor, let’s say. Kapacitor used something called TICK Script and Influx was using InfluxQL. And InfluxQL is a SQL like language and it was an easy onramp for gaining access to the data within InfluxDB. And we’re certainly not abandoning that, but what we’re doing is we’re going to create the ability for you to use InfluxQL, but behind the scenes and on the fly, we will transpile that into something that’s executable by the Flux engine, and then return the results to the customer. And this should allow for existing dashboards and queries just to work natively and seamlessly with InfluxDB 2.0 and obviously InfluxDB Cloud 2.0 when we offer it there. And when we reach that point, then we will declare beta, which means we’ll have feature completeness which include those migration tools and then that’s the time for the community really to start piling in and trying it out, and then we’ll be focusing more on performance and sort of user experience changes between the beta declaration and when we declare 2.0 GA.

    But in the meantime, we’ll continue to add features to the cloud edition, they will show up in the open source through our alpha cadence. But I would say that those three big things, you know, the migration tooling, InfluxQL support, and just completing Flux as a generally available language in version one is where we’re focused.

    Ashley: Okay, very good. Well, certainly a well thought out strategy and you’ve done a lot of work that we’d look forward to seeing in the open source path as well.

    You mentioned buckets, remind me to ask—what types of clouds are you offering the Cloud 2.0 service in?

    Hall: Yeah, great question, Mitch. So, currently, we’re available in AWS and there’s two regions we’re available in—U.S. West 2, which is up in Oregon, and we just started the beta process in Europe, which is available in Frankfort, and it’s currently using the same code, but we just haven’t turned on the ability for folks in Europe to pay us yet. But we’ll run that beta probably for four to six weeks in Europe before announcing GA there.

    In addition, we announced at Google Next earlier this year that we would be available on GCP. And so, our intent is to continue to drive forward. Most likely, that will appear in the U.S. East region, just to sort of fill out a geographic strategy, but we’re very excited about working with the team at Google. Some of my former colleagues from Oracle are now there, which is fun to reconnect with them.

    Ashley: Mm-hmm.

    Hall: And we’ll see that land hopefully by year end, and then we’ve been spending time with Azure on the Microsoft team, and so, I would look for that in Europe and Cuba.

    Ashley: Okay. Very good. I’m really tying back to the very beginning when you told us about your background—I’m really interested in what do you see are the differences in the approach with InfluxData takes with the developer community being such a focus for the company versus a more traditional database company like you worked at at Oracle? Not that they don’t have developer products, they do, but clearly, this is the center of the universe for you. What are some of the things that you’ve learned or maybe differences in the approach you’re taking now?

    Hall: I’d say it’s a mix, Mitch. Like, there are definitely some things in the Oracle landscape that were powerful and can apply to the open source community.

    So, for example, feature flag items—so, don’t ship a new version of the database with all your latest and greatest and coolest features turned on by default. That really prevents people from moving rapidly from, you know, from older editions to newer editions. So, we try to keep those things turned off and provide highlights to folks and release notes about how to activate them and turn them on. And it is a double-edged sword, because people who are new that come in, you’d like them to use the latest and greatest and best technology and we do try to highlight that. But we also want people to move and eliminate that sort of long tail of support as quickly as possible.

    So, a very smooth process. And this sort of came from Oracle—yeah, feature flag those new things, definitely highlight it so they can take advantage. But let the developers—they’re smart. They’ll take advantage of those new features if you just tell them where they live and how to turn them on.

    Ashley: Mm-hmm.

    Hall: On the flip side, on the community side, you know, listen to the community feedback. And I would say that’s one thing that’s been top of mind, if not totally obvious is, you know, the motivation behind introducing Flux as a new query language was not done lightly. But it really is in response to all of the outstanding query related questions and issues that were opened by our community members. They could not move certain workloads, they couldn’t do certain kinds of functions, they needed more flexibility with the language. And so, Flux is the genesis of listening to all that feedback.

    You know, unfortunately, some of those issues have definitely stood out in the community for, you know, now going on two or three years. In some cases, I finally feel like we’re at the point where we can deliver on those requests from our community members.

    Ashley: Very good. I feel like we’ve covered a lot of ground on our podcast today, certainly giving our listeners a great feel for what’s unique and new and also how to get started with Cloud 2.0 InfluxDB. I imagine on your website, there’s a pretty easy place to get started?

    Hall: Yeah, absolutely. I think it’s the primary call to action on the website today. If you go to InfluxData.com, it’ll pop up and say, “Get started with Cloud 2.0 immediately.” Also, you can get the open source by going to our Downloads page, or you can find us on GitHub.

    Ashley: Excellent. Well, Tom—Tim, [Laughter] I knew I would do that—Tim, thank you for being on the podcast today.

    Hall: [Laughter] You’re welcome, Mitch.

    Ashley: Okay. I owe you one. [Laughter] I’d like to thank my guest, Tim Hall, Vice President, Products, at InfluxData for joining us today—and, of course, thank you, our listeners, for joining us as well. This is Mitch Ashley with staging-devopsy.kinsta.cloud. You’ve listened to another DevOps Chat podcast. Be careful out there.

    — Mitchell Ashley

  • DevOps Chats: Robotic Process Automation with Tricentis

    DevOps Chats: Robotic Process Automation with Tricentis

    Robotic Process Automation (RPA) has garnered a great deal of attention as organizations pursue automation routine tasks through automation software. RPA automates tasks integrated into digital business process automation.

    Wayne Ariola, Tricentis general manager of RPA, joins DevOps Chats to discuss how Tricentis brings its strengths and heritage in software testing automation into the world of RPA. The expansion into RPA makes a lot of sense when you consider testing technologies that operate on UX domain and data models, avoiding the pitfalls of brittle screen scraping approaches.

    Together, Wayne and I explore how business process automation using RPA intersects with DevOps tools, processes and teams, and how RPA benefits from the experiences gained in the DevOps community. Join us on this episode of DevOps Chats as we explore RPA and DevOps.

    As usual, the streaming audio is immediately below, followed by the transcript of our conversation.

    Transcript

    Mitch Ashley: Hi, everyone, this is Mitch Ashley with staging-devopsy.kinsta.cloud, and you’re listening to another DevOps Chat podcast. Today, I’m joined by Wayne Ariola, who is general manager of RPA at Tricentis. Our topic today is RPA, robotic process automation, and DevOps. Wayne, welcome to DevOps Chat.

    Wayne Ariola: Thank you so much for having me on this beautiful Friday. I’ve been looking forward to it.

    Ashley: Best way to spend a Friday—well, second best, maybe. There might be a few other things—

    Ariola: Exactly.

    Ashley:– that might be above the list of this, [Laughter] but yeah, what a great day to do that. And I know we’ve had Tricentis, your CEO has been on before. Welcome back. Glad to have you on. Would you introduce yourself to our audience—

    Ariola: Sure, absolutely.

    Ashley: Tell them about you and, for those that don’t know, tell them what Tricentis does.

    Ariola: Absolutely. So, my name is Wayne Ariola, as you stated, Mitch, and I’m the general manager of our RPA solutions at Tricentis. And I’ve done—I think I’ve done at least three of four webinars on staging-devopsy.kinsta.cloud.

    Ashley: Oh, yes.

    Ariola: So, I’m certainly no stranger to the audience, the topics, nor the concerns, yet I wanted to actually bring forth the news about RPA and see if I can’t connect it to what’s going on in the end to end DevOps world today.

    Yet, just a little bit about Tricentis, for folks who haven’t heard about Tricentis—you know, Tricentis is, today, considered one of the automation leaders out there in the world. You know, consistently across all major analysts who are always ranked as either (a) the number one leader, or, you know, right up there as number one. So, it’s been an interesting ride at Tricentis for the last five years as we’ve begun to kinda notch our way into the DevOps world.

    The one thing we are known for us really making software testing more productive. And we do that via automation. So, we know as we look at scaled agile that things need to happen quicker and then you add in the idea of CI/CD or these kinds of concepts, and then you really wanna be fast and agile. And we’re the company that makes sure that testing is not your barrier to speed. And that’s what we’re best known for out there.

    Ashley: Well, let’s start out by, I know RPA, robotics process automation, is kinda—is the hot thing or the new thing, however you wanna put it. How do you define RPA?

    Ariola: So, I mean, poor DevOps people, right?

    Ashley: [Laughter]

    Ariola: I mean, DevOps was the sexy thing for such a long time.

    Ashley: Oh, it still is. No, don’t break my heart, Wayne—come on, no.

    Ariola: Good, good. You know, you guys haven’t aged out yet, so.

    Ashley: No.

    Ariola: And all of a sudden, you know, RPA is like the Millennial of topics, right, these days.

    Ashley: Okay. [Laughter]

    Ariola: It is definitely on fire. In fact, I have a slide that I use in presentations that I compared DevOps to RPA. And the DevOps folks, it looks like there’s a bunch of people dancing, but then I show a picture of RPA and it looks like a rave—you know, a big, huge party at a rave with lasers and lights and a whole crowd thumping to the music.

    And that’s really what’s happening right now. I mean, the RPA world and the RPA market is just on fire, and it’s a bit crazy. But the reason why it’s a bit crazy is that the ROI associated with deploying automation or what is defined as RPA automation is really high. You get immediate value out of it, so it’s kind of like this—especially if you talk to CIOs, you know, they’re kinda saying that it’s kinda the no-brainer of today’s world.

    But this is, I mean, the DevOps folks are no stranger to this, right? So, we did, we went through this when we kinda went through the whole DevOps motion beginning with build automation, right?

    Ashley: Right, yeah.

    Ariola: You know, we kinda took a breath and we said, “Hey, this build process is manually painful and it’s difficult and it’s arduous.” And, you know, once we got through kinda the initial thing around the automating build, you know, the imagination associated with what could be automated then exploded, you know, and then you got into the DevOps motion, right?

    And I would say it’s a perfect parallel. RPA is doing exactly the same thing. It’s kinda taking these preliminary type of motions associated with manual processes and the automating them.

    Ashley: Well, you know, in a way, the way I think about it, Wayne, is, you know, to step back for a minute, the best developers I’ve worked with in my career have always been the laziest, and that’s a compliment.

    Ariola: [Laughter]

    Ashley: Meaning, they write the least amount of code, the automate everything—they don’t wanna do that manual stuff, right?

    Ariola: Yes. So true.

    Ashley: They wanna spend time on the things they enjoy working on. So, you know, in the DevOps world, of course, we’re big on automating things and maybe all of it isn’t automated, but it’s taking that same kind of, “Let’s take those things we don’t need to do manually and we really can automate it and look across IT, across business process automation—you know, the whole organization.”

    Ariola: Yeah.

    Ashley: That’s probably why it’s hot, because it’s not only in IT, but it can affect other parts of the business. So, I’m gonna lay claim to RPA’s roots being in DevOps. Now, you can tell me I’m totally wrong, but I won’t—

    Ariola: No, absolutely not, given the audience I’m speaking to. I agree.

    Ashley: [Laughter]

    Ariola: I totally agree. But it’s true, though—I mean, think about this, Mitch. Let’s unpack that just for two seconds, right? So, take the idea of that developer who’s lazy, right, and the fact that—what are you doing? And if you would, say, ask the person what they’re doing at any point in their day, it’s like, “You know what, I’m writing a script to do X, Y, Z because I’m so sick of provisioning this particular application to this environment,” whatever it might be, right?

    So, their knowledge base and their understanding of how to construct a script gave, especially in the DevOps world or the Dev world, these guys a leg up in terms of producing this automation to eliminate manual tasks.

    Ashley: They’re the most productive, yeah.

    Ariola: Yeah, exactly, and make that the most productive. But RPA is really trying to—and I hate to even use this word, but it’s the best way to describe it succinctly, is democratizing that, right? It’s giving people who are not as talented as your laziest developer the ability to actually create the automation to get rid of the mundane or manual tasks that are slowing them down, right?

    Ashley: Mm-hmm.

    Ariola: So, I mean, even by—Mitch, even by definition, you know, RPA is defined as a technology that predominantly leverages a combination of maybe UI and API or surface level features and applications to create more automated, routine, predictable data transcription or enablement work, right? So, automation work or executable work.

    So, if you think about it, it’s a technology that sits on top of—a super structure that sits on top of apps that allows you to do things more proactively across applications. So, making sure that a manual task to transpose data, to move data, to initiate something, to do data entry is eliminated by the automation itself, which is really cool.

    Ashley: You know, I kinda think of it as—I think of it as, an analogy would be just like sysadmins write scripts, right, to do things that automate it.

    Ariola: Yes, yes.

    Ashley: This is for other IT folks, but also can be end users, right, that automate filling out forms on online screens or doing some processing of “load this data into a spreadsheet, export it out to here, send it to here.”

    Ariola: Yep, yep.

    Ashley: Now, why do all that stuff by hand? Here’s an easy tool. Now, do you see RPA as primarily being rule based, or—I know you have an orchestration product. I don’t know if that’s primarily rule based, but how do you do RPA in a way that it’s accessible by as many people as possible?

    Ariola: Yeah. So, I mean, also knowing the demographic of your audience, RPA certainly falls into the bucket of, nothing new has be invented here, right?

    Ashley: Mm-hmm.

    Ariola: This is [Laughter]—this is definitely technology that we’ve had awareness of in the past. But it’s been uplifted for our era, right? Is really what makes the difference, right?

    So, you know, whether it’s rule based, whether it’s AI, whether it’s—it really depends on the task you’re trying to achieve, because not everything fits every scenario, right?

    Ashley: Mm-hmm.

    Ariola: So, if you—let’s try it this way, though. If you look at RPA and you take, let’s just take roughly 75% of the use cases out there, it falls into one of three buckets.

    Ashley: Okay.

    Ariola: One bucket is what I call poor man’s integration. Which is, you’re scraping information out of one system and you’re putting it into another system and then validating it in another or something like that, right?

    Ashley: Okay.

    Ariola: Meaning that you’ve gone to your IT group and you’ve said, “Hey, I got this integration project that I need” and they say, “I don’t have enough time.” And then you’re saying, “You know what? I’m just gonna do it myself with RPA” and you’re just gonna do it by scraping screens and moving and pushing data across multiple systems. Or maybe even—

    Ashley: “Here’s some vise grips and a crescent wrench, go fix it up.”

    Ariola: Exactly! Yeah, exactly.

    Ashley: Okay, yeah.

    Ariola: By the way, I call this poor man’s integration, or these kinda RPA use cases the best identification of backlog that you could ever have, right? You know there’s a problem here when you use it for this problem.

    Ashley: Okay.

    Ariola: So, the second group of activity that has been used is the augmentation or automation of manual tasks.

    Ashley: Okay.

    Ariola: And this is what you primarily hear about, right? So, instead of having me do the mundane task of pointing and clicking, going to another system, copying a number, bringing it over, taking—finding the core account number from a subsystem and pulling it into the work order, you know, pulling information from SAP and pulling it over to SFDC. You know, all that kinda manual stuff that’s going on in between. This is the second scenario that you’re just helping the human do something as part of their day to day activity, whether they’re processing an invoice or monitoring inventory or onboarding an employee or, you know, any kind of horizontal process that’s happening in the organization.

    And then there’s the third scenario that it’s being used for, really, which is—validate critical checks. So, when I say validate critical checks, let’s take the scenario, like, that we all got indoctrinated to in the last two years, which is GDPR.

    Ashley: Mm-hmm.

    Ariola: Where, if someone opts out from an e-mail perspective of your system, you want to make sure that they’re opted out of all systems of record. So, you might want to build a queue of opt-outs, which, the RPA engine comes in, grabs a unique identifier such as an e-mail from that system, and then goes into all other systems or potential systems of record and making sure or giving us, you know, taking a visual snapshot that the individual has, in fact, been opted out.

    So, those are kinda like the three general cases. And, of course, we then transist multiple horizontal, vertical, and technical domains for those particular uses cases.

    Ashley: Well, I thought with that third example, you were gonna tell me that I wouldn’t have to click on that. I accept cookies any more on all these websites, but—

    Ariola: [Laughter] You can set up an RPA [Cross talk] for that, no problem. You could do it.

    Ashley: [Laughter] You could, actually.

    Ariola: You could definitely do it—yeah, yeah, yeah. [Cross talk]

    Ashley: So, talk to us about what would you say are the top one or two use cases that customers come to you and say, “Oh, you do RPA? I’m trying to solve this situation”—what is it?

    Ariola: You know, it’s really funny, because it’s really across the board. And the industry is starting to sub-segment themselves now into vertical and horizontal type processes or technical—so, there’s three ways that the whole industry will ultimately segment. You’re gonna have pure plays, right, that are gonna be cross industry, cross horizontals. You’re gonna have some horizontal folks who are gonna go across—like, so, for example, we’re seeing new companies pop up that do just kind of HR RPA. Meaning, getting information in and out of HR systems, sensitive data—blah blah blah. And then you’re getting some vertical stuff. Like, you know, you have folks like, you know, Statenical and Edgver, which are Mphasis companies that are in the financial vertical that are really, really good at doing some stuff in the financial vertical.

    Ashley: Mm-hmm.

    Ariola: But in terms of us, you know, there’s really only—I would say there’s only about four or, three or four sure plays out there. Tricentis RPA is certainly one of those pure plays, and you know, we see a lot of insurance scenarios where you’re doing payouts, payment validation. We see a lot of things like supply chain data verifications a ton, a ton of that stuff.

    We’re seeing a lot of IoT type scenarios where a sensor is collecting information from one particular device and it needs to be married up with other IoT type of devices out there collecting information, and via an API, it’s being all sucked together and orchestrated together with RPA in order to produce an outcome, right?

    Ashley: Mm-hmm.

    Ariola: We’re seeing a lot of things around tax and payroll type steps and validation. A lot of stuff going on around SAP as well, by the way, for some reason.

    Ashley: Mm-hmm.

    Ariola: So, getting stuff in and out of SAP. And this is, there’s two major things going on, here. First of all, from a merger/acquisition perspective. So, in order to gain the short term benefits of any M&A activity without necessarily having to invest in the entire integration plan, you know, you can use RPA to assist you in moving data or updating records or updating customer numbers, right?

    Ashley: Mm-hmm.

    Ariola: And then you’re also seeing scenarios like it being used for S/4HANA migration, which is also really interesting. So, those are kinda the things you’re seeing, but it’s all over the place.

    Ashley: Yeah. Yeah, you know, I love the IoT example, because—I know it’s not your technology, but everybody knows “if this, then that,” right? That’s sort of the universal—

    Ariola: Yes.

    Ashley: Doesn’t work for every scenario, but it certainly is a great example of a really simple way to do RPA for IoT events and happenings. Talk a little bit about why did make sense for Tricentis to expand from automating testing to more generally providing an RPA solution.

    Ariola: Yeah. So, thank you for asking that, because I do happen to get that question quite often. So, at the core of—Tricentis’, you know, a real differentiator in the marketplace is this technology we have called model based automation.

    Ashley: Mm-hmm.

    Ariola: And most of what you’re gonna see today out there from RPA vendors are scripted technologies or technologies that use image based controls to understand screens. And the script based technologies and this image stuff, by the way, is the prime reason why software test automation has been so difficult. Because scripts fail.

    Ashley: Yeah, I was just gonna say. [Laughter]

    Ariola: Yeah, they’re brittle. They basically are hard to update. It’s hard to maintain. You know, it’s the old adage of keeping in sync with the code base and, you know, RPA is highly susceptible to changing UIs. And this is where we shine.

    So, our model based automation is actually, takes a much, much more technical approach. So, where most of the vendors out there, if not all of them, interrogate a UI from a screen perspective, we actually interrogate the implementation.

    Ashley: Mm-hmm.

    Ariola: And understand the UI from its more technical implementation perspective. So, if you take SFDC or if you take SAP or you take ServiceNow or any of these big, large vendors, the complexity of the UI even through a browser is pretty significant. So, the fact that we actually interrogate it at a technical level to build an abstracted model gives us the ability to deliver real, resilient automation.

    And what I mean by that is that, you know, we—universally, in our platform, you cannot script. If you wanted to script, you’re gonna have to go somewhere else, because you can’t script, because ours is model based. So, everything that we’re doing is actually produced from technical observations of the application itself. And then we allow you to put those observations or components that you discover—reuse them as LEGO blocks, right, and put them into different flows.

    So, the good thing is that, when the UIs change that are part of the automation, (a) if you have to make a technical change, you only make it in one spot, and that change propagates throughout all instances or all bots that are using that automation, and (2) because we actually do a highly resilient technical implementation, you know, we self-heal our technology or our bots in roughly 80 percent of the change cases.

    So, this is why were successful and still are successful in software testing, and this is why we decided to bring this to RPA, because it faces the exact same challenges.

    Ashley: Well, it’s almost a knowledge based approach, thinking about it as a model oriented solution, if you hearken back to screen scraping days, which is the brute force—

    Ariola: Yes.

    Ashley:—this vector is where the information is at that you’re testing—

    Ariola: Exactly.

    Ashley:—to a model based, which is more of a canonical based approach, which is a defined—

    Ariola: Absolutely.

    Ashley:—you have a way of determining what the information is and then you can detect when it’s changed. “Hey, there’s a new field. There’s a new—this field’s changed,” you know, whatever, it’s not in place.

    Ariola: Yes.

    Ashley: Now you have some reference model to pull from. It makes a lot more sense now. You can put logic based or flow based logic into an automation process rather than having to script it into everything as manual. I think that’s what you’re saying.

    Ariola: Absolutely.

    Ashley: Do I have it right, here?

    Ariola: You have it 100 percent correct. So, it also gives us a broader reach within the organization as well. So, you don’t have to rely on the highly technical skills to actually produce the automation or, even more importantly, maintain the automation that’s required by your business.

    Ashley: So, does this open up different people that you’re selling to in the organization? Are you still selling through the DevOps test organization in IT and you kind of expand what other places you can apply your RPA technology? Are you now calling on business units, end users, other application areas? What’s that mean for your business?

    Ariola: So, whoever has the problem, we certainly want to talk to about assisting them to solve it. You know, within the DevOps space, there’s a lot of conversation about maintaining the scripts that are automating everything around my infrastructure is painful, right? So, there’s conversations that are even in that horizontal area. However, you know, the skill set that’s mostly in that domain is highly technical, so, you know, there’s a lot of preferred—a lot of preference to actual scripting there. But within the line of business, there’s certainly a lot of conversation going on, and that broadens the conversation.

    But believe it or not, you know, when you’re talking about tests and test automation, although we do think of it as kind of a Dev test in motion, when you look at what Tricentis does, which is more end to end testing and protecting the end user journey because we cross applications, we have a lot of conversations with the business analysts. We have a lot of conversations with the line of business.

    Ashley: Mm-hmm. I can see that.

    Ariola: RPA tends to skew that more to the line of business, but what we’re noticing distinctly these days is, RPA is coming back to IT. And there’s a real, really interesting reason for that, and it’s what’s called the RPA death spiral. Which, it goes something like this. You asked IT to do an integration for you, but they were too busy. You decided, “Hey, by the way, I’ll use this RPA thing to bridge the gap right now” and you called in an RPA vendor and a service partner and you got your initial flow or bot stood up.

    Ashley: Mm-hmm.

    Ariola: But then again, you just don’t wanna—you don’t wanna spend a lot of money keeping that service partner on site because it’s expensive, and they go away. And once they go away—guess what? The interface changes on one of the automation sequences or bots that you have enabled and the bot breaks. And as soon as that bot breaks, [Laughter] where are you gonna go to? Well, you’re gonna go to IT.

    Ashley:Well, I can imagine.

    Ariola: And then IT—yeah.

    Ashley: Yeah, people come back to IT because the problem that they’re automating gets bigger than they want to handle or can handle.

    Ariola: Absolutely.

    Ashley: Or they need access to resources—“Hey, I need single sign on here to get to this information. I need this database that isn’t—

    Ariola: Absolutely.

    Ashley:—easy for me to get to as an end user.” So, lots of reasons.

    Ariola: Absolutely.

    Ashley: You know, things happen outside of IT because they’re easier, but then we need to go back to IT because they’re necessary to get what you need.

    Ariola: Absolutely. That’s why it’s kinda swinging back to IT. Now, also what’s happening is, when the line of business acted kind of like a shadow IT or kind of a rogue IT shop and adopted the technology, you know, now the CIO is essentially inheriting the maintenance tab.

    Ashley: Mm-hmm.

    Ariola: And now this is why, you know, the conversation is swinging right back to IT and the CIO.

    Ashley: I can see that. Yep.

    Ariola: And we’re back—from Tricentis’ perspective, you know, we’re back on home turf talking to the nerds that we like to talk to.

    Ashley: Mm-hmm. Well, this has been great. I really appreciated having you on, and a new topic for us on DevOps Chat with Tricentis. I don’t believe we’ve talked with you about this before. It sounds like you’re a big enough believer, you’d change your job to be the general manager of RPA.

    Ariola: [Laughter] Yeah.

    Ashley: I can tell you’re committed.

    Ariola: There’s no goin’ back—there’s no goin’ back now, yeah.

    Ashley: Exactly.

    Ariola: I’m committed.

    Ashley: Exactly. Well, it’s been fantastic, Wayne. Thanks for being on the podcast.

    Ariola: Absolutely my pleasure, Mitch. Thanks for having us.

    Ashley: Well, I hope you come back and we get a chance to talk to you again and maybe we can dive into some of the use cases you’ve solved for customers and some of the learning, so we’ll save that for another time, okay?

    Ariola: Love to do it.

    Ashley: Great. Well, you’ve listened to another DevOps Chat podcast. I want to thank my guest today, Wayne Ariola, who’s general manager of RPA at Tricentis. And, of course, thank you to our audience for spending your time with us today—it’s valuable, and we appreciate it. My name is Mitch Ashley from staging-devopsy.kinsta.cloud and you’ve listened to another DevOps Chat. Have a good day and be careful out there.

    — Mitchell Ashley

  • DevOps Chats: Demystifying Spinnaker VM Baking, with Salesforce

    DevOps Chats: Demystifying Spinnaker VM Baking, with Salesforce

    Spinnaker Summit 2019 Preview: As you use an open source tool like Spinnaker, the more you develop best practices and learnings. These are useful to share with others internal to the enterprise and other open source users.

    Jing Vergara, principal software engineer at Salesforce, joins DevOps Chats to share a preview of her upcoming talk “Demystifying Spinnaker VM Image Baking and Deployment.” Jing shares with us what to include, and not to include, when baking a VM image, how to avoid configuration drift and how to deploy to multiple cloud providers using Cloud-init files. Jing also shares how to use open source Packer templates with Ansible, Chef and Docker, and how to build your own Packer templates.

    Jing Vergara’s talk is on Sunday, Nov. 17th at 10:45 AM PT. Spinnaker Summit 2019 is San Diego on Nov. 15-17.

    Transcript

    Mitch Ashley: Hi, everyone. This is Mitch Ashley with staging-devopsy.kinsta.cloud and you’re listening to another DevOps Chat podcast. Today, I’m joined by Jing Vergara, who is principal software engineer at Salesforce. Now, Jing is going to be talking at the Spinnaker Summit in San Diego in November, and she’s gonna be talking about a really interesting topic, “Demystifying Spinnaker VM Image Baking and Deployment.” Now, don’t think we’ve changed into a baking show. We’re baking VM images, and that’s what Jing’s gonna tell us a little bit about, that whole process—pipelining and creating VMs and baking them. [Laughter] So, her talk is on Sunday, November 17th, at 10:45 a.m.

    Jing, welcome to DevOps Chat.

    Jing Vergara: Thanks, Mitch.

    Ashley: Great to have you.

    Vergara: Thank you for the opportunity to have me talk here.

    Ashley: Well, I’m honored you’re on. Thank you for being here. Tell us about yourself. Tell us a little bit about what you do and maybe what parts of Salesforce product or code that you work on.

    Vergara: Sure. So, I’m part of the continuous deployment team at the company. We offer Spinnaker as a Service for all of the organizations within the company. And I help teams onboard their delivery pipelines into Spinnaker.

    Ashley: Oh, interesting.

    Vergara: Yep, and as part of my job, I also build custom extensions to Spinnaker to support their internal business needs.

    Ashley: So, you’re kind of the Spinnaker service center expert person or group that everybody goes to for help implementing Spinnaker. Is that close?

    Vergara: Yeah—very, very close, yes. So, my role within the team has been changing, but yeah, it has transformed into that one. Another thing that I do for Salesforce is that I represent Salesforce at a couple of the Spinnaker special interest groups for stakes insured so that I could learn more about upcoming features of Spinnaker, so we can leverage those in our roadmap as well so that we know we can offer it to the business.

    Ashley: Mm-hmm.

    Vergara: And then we also contribute back to the ________ and let them know what would and would not work for us as users of Spinnaker.

    Ashley: Well, and I noticed in your bio on the Spinnaker site, you have a Master’s in Computer Science from the University of Illinois at Urbana-Champaign. I used to live not too far, from there, though I didn’t go to that school. [Laughter]

    Vergara: [Laughter]

    Ashley: That’s a very strong C.S. degree, and it sounds like you’re putting it to good use with Spinnaker and DevOps style deployment with these kinda pipeline creation.

    Vergara: Yes, right.

    Ashley: Excellent. Well, tell us a little bit about what—so, we’re talking about your talk. So, describe for us a little bit about what are some of the things that are maybe a little mystical that need demystifying about Spinnaker VM image baking and deployment. What are some of the things people don’t know that you usually have to help them understand with, “Here’s how it works and here’s how we do it”?

    Vergara: Sure. So, for companies that are traditionally in a bare metal world where you don’t use VM images, it’s kind of a definite shift in the mindset, right?

    Ashley: Mm-hmm.

    Vergara: So, the way I usually explain it to them is using the baking metaphor, right? So, when you talk about baking, you think of putting the ingredients together like flour, salt, what have you, right? You form a dough, shape the dough, then you cook it. You start with a base ingredient—like, again, the flour—and you mix it up.

    So, in the VM world, when you bake an image, you start with that base OS image. You add your other ingredients, which means installing your applications or your agents to run on that VM. And then you have your standard or common configurations and possibly some data, right?

    Ashley: Mm-hmm.

    Vergara: You add them all up in the VM, you create a snapshot. So, that snapshot, you can now call it the VM image. So you can save this image into a store, and once you have that image, you can use it for deployment. So, deployment, it’s basically building or provisioning one or multiple VMs out of that same image.

    Ashley: Okay. Now, leading into this, the thing that kinda kicks it off is that, I’m not sure how it works at Salesforce, but is it something like a Jenkins process that then kicks off this creating of the image? How does that work?

    Vergara: Right. So, we have our continuous integration system. So, we have a few and we’re standardizing that at Salesforce as well, right? So, we have those artifacts built by those CI systems. We have our Docker images so they get stored, we put them into the cloud, and then we have triggers. So, we listen for those changes. So, Spinnaker supports a Pub/Sub subscription that triggers the pipeline. So, then, once we have those architects, we can start the baking process.

    Ashley: Great. Excellent. Well, that makes a lot of sense. Think of it as all of the pieces of software, OS on up, that are ingredients, right? Just elements of what’s gonna go into that image.

    Now, I imagine this is something that’s happening frequently if you’re doing a lot of deploys, right? Is this happening every time that you do a deploy or just as part of the integration process, and then it may go through a deployment process, too, or how frequently would you see this running?

    Vergara: Yeah, so it’s for every code change that gets checked in, right? Ideally, you would want that change to go all the way, so for continuous—again, for continuous delivery and deployment, that’s the ideal process. Check in, generate the artifacts, and deploy that artifact, create a baked image out of it, and then go and deploy it.

    Ashley: Okay, great. So, this is—

    Vergara: Right, so, but—yeah, yeah, right. But like, when we say delivery and deployment, of course, you know, there’s release management and change management policies at every company that says, “Okay, this is actually good for production deployment,” right? So, some of them are continuous deployment up to development or test environments, but for production, we still want to have that gate system in there.

    Ashley: Mm-hmm, yeah.

    Vergara: And Spinnaker actually supports it really well, you know? It’s got those manual judgment stages that you can add into the pipelines to get those kinds of approvals.

    Ashley: Yeah. Well, we’ve had some other folks talking about some of those stages like Run Job and Webhook and some of the capabilities that are built in.

    Vergara: Yes.

    Ashley: Easy to extend as well in Spinnaker. So, I think one of the things that you’re planning on talking about in your talk at Spinnaker Summit is maybe some lessons learned or best practices that you’ve figured out from doing this a lot at Salesforce. Any thoughts on that you want to share with us about the kind of best practices you’re gonna get into in your talk?

    Vergara: Sure, yeah. So, some of the best practices that I’d like to share with you is more around the baking and deployment. So, for example, when you do baking and never, ever bake secrets during bake time. [Laughter]

    Ashley: Yeah. No tokens, no keys, no certificates, no—

    Vergara: Exactly. Yes, yes. So, that’s very, very important. So, those kinds of sensitive data, you can actually perform that during the deployment stage. There is a trap stage there during deployment where you can go and, you know, connect to a secret store, like, Vault, HashiCorp’s Vault, right? So, you can pull down those secrets and install it on the resources or the VMs, right? So, that’s the safer way to go about it rather than baking it during bake time.

    And then another thing as well, you know, as much as possible to avoid configuration drift, right? So, if you have any common or standard configuration, be it your application configuration or your OS configuration, right, do it during the baking step, right? So, as much as possible, limit those VM configurations towards them. Again, you can do it during deployment.

    Ashley: Mm-hmm.

    Vergara: And then, for companies that support multiple cloud providers—so, for example, you need to go and deploy to, say, Amazon Cloud or Google Cloud or other clouds that are available and supported by Spinnaker, you can use Cloud-init files so that it’s not tied to one particular provider and you don’t have to repeat your bootstrap steps during that same time. So, you write your script once, your Cloud-init file, and then you can use it in those different cloud providers, so that’s, I think, that’s very, very helpful.

    Ashley: Now, is that capability something that’s built into Spinnaker if you write your script corrects, and I think you said init file for each—

    Vergara: Cloud-init.

    Ashley:—each cloud, Cloud-init.

    Vergara: Yes.

    Ashley: Is that something that, it’s a feature if you know about, you can use it up front and design that into how you’re writing your scripts?

    Vergara: Yeah, so, there’s—there are fields within the Spinnaker deploy stage where you can put that in. So, in, I think in GCEVMs, you can use the user data field in the deploy stage, and then in AWS, there’s the metadata field.

    So, the field kind of differs by provider, but it’s the same concept, right?

    Ashley: Mm-hmm.

    Vergara: So, you put it in there, you plug it in. And then, during the bootstrap process of those VMs, I think the OS has the support Cloud-init, though, so when they choose an OS, make sure it’s got the Cloud-init, and then Cloud-init will pick up those files and then launch the steps or commands from what you have in those Cloud-init scripts.

    Ashley: Mm-hmm, great. What are some of the things that people that are maybe new to Spinnaker or new to creating VM images and doing baking, what are some of the typical struggles someone might have that would be helpful to come to your talk for?

    Vergara: Right, so I experienced the same thing when I started out with Spinnaker. I’m like, “Oh, this is all new to me! How do I go about this?” [Laughter] So, one of the things that I’m gonna share in the Summit is, like, “Hey, you know what? Spinnaker actually uses Packer, it’s a HashiCorp tool that actually performs the baking behind the scenes.

    Now, Spinnaker offers a set of Packer templates, right? So, those are the very basic ones. And it actually varies per provider. So, you can only use one at a time. So, if you’re building for, say, GCP, then use the GCP Packer template. And if you’re using AWS, you have to use a different AWS Packer template.

    Ashley: Mm-hmm.

    Vergara: So, and then, if you have a more advanced use case, you can actually customize and build your own Packer template, so that’s one of the powerful teachers of Spinnaker as well, right? So, if you end up—like, you know, so for in-house, you use Chef or Ansible or Puppet. There are Packer templates that can use those kinds of provisioners, but you have to go and deploy those into Spinnaker.

    So, at the talk, I’m gonna share how you can deploy those custom Packer templates. So, there are a couple of ways on how to do it. So, one is the simpler way, but it’s kind of limited. Those Packer templates are stored like Kubernetes secrets in Spinnaker, so it’s got a limitation in size. So, if you have complex scripts or Packer templates, then you have to deploy it as a sidecar container.

    Again, I’m gonna share those tips at the Summit and give more details around it.

    Ashley: Mm-hmm. Yeah, folks may be familiar with Packer, too, it’s an open source project in and of itself for creating images, machine images.

    Vergara: Yes.

    Ashley: Great. Are there some other struggles that you find folks typically have, or may have when they’re new to this?

    Vergara: What else? So, Spinnaker by itself mostly focuses on releasing software changes, but it has a very flexible pipeline management system. Like, at Salesforce, we actually use it for building resources such as load balancers, DNS records, topics, queues. You can actually create custom job stages to perform those.

    Ashley: Mm-hmm.

    Vergara: So, we found it really awesome, because we can rebuild a full environment carrying down all of those existing resources, running the same set of pipelines that we can deploy together with those infrastructure resource setup. So, we can build it up right away, right, and then it also answers our scaling requirements. Because, you know, large scale enterprises such as Salesforce, you know, we have to build all of these resources and deployments. We’re talking about thousands of servers across different regions and geographical locations.

    Ashley: Mm-hmm. You know, sometimes tools, open source, or even commercial tools that are highly flexible and very powerful, sometimes with that comes complexity, so it helps to have some guidance, some experience mentoring from someone like yourself that’s gone through that process and can help you kinda get through it and understand how to use it, how to get things done, right?

    Vergara: Right. So, Spinnaker has a Slack channel. So, you know, when people like me or others who are new to Spinnaker have any questions, we go and ask in that channel. There’s a few channels there that are specific to certain features. So, for example, if you’re struggling with templatizing pipelines, then you go into the Pipelines as Code channel, right? Or if you have questions about how to use Kubernetes in Spinnaker, then you use the Kubernetes Slack channel.

    That was very, very helpful for us, especially in the beginning when we were trying to figure out how things are working.

    Ashley: Mm-hmm, I’ll bet, I’ll bet. How long have you been using Spinnaker, then, at Salesforce?

    Vergara: We started our journey last year.

    Ashley: Okay. [Cross talk]

    Vergara: Late last year, so almost one year now. [Laughter]

    Ashley: Wow, okay.

    Vergara: Yeah.

    Ashley: Wow. I bet a lot has happened in that one year. I know one of the things that you’re planning on also talking about is some contribution ideas for those that might want to help make Spinnaker better. And those are things that you—ideas you had or lessons you’ve learned implementing Spinnaker there, I assume?

    Vergara: Right. So, some of the contribution ideas that we have is with regards to a new effort called Kill. I’m not sure if you have interviewed the speakers.

    Ashley: I haven’t yet, but I may. [Laughter]

    Vergara: Okay. Yeah, so I’ll let them—you know, I won’t steal their thunder.

    Ashley: Don’t steal their thunder—okay, great. [Laughter]

    Vergara: [Laughter] Right, so, it’s one of those, they’re focused more around the managed delivery, so that’s when. Another contribution idea is centered more around each of the providers. So, again, if you’re building for one provider, it’s a different, totally different stage. So, for example, like GCEVM, right? So, they have their own deploy stage for that. And then AWS, they would have their own deploy stage for that. They’re kind of not, like, at par right now. Like, you know, they focus on certain things.

    I think another contribution there is trying to add more options or features into that, so it can fully support that provider. And so, for example, for GCE, I think one of the features lacking there, if I remember correctly—or they might have fixed it. But it was, say, so if you have a persistent disk and you want this particular persistent disk, then it’s not supported. It always creates a new one. So, it doesn’t give you an option to actually reuse an existing one that you’ve built earlier. So, things like those. I think that would be important.

    Same thing with AWS, right? So, there’s some features for EC2 VMs during deployment that could be added in there as well.

    Ashley: Great. I’m wondering, too, thinking about your talking that’s coming up, who are the kinds of folks that would be, get the most out of it, coming to your—obviously, people that wanna perform this function, but is that usually, is it developers, is it DevOps engineers, is it CI/CD automation specialists? Who are folks that would get the most out of your talk?

    Vergara: It would be around the Spinnaker administrators within their companies. So, if they want to learn how to configure Spinnaker to enable features for baking and deployment, then they would be the ones who would learn the most out of it. But I would say, like, all the DevOps engineers who want to figure out how they actually use Spinnaker for baking and deployment as well. Like, how do they set up their pipelines?

    I’m gonna give a demo on how it actually works from start to finish so that they can have a visual. Sometimes, you need that visual to say, “Oh, yeah! Oh, that’s how it works!” Right?

    Ashley: [Laughter] Mm-hmm. Well, that guidance is very helpful. And I can see, too, what you’re saying about the folks that are obviously administering Spinnaker in that deploy process are ideal to attend, and of course, other DevOps engineers automating and wanna know how this part of it works certainly would be good to go as well.

    Well, I wish you the best of luck. I know you’ll do real well. You obviously know, have learned a lot and know a lot from your experience at Salesforce. So, I wish you the best.

    Vergara: Well, thank you. Thank you, Mitch, again, for the opportunity for having me talk here.

    Ashley: Well, absolutely. Thank you for being on. I wanna thank my guest today, Jing Vergara, who’s princ—I can’t say it [Laughter]—principal software engineer at Salesforce. There, I got it out. Now, Spinnaker Summit 2019 is November 15th through the 19th in San Diego, and Jing’s talk is on Sunday, November 17th, at 10:45. It’s titled, “Demystifying Spinnaker VM Imaging Baking and Deployment.”

    So, thank you for joining us today for this DevOps Chat podcast. This is Mitch Ashley with staging-devopsy.kinsta.cloud. Have a great day, and be careful out there.

    — Mitchell Ashley

  • DevOps Chats: Canary Deploys Using Istio, with Autodesk

    DevOps Chats: Canary Deploys Using Istio, with Autodesk

    Spinnaker Summit 2019 Preview: Canary deploys give us a window into how new code deploys perform in production on a limited basis. The technology holds great promise in helping us learn the positive or negative effects of deploys without putting the more extensive set of microservices, application functions and beyond at risk.

    How do you route traffic? How do you apply a virtual service manifest? Is there a quick way to perform a smoke test to find server errors? When comparing metrics from new code to old, how do you know which results are good or bad? There is a lot to know and learn about how to deploy and use Canaries.

    Our DevOps Chat guest Omar Al-Hayderi, engineering manager and principal engineer at Autodesk, is giving a talk on “Canary Deploys with Istio: Lessons Learned” at the Spinnaker Summit 2019. His talk is on Sunday, Nov. 17, at 1:30 PM PT.

    As usual, the streaming audio is immediately below, followed by the transcript of our conversation.

    Transcript

    Mitch Ashley: Hi, everyone. This is Mitch Ashley with staging-devopsy.kinsta.cloud and you’re listening to another DevOps Chat podcast. Today, I’m joined by Omar Al-Hayderi. He is engineering manager, formerly principal engineer at Autodesk. And Omar is gonna be speaking at the Spinnaker Summit 2019 in San Diego. His talk is on Canary deploys with Istio. It’s happening on Sunday, the 17th of November, 1:30 p.m.

    Omar, welcome to DevOps Chat podcast.

    Omar Al-Hayderi: Thanks for having me, Mitch.

    Ashley: Awesome to have you here. I’m excited to hear about your talk. But first, tell us a little bit about yourself, introduce yourself, and what you do at Autodesk. I think we know about Autodesk, maybe what part of Autodesk that you work in.

    Al-Hayderi: Right, yeah. So, I’ve actually gone through a lot of change recently. I was working at a startup called PlanGrid, and we were a small construction software company four years ago, grown a lot since then. I started off at that company as sort of a backend generalist. You know, there was like 25 engineers, really, at the company, and have grown a lot since then, since the acquisition. We got purchased at about 400 employees.

    In that time, I’ve sort of evolved naturally into more of a DevOps infrastructury engineer role, and—yeah, it was kinda the best four years of my life. Really learned a lot there. Post-acquisition, I moved into a principal engineer role on the infrastructure team. We then went through sort of a re-org and I became the engineering manager of the back end platform team. And now, that team, we sort of specialize in building a lot of tooling around our sort of DevOps and infrastructure cloud.

    Ashley: Mm-hmm, great. Well, great experience to bring with you to Autodesk. It sounds like you’re working on some fantastic things. How did you come to decide that you’re gonna talk about Canary using Istio at the Spinnaker Summit? Was that something you chose to kinda bring in and start up at Autodesk, you joined a team already using it? What’s kinda the background on how you kinda came to this place to wanna talk about it?

    Al-Hayderi: Yeah. So, I was sort of, you know, dipping my toes into Spinnaker at the early stages of it, and I was really fascinated by the tool. I loved using all sorts of features and playing around with it. And at PlanGrid, actually, before we were even acquired, we did yearly hack weeks, and we got to sort of move off of our regularly scheduled program and work on whatever we felt like, just had fun.

    And one of the things that recently came out for Spinnaker was this idea of Canary deployments. So, I experimented with it, never got it off the ground, fast forward one year later to the next hack week, Kayenta, the Spinnaker Canary component, was a lot more built out and user friendly. So, I managed to get it off the ground then.

    We sort of have a monolith API at PlanGrid. And so, a huge problem we had was, we would do weekly releases, and after the release, stuff would break and people would get paged and we’d have to roll back, so.

    Ashley: [Laughter] Vicious cycle.

    Al-Hayderi: Common story I’m sure everyone’s familiar with.

    Ashley: Mm-hmm.

    Al-Hayderi: So, I really wanted to find a way to solve that, and Canary just seemed like a great way to do that. So, after the hack week, I sort of kept working on it full time, and I started doing, you know, presentations to the rest of the company. Because Canary is sort of this new idea that a lot of people aren’t familiar with. I started, you know, being active in the Spinnaker Slack channels, and it actually turns out that an ex-PlanGrid employee moved to Netflix to be a lead on the Spinnaker open source software from there.

    Ashley: Ah, interesting.

    Al-Hayderi: Yeah.

    Ashley: Small world.

    Al-Hayderi: Yeah. [Laughter] And he noticed I was talking a lot about it and he brought up the point that they’re running the Spinnaker Summit in November and he asked if I wanted to talk. And I’ve actually never given a talk at a conference.

    Ashley: Oh, excellent.

    Al-Hayderi: Yeah, and I do a lot of talks internally. So, this is something I was super excited about and, yeah, I jumped on it.

    Ashley: Great thing to do for your career, too, and it sounds like those internal talks have set you up well to do this.

    Al-Hayderi: Yeah.

    Ashley: So, how did you go about deciding, is this gonna be kind of an overview of how to do it? Is it going to be, “Here’s the metrics to watch” and the value that comes out of it is a combination of that? Tell us a little bit about what part of this you’re—kinda what take you’re taking in your talk.

    Al-Hayderi: Yeah. So, public speaking was honestly never my strong suit when I was younger. And some great advice I got was that, when you’re giving a speech to a bunch of people, you wanna tell it like a story.

    Ashley: Mm-hmm.

    Al-Hayderi: So, since then, I’ve really focused on storytelling for even these tech talks I give internally.

    So, the story I’m trying to tell here is how we leveraged Istio to finally solve some of the deep technical problems of running Canaries.

    Ashley: Mm-hmm.

    Al-Hayderi: And once we got it out there, sort of the growing pains and, you know, first obstacles we hit and problems we’ve got. I’m hoping people come out of this with a better idea of how to actually get this running in your infrastructure and some common pitfalls to avoid.

    Ashley: Mm-hmm. Great. Well, tell us a little bit of that story. Kinda get us started down that path.

    Al-Hayderi: Totally. So, the first thing I did was, I went in and learned exactly what Canaries are. And the basic point of it is that you want to expose a small amount of traffic to something new. You want to analyze the deltas and some sort of metrics or something between the new and the old. And then you want to make a decision if this new is ready for the majority of traffic.

    Ashley: Mm-hmm.

    Al-Hayderi: So, there’s a lot that goes into that. Anyone that’s worked with Spinnaker on production scale knows that managing all these pipeline definitions is very cumbersome, and trying to, if you end up changing something, especially changing production pipelines, it can be really dangerous.

    So, a lot of the things we learned about our, you know, how we tested this, even just testing the Canary process in our sort of R&D Dev staging environments and the problem there, even though we got confident with how the pipelines worked, it was very difficult to get a signal on how to actually tune your Canaries. And when I say tune, I mean, when you’re comparing metrics from the new and the old, how do you know if it’s bad or how do you know if it’s a good decision to move forward? What are the thresholds, what metrics are you looking at? A lot of our iterations were focused on that.

    Ashley: Interesting. Now, I’m curious if—are you approaching this problem of using Canary to understand the service mesh or the services itself, or are you trying to also understand kinda Istio and it acting as a sidecar proxy and how that’s performing coordinating traffic load balancing across that? Is it one or the other or both, or what part of the problem were you looking to solve?

    Al-Hayderi: Right. So, actually, we were kind of blocked on the fact that we couldn’t route traffic, you know, with small granularity coming in from the Internet. And we actually only use Istio for its API gateway feature at the moment.

    Ashley: Oh, okay, okay.

    Al-Hayderi: Yeah, there are some blockers right now, at least in our version of Istio, that we’re running around connecting over SSL to Redis. So, that’s blocked having sidecars or application pods. So, when we do canaries, we’re just canarying end user or ingress traffic into our monolith.

    Ashley: Okay.

    Al-Hayderi: Yeah.

    Ashley: Interesting. So, tell us the story. How did you, you went about implementing it—what did you learn from it?

    Al-Hayderi: So, I mean, first off, we had our sort of snowflake-y pipeline that would deploy our own little service and change traffic and compare metrics. So, I guess the first part about it was how do we route traffic and what we actually do is, in a single pipeline, we’ll deploy containers to a new supergroup and we’ll, that’s using the V1 provider of Kubernetes.

    Ashley: Mm-hmm.

    Al-Hayderi: And then we actually used the V2 provider of Kubernetes to apply a virtual service manifest. Virtual services are just a kind of way of routing ingress traffic based on some set of rules. So, if the host header matches this and the path has this Regex, route it to this service.

    What you can also do is, when you match a rule based on some ____ path prefix is route to several services based on weights. So, we first experimented with how are we gonna get a small amount of traffic, yet a meaningful enough amount of traffic to get a good enough sample size of data.

    So, a lot of our, at the start, what we were doing was just sort of like load testing and seeing, “Okay, is 5 percent of the traffic to the Canary deployment enough? How about 10?” et cetera, et cetera. We also looked at some common best practices around Canaries and how they recommend running baseline deployments. So, what we would actually do is deploy two new deployments. So, one with the new code on a small sized server and one with the old code on a brand new, deployed small server. That way, you sort of take away a lot of variables from the experiment, like long running processes, et cetera.

    So, we got there, got traffic routing correctly. Now, the next part we had to do was metrics, right? How do we analyze the delta between the two deployments? At this point, Kayenta only supported Datadog, at least out of the metrics providers that we used.

    Ashley: Okay.

    Al-Hayderi: So, what we really wanted to do, you know, going back to the whole point of getting Canaries out was stop deployments that are just plain broken, right? So, kind of getting, like, a cheap, easy way to get smoke tests. And so, we were just looking for server errors.

    So, as we were trying that, the problem is, is that the server error metric that Istio gives you will only send metrics of a server error happens. So, between two deployments, if you only get two server errors on one, that will trip the Canary as a failure. It doesn’t give you the option to aggregate metrics, so we couldn’t do things like check error rates. So, that just really never worked.

    And actually, in my talk, I go into sort of the—this was kind of a big, big problem the first time we actually released this into production is that Canaries would just fail a lot for no reason, right? And Canaries running for an hour long, you know, our release engineer is sitting there waiting for an hour and then has to trigger it again just because, you know, one server error happened on the Canary deployment.

    Ashley: Well, it sounds like maybe one of the lessons you learned or concept you came up with is, don’t cast a big net. Start out with sort of higher, larger grain, most fundamental issues, kind of a ____ kinda concept. Start there, get those—get that quality fixed, improved, so you’ve got those things working better, probably tighten the net a little bit more and filter out some more specific things that are causing issues. Does that sound like the approach that you learned how to take?

    Al-Hayderi: That’s exactly it, and actually, what I would recommend is, the first time you give out Canaries, don’t have it actually fail or roll back the deployment if the Canary was unsuccessful.

    Ashley: Mm-hmm, okay.

    Al-Hayderi: What we actually ended up doing was actually running this for a month and just gathering data, seeing which metrics worked, which metrics didn’t, and mapping up Canary reports with what we actually saw as problems in production. That way, that gave us more information on which metrics to use. We ended up going with average latency and 95 P latency and once that was out, we started actually aborting and rolling back deployments and then we started to get a lot of great value of this tool.

    Ashley: Mm-hmm. Yeah, that’s actually very similar to kind of a quality process, right? You don’t solve all problems at once. You start, you prioritize, you filter, you work done—you know, improve it one step at a time, rather than just throwing a canary out there and seeing what happens, which sounds like you’ll get results, but you don’t know what or why, right?

    Al-Hayderi: Yep, exactly.

    Ashley: You’ve gotta dig into it and be a little more thoughtful about it, it sounds like.

    Al-Hayderi: Yep.

    Ashley: Good. Anything else? Any kinda other big things you’re expecting to talk about as part of this?

    Al-Hayderi: So, another big thing we learned was, even after this, when we had, you know we were confident that our Canaries were failing when they should and passing when they should. The problem we got was, when a Canary failed, people didn’t really know what to do.

    Ashley: Mm-hmm.

    Al-Hayderi: Because, you know, when a deployment fails, we have a lot of data and we have a lot of observability into that. The Canary is getting a lot less traffic, and it’s hard to know exactly what failed just by looking at the Spinnaker UI Canary report.

    Ashley: Mm-hmm.

    Al-Hayderi: So, we then, you know, sort of invested more into our—we used New Relic for our APM and we piped our Canary into that so that we could actually see transaction data and the delta between it and the baseline metrics to see, you know, which end point, for instance, was misbehaving or was some new database core utilizing this—that helped us a lot, too.

    Ashley: Mm-hmm. Interesting. Well, it sounds like you’re using—think about this in a systemic way, not just Canaries and Spinnaker, but also you mentioned Datadog, New Relic, a number of tools that you’re using as part of your environment, which all have to come together to figure it out, right? It’s not just one thing.

    Al-Hayderi: Yep.

    Ashley: Well, good. I wish you the best in your talk. The folks that would go to this talk, are they gonna be, tend to be software developers, engineers, architects, operations focused on the DevOps team? Who tends to be more interested in this than others?

    Al-Hayderi: I would say more the DevOps operational people. People who work with the Spinnaker pipeline definitions at their company would get the most value out of this. But, you know, back when I was just a developer, I heard a talk about Canaries, and that sort of triggered me to bring this in. So, I mean, anyone who’s interested in resiliency and, you know, release confidence would get value out of this talk.

    Ashley: Okay, excellent. Very good. Well, thank you for being on the podcast, Omar.

    Al-Hayderi: Thank you very much, Mitch.

    Ashley: It’s been great to have you, Omar Al-Hayderi, who is engineering manager, formerly a principal engineer, at Autodesk. He’s gonna be speaking at the Spinnaker Summit 2019 in San Diego. The dates for that is November 17th through the 19th, and Omar’s talk is on Sunday, the 17th, at 1:30 p.m.

    So, thank you, all of you, for joining us and listening to this episode of DevOps Chat podcast. This is Mitch Ashley with staging-devopsy.kinsta.cloud. You’ve listened to another DevOps Chat. Be careful out there.

    — Mitchell Ashley

  • DevOps Chats: Plugins Extend Spinnaker for the Enterprise

    DevOps Chats: Plugins Extend Spinnaker for the Enterprise

    Spinnaker Summit 2019 Preview: Secrets management is vital to any vibrant DevOps and GitOps toolchain. As part of its open source maturation and thanks to the contribution of our DevOps Chat guest, Spinnaker now has a secure secrets management capability enabling GitOps, taking non-application secrets out of plaintext in hal config files.

    Cameron Motevasselani, software engineer at Armory, is giving his talk “Spinnaker Plugins: Extending Spinnaker for the Enterprise” on Sunday, Nov. 17, at 10:45 AM PT.

    In addition to his contributions to Spinnaker’s new secrets management facility, Cameron shares the work in progress to create an extensible Spinnaker plugin system. The current work is an early MVP to be followed by the implementation of PF4J and concrete extension points as part of Spinnaker stages.

    As usual, the streaming audio is immediately below, followed by the transcript of our conversation.

    Transcript

    Mitch Ashley: Hi, everyone. This is Mitch Ashley with staging-devopsy.kinsta.cloud and you’re listening to another DevOps Chat podcast. Today, I’m joined by Cameron Motevasselani—Motevasselani, thank you [Laughter]—a software engineer. He’s also a speaker at Armory. He has a couple of talks that are at the upcoming Spinnaker Summit 2019 in San Diego. First one—I guess this first one is on the plugin system and the second one is on secrets management.

    So, Cameron—welcome to DevOps Chat.

    Cameron Motevasselani: Hi. Thanks for having me.

    Ashley: Glad to have you on. Well, tell us a little bit about yourself. Introduce yourself and tell us a little bit about what you do at Armory and what Armory does.

    Motevasselani: Oh, totally. Yeah. So, I’m Cameron Motevasselani. Thanks for the quick introduction. Let’s see, I’ve been at Armory for about a year now. I joined and hopped feet first and started working on this secrets management project. So, that was pretty cool. And then, now, recently, I’ve been working on this plug-ins system. Yeah, that, I think, about sums it up.

    Ashley: Great. What does Armory do? Yeah.

    Motevasselani: Yeah, Armory, we basically have an enterprise offering for Spinnaker. So, we have some extensions that we’ve written on top of Spinnaker, so we don’t fork and have to maintain all that, we just—we extend Spinnaker and provide some extra features that way.

    Ashley: Very nice. Now, the way you phrased that, you’ve been working on secret management and the plugin. Does that mean you’re actually contributing, writing code and contributing it to the open source, or are you primarily an implementer of what’s already been built?

    Motevasselani: Oh, no, that’s a great question. Yeah, for the secrets management as well as plugins, we both, we’re contributing directly to open source for both of those projects.

    Ashley: Awesome. Excellent. Well, let’s start with secrets management. I mean, I think we all kinda have the idea of where we’re gonna keep our keys, our passwords, certificates, all that good kinda stuff. We don’t want that heading into the source code management system or somewhere in the cloud or in someone’s software. Tell us a little bit about the secrets management part of Spinnaker and what you’ve been doing, what you’re gonna talk about.

    Motevasselani: Yeah, totally. So, secrets management really came up as we started talking with customers about their Spinnaker setups and how they deploy and manage the configurations for Spinnaker.

    As you know, Spinnaker has the keys to the kingdom, essentially. It has your cluster keys and your credentials in order to deploy cloud resources that cost you money, right? And those are running—it runs the code that your organization arrives on. You can’t have that in source control. That is not, not good to have.

    Ashley: No-no—a big no-no, though a common mistake. [Laughter]

    Motevasselani: That’s right. Right. [Laughter] I think GitHub even yells at you now, if you have keys—

    Ashley: Oh, does it? [Laughter] Nice.

    Motevasselani: – in there. So, they’ve been adding nice little security features here and there. People would like to move towards this idea of GitOps so that you can manage your configuration via pull requests. It’s a great idea, and a lot of people can’t do that, because the keys are quite literally in the Spinnaker configurations.

    So, we pull those keys out and are essentially referring to them with our secrets management solution. So, you can use different backing stores such as S3 buckets or Vault and store your secrets there and then refer to them in the Spinnaker configuration that way.

    Ashley: Mm-hmm. Excellent, great. So, in your talk, you’re gonna be walking through what you’ve been doing, the functionality of it, how to implement it via GitOps?

    Motevasselani: Yeah, so—that’s right. The functionality, how it differs from, there are other solutions coming into open source now that tackle a slightly different use case. So, discussing the differences there as well as the future, essentially, of secrets management.

    Ashley: Okay. Very good. Great. Anything else to tell us about that before we move into the plug-in system?

    Motevasselani: Yeah. You know, secrets management, it is about the secrets for Spinnaker in particular. It’s not actually tackling the application secrets themselves.

    Ashley: Oh, okay.

    Motevasselani: So, I just wanted to make that clarification. And that’s at the current time. It’s possible that we might have a story for that in the future, but the talk that I’ll be giving is around Spinnaker secrets, not application secrets.

    Ashley: Okay. So, this is just what’s required to run, operate, manage the Spinnaker environment, not your application secrets?

    Motevasselani: That’s correct.

    Ashley: Okay, great. Well, that’s an important clarification. Great. Well, let’s talk about the new plug-in system. What is new about it, other than being new code?

    Motevasselani: Yeah. So, the plug-in system is, we’ve been developing it as a very—it’s been a very alpha, MVP type feature. So, right now in open source, we do have a method for users to create custom stages in Orca. There is quite a bit of knowledge and expertise that you have to have in order to do that, so I don’t quite recommend doing it right now. But we’re building on this plug-in system and we’ll be testing it out with users hopefully in the near future.

    So, right now, there’s currently no UI, but we’re working on a UI and then that should be getting open sourced as well in the future.

    Ashley: Now, that sounds like kind of a tough assignment. How do you create a plug-in system that anybody can do anything they want with, because that’s probably one of the requirements is, you want a lot of flexibility. How do you go about architecting and designing that so what you’re building can be extensible and you’re not gonna have to re-do it in a year?

    Motevasselani: Yeah, that’s such a great question. Because we are going with this MVP model of just getting something working and then getting feedback on it to iterate, we actually are already planning on deprecating this first, let’s say, method of loading plugins into Spinnaker. We’ve already been doing some internal revs on it and we’ve been working with the community, Google and Netflix in particular, on this project. So, we are going to be moving away from our current implementation into another full-fledged plug-in framework, and we’ve chosen PF4J for that.

    Ashley: Okay. Now, I’m not familiar with that framework. Can you tell us a little bit about that?

    Motevasselani: Yeah, so the way we currently did it—and hopefully, this isn’t giving too much of my talk away [Laughter]—but the way we’re doing it right now is essentially shoving JARs into a ClassLoader for Orca. Well, via Quark. So, these are different Spinnaker services and libraries that make up Spinnaker.

    Ashley: Mm-hmm.

    Motevasselani: Right now, it’s a pretty dumb way of doing it, just adding JARs onto a Classpath in order to load it onto the Spinnaker spring application context.

    Ashley: Makes sense. Pretty easy way to do it.

    Motevasselani: Yeah, exactly. But like you’re saying, there’s no—it’s very open ended right now. We are going to, as we move to PF4J, define concrete extension points within Spinnaker so that there’s actually not a sprawl of extension points, there’s clearly defined extension points throughout the code.

    Ashley: Mm-hmm.

    Motevasselani: That will help us iterate on this feature and make sure it’s a good solution as we continue developing it.

    Ashley: Okay, and this is the PF4J that’s on GitHub? I think it’s an Apache license, right?

    Motevasselani: Yeah, that’s correct. I believe it’s Apache. That’s correct. And so, moving into PF4J will give us a lot of really nice features that our current implementation doesn’t have, such as using a different ClassLoader for every plug-in so that it can have its own dependencies, et cetera.

    Ashley: That makes a lot of sense. Are there any unique challenges for implementing a plug-in system in an environment like Spinnaker? You know, this is software to create software, software to deploy software. You know, it’s not an application in and of itself, it’s something that you’re gonna make a key part of your kinda tool chain, your workflow pipeline. I wonder if there’s any unique challenges that come about in trying to build this.

    Motevasselani: Yeah, it’s interesting, because we have been looking at, you know, as you develop something new, you look at prior art in order to determine how to build what you’re building, I guess. [Laughter]

    Ashley: Mm-hmm.

    Motevasselani: For lack of a better word.

    Ashley: Technical speak, I know. [Laughter]

    Motevasselani: That’s correct. [Laughter] So, looking at something like Jenkins or Eclipse that both have plug-in systems, same with WordPress, those are all kind of monoliths, right?

    Ashley: Mm-hmm.

    Motevasselani: Spinnaker is definitely many different microservices composed into one application. Essentially, we’re just kind of turning it into a platform, kind of, by making it extensible and pluggable.

    But that is, to answer your question, a major challenge is, you have all these different microservices. You know, depending on how you deploy and update your Spinnaker installation, they might be on different—I don’t wanna say different versions, but if you’re deploying master all the time, you might have a different set of microservices in production compared to if you’re all running the same release version.

    So, deploying plug-ins into that environment is not an easy task and how you get the versioning done is important, as well as rollbacks.

    Ashley: You kinda have a recursive problem here, right? Because you have to version control the plug-in manager, which then version controls of what plug-ins can work with it, c?

    Motevasselani: Yeah, exactly. Plugins might rely on some functionality that is in one version of a microservice that’s not in another version, and if you need to do a rollback, how does that work, et cetera. That’s correct.

    Ashley: Well, you kinda took my breath away for a moment. You scared me when you said, you mentioned the WordPress plug-in system.

    Motevasselani: [Laughter]

    Ashley: Please, don’t do that! [Laughter]

    Motevasselani: I do think, though, that WordPress is an example of a project that has done very well—

    Ashley: They have.

    Motevasselani:—because it has such as good ecosystem of plugins.

    Ashley: It’s come a long way. It has come a long way. It’s much more, I don’t know, not self-sustaining, but kinda auto updating much better than it used to be. [Laughter]

    Motevasselani: Right, right. And we do hope to take learnings, right, from past projects, let’s say, as well. [Laughter]

    Ashley: Absolutely. I don’t mean to dis WordPress, you know? You gotta every once in a while. Well, tell us a little bit about more about your talk. What are you gonna cover, then? Is this more how to use it, or you’re gonna lay out kinda where you’re going in this next version and you’re looking for feedback? What are you trying to accomplish or get out of it? What would you like to have happen in the talk?

    Motevasselani: Yeah, you know, that’s a really good question. I think there’s a lot of stuff to cover, all the way from our journey running and creating this project, working with the community and those efforts, to showing users, “Hey, this is actually a nice and easy way of modifying or extending Spinnaker’s functionality.”

    The idea here is to—there’s kind of a few ideas here with this project, because we really want to lower the barrier of entry for users, particularly in enterprises, to be able to create extensions for Spinnaker. The method we chose to do this at first was a way that most users interact with Spinnaker, and that’s creating pipelines and defining stages.

    So, let’s say an enterprise has some custom infrastructure at their business running their data centers, running their cloud. If they want to, let’s say, interact with that in a native—in a Spinnaker native way, you would want to have a stage that their users can just choose. So, rather than having to learn how Spinnaker works on a very deep level, you can implement this nice, simple interface and get it going.

    So, at the end of the talk, it would be great for the audience to have an understanding of how to create and implement their plugin on their own.

    Ashley: Oh, excellent. Well, then, hopefully, you can, you’re creating an audience and a community to go out and try your new plug-in system and the variations of it.

    Now, you talked about using plug-ins as extensions to Spinnaker. Do you also see this plug-in system being kind of an integration interface to third party tools, maybe other open source, other commercial tools, or is it more things that just extend the functionality of Spinnaker?

    Motevasselani: Oh, that’s a really good question. So, there is a push to move Spinnaker to be a lean core and an expansive ecosystem, I believe, is the term being used, and this goes directly towards that goal. There’s certain functionality in Spinnaker, the core, Spinnaker core, that doesn’t quite make sense.

    For instance, there’s some stages that are specific to companies. Don’t wanna call Gremlin out, I think they do great work. There’s a Gremlin stage in Spinnaker that should probably be a plug-in. So, hey, Gremlin guys, hit me up. [Laughter] But not quite yet, maybe for the end plug-in system.

    Ashley: It’s that evolution of software, right?

    Motevasselani: Exactly.

    Ashley: I mean, that’s how you get to this point where you need to do that, you need to spin things out, so it’s part of its growth.

    Motevasselani: That’s right, that’s right. Yes. It would be really cool if you could have a Spinnaker installation where you can essentially pick and choose which cloud providers you want to have installed and running. If you only—you know, if you interact with, let’s say, both Google and Amazon’s cloud, maybe you don’t need Oracle’s cloud software to be on your cloud driver. Or maybe you want to have a cloud driver in each of those data centers—sorry, in each of those cloud accounts that only has the code necessary to deploy to that cloud environment.

    Ashley: Mm-hmm.

    Motevasselani: So, then you have a smaller attack surface, just leaner and smaller binaries, et cetera, et cetera.

    Ashley: Great. Well, I think, unless there’s anything you wanna add, I think you’ve given a really comprehensive look to kinda what that talk, both talks are gonna be about.

    Motevasselani: Yeah, that’s—you know, that’s hopefully, just getting the community excited and get some ideas flowing, I think, would be great with those talks.

    Ashley: Well, and I also wanted to just throw in my little thanks just to appreciate what you do and all the other contributors to the projects in software. And I always encourage people who are just using open source, and there are folks at the Spinnaker Summit who are implementing it, not writing code for Spinnaker, but talking about it at a summit like this is another way to contribute to the community. So, great to talk with you who are actually writing code as well.

    Motevasselani: Yeah, thanks. Nice talking with you, too, Mitch.

    Ashley: Great. Well, it’s been my pleasure. You’ve listened to another DevOps Chat podcast. I’d like to thank today’s guest, it’s Cameron Motevasselani, software engineer at Armory, and he is speaking at Spinnaker Summit 2019. The dates for the conference are November 15th through 19th in San Diego, and Cameron has two talks. His talk on the plugins is Saturday, November 16th, at 2:30 p.m. And then he’s also talking bright and early about secrets management—boy, you’re really gonna wake ‘em up here—Sunday morning at 9 a.m., Cameron, so. [Laughter]

    Motevasselani: [Laughter]

    Ashley: So, you know, maybe you serve mimosas at the morning talk. [Laughter] So, it’s been my pleasure to have Cameron on, and thank you to all of our listeners today for joining us. This is Mitch Ashley with staging-devopsy.kinsta.cloud. You’ve listened to another DevOps Chat podcast. Be careful out there.

    — Mitchell Ashley

  • DevOps Chats: Debugging Spinnaker Apps, With Salesforce

    DevOps Chats: Debugging Spinnaker Apps, With Salesforce

    Spinnaker Summit 2019 Preview: Debugging production issues in any environment can be challenging, and Spinnaker has its production learning curve. Problems aren’t always replicable in a smaller environment, and debug messages can be verbose and confusing to triage what’s happening.

    Our DevOps Chat guest Chuck Lane, Salesforce lead software engineer, is giving a talk on “Debugging and Profiling Spinnaker Applications Live” at the Spinnaker Summit 2019. In Chuck’s talk, you’ll learn skills like remote JVM debugging, custom profiling builds and the magic of figuring out what’s going on with a multithreaded microservice using htop.

    Chuck’s talk is on Saturday, Nov. 16, at 3:40 PM PT. Spinnaker Summit 2019 is Nov. 15-19 in San Diego.

    As usual, the streaming audio is immediately below, followed by the transcript of our conversation.

    Transcript

    Mitch Ashley: Hi, everyone. This is Mitch Ashley with staging-devopsy.kinsta.cloud and you’re listening to another DevOps Chat podcast. Today, I’m joined by Chuck Lane. He’s a lead software engineer with Salesforce.com. He works on Sales Cloud, which is the Salesforce you and I use, know and love, or our CRM capabilities. Chuck is talking at the Spinnaker Summit 2019 in San Diego, and his talk is about debugging and profiling Spinnaker applications live. And I think live means, like, live while they’re running, in production, and see what’s going on. We’ll explore what that is. This talk is on Saturday, November 16th, at 3:45 p.m. in San Diego. Chuck, welcome to DevOps Chat.

    Chuck Lane: Thank you. It’s nice to be here.

    Ashley: Excellent. Awesome to have you here. Would you start by just introducing yourself? Tell us a little about you and what you do at Salesforce.

    Lane: Yeah. So, my name is Chuck Lane. I am, as you mentioned, a lead software engineer. And basically, my chief responsibility at Salesforce is to help bring Salesforce into using Spinnaker for our public cloud based deployments. So, as we make a transition to try to move away from first party architecture and move over into the public cloud, one of the technologies that we’re using to do that is Spinnaker and I’ve established myself as a subject matter expert in Spinnaker. And so, I kinda help to bridge the gap between the traditional development and deployment cycle and what that looks like using Spinnaker in a containerized world.

    Ashley: Okay, great. You used a term there I don’t know if I’m familiar with. You mentioned something about moving from first party architecture to the public cloud. What’s that mean about the environment you’ve been in that you’re, and what you’re moving to?

    Lane: Yeah. So, basically, what I mean there is that, historically, Salesforce has owned their own data centers, and we’ve deployed to those data centers. We are moving, embracing in a big way public cloud architecture, whether that’s through AWS, whether that’s through GCP, or any of the number, any other number of cloud offerings.

    And so, as opposed to doing what some people would call a lift and shift where you just take your software that’s meant for, to be run on hardware that you own and move it into the cloud, we’re doing a fundamental re-architecture of the software to make use of all of the great things that cloud architecture allows us to take advantage of.

    Ashley: Mm-hmm. Fantastic. That’s interesting to hear a little bit about your own evolution at Salesforce. So, your talk is about using Spinnaker, you know, really about debugging and finding out problems that are happening in production, but sometimes you might struggle with trying to replicate it into a smaller environment to get the same problems working.

    Can you tell us a little bit about what sort of led you down this path to figure out how to do this? Was it a big problem that was happening or something you saw repeated over and over that there wasn’t a good solution to and you figured out how to do that? How did this come about?

    Lane: Yes. So, basically, as you’re taking workloads and migrating them over, I mean, when—you know, when you just start with a small subset of workloads, you know, the path is relatively straightforward. And, you know—hopefully, anyways—and you don’t run into too many issues that can’t kind of easily be solved.

    But the reality is, as you transfer more and more workloads over to the public cloud, that’s—you know, the devil’s in the details, right? So, that’s really where things can pop up where you’re hitting various limits or software isn’t performing in the way that you would expect it to perform. And these are the situations where we—where it can be very beneficial to jump into production software and really, you know, sometimes slap a debugger on it and see exactly what it is that it’s doing that differs from kinda what you expect. You know, and that can help you kinda tailor your workloads in such a way that you can make things run more smoothly.

    Ashley: Mm-hmm. I know sometimes using debuggers, enabling them—that can be helpful, but it also can be too much information, trying to sort through what all is happening, trying to find what should you be looking at. What are some of the challenges you’ve found by turning on debugging or using debuggers?

    Lane: Yeah. Well, so, I mean, there definitely is that problem that you’re saying just as far as too much information that’s coming out. One of the nice things about Spinnaker and the way the services are written is, you have a very fine grained tuning over which libraries inside of Spinnaker you tell it to print out and debug information.

    Ashley: Mm-hmm.

    Lane: So, if we know that there is a hiccup with one of the data binding layers or a hiccup with one of the authorization layers, then we can, through the config files, we can—and Java settings—we can really target that explicit directory, or I’m sorry, that explicit library and say, “Give me all the debug information that you have.”

    Ashley: Mm-hmm.

    Lane: But, you know, failing that, I mean, there are definitely times when it’s been advantageous and the best course of action is just to go down and actually look at the code and see what it’s doing. And, again there, it’s best to do that under—you can’t do it in production, usually, but what we can do is simulate a load that’s similar to what we would see in production and really, really take a look at what’s going on to help us identify those algorithms that might be o event squared instead of a login or something like that.

    Ashley: Mm-hmm. Now, are there some specific techniques—I know in your description of your talk, when you talked about using remote JVM debugging, custom profile builds, even using htop, you know, a UNIX command or a LINUX command to help you figure out what’s happening with multi-threaded microservices. It sounds like a variety of different approaches that you’ve kind uncovered and learned that you can use to figure out what’s happening.

    Lane: Yeah. And so, I mean, in general, we like to run as closely to the open source build as we can, the ones that are provided by Spinnaker, which is ultimately provided by Google. But there definitely are circumstances where we need an extra tool, be it htop, be it something like Glowroot, which is a Java application profiler where what we’ll do, then, is we’ll go in and we’ll build a custom image using the Spinnaker images as a base and then tack on those additional libraries and put them into the Spinnaker ecosystem and then launch them up just to see what additional information we can get out of there.

    And so, you know, one scenario where that was really useful to us, we ran into a scenario where cloud driver, which is—cloud driver is the main tool that talks back and forth to all of our different cloud infrastructures.

    Ashley: Mm-hmm, mm-hmm.

    Lane: Sometimes calls to cloud driver were taking upwards of, well, two to three minutes to respond. Now, you know, timeouts are set in such a way that if it doesn’t hear anything back in about 30 or 60 seconds, then it just disregards that load and, you know, so that created quite a bit of problems for us.

    By building a custom cloud driver image that had htop in it, we were able to see that the majority of the processes that were running were actually running basically commands to reach out to the Kubernetes clusters and get a list of all of the name spaces that are in the Kubernetes cluster. And, talking it over with some of the Spinnaker developers, what we found is that if you have 15 or so clusters that you’re connecting to Spinnaker, then that’s not so much of a problem to go and query each of those to get the list of name spaces. But as you scale up, on the order of 400 or 500 different clusters the way that we do, then a lot of the delays and a lot of your time can be just essentially Kubectl calls waiting to return to you the lists of name spaces that you need to go and scan.

    So, we implemented a workflow based solution for that, basically, where we let our teams know that when they’re creating their clusters, they should use Terraform to go ahead and create the clusters—I’m sorry, the name spaces that they’ll be using. And then, if they need to use a name space after the fact, then we provide a pipeline that they can use that will dynamically add in an additional name space that their cluster will start to scan. So, that saves us from scanning all the clusters all the time, which was causing quite a bit of performance bottlenecks.

    Ashley: Yeah, I would imagine that would have some overhead, maybe a lot of overhead if you’re doing that frequently. So, it sounds like a way to both reduce overhead, but also to get the information faster.

    Lane: Exactly, exactly—yep, yep.

    Ashley: Mm-hmm. Very cool. So, in your talk, are you going to be doing any demos? I know there’s always the demo Gremlin, or are you gonna be just showing, you know, talking about what some of these techniques are?

    Lane: So, I plan on doing some demos. As far as—

    Ashley: That’s really cool.

    Lane: Yeah. You know, I may do the, what’s the—the Easy Bake Oven a little bit as far as, “Here’s the behavior and here’s what we’ve found in code.” But from what I understand, a lot of that stuff can kinda be dependent on Internet connectivity at the actual site. So, if it’s something where we have good Internet connectivity, then by all means, I plan on walking through a couple of debugging scenarios, as close to what we would do in real time as possible.

    Ashley: Yep. Well, you know, someone somewhere—it’s been a while ago—gave me the great advice of, you know, “Have your live demo and have your disconnected demo ready in the background.” [Laughter] So, you can always at least show something locally if you can’t get on the net. So, if you’re depending upon the network—I’m not sure if you are for your demo, but always a good lesson, right?

    Lane: Sure, exactly.

    Ashley: Great. Are there any other kind of lessons learned, common mistakes, or mistakes that you or others might have made along the way, kinda so there’s hard things you learned by trial and error that you plan on sharing?

    Lane: I mean, well, there are—whew. You know, I’ve been working on Spinnaker for a couple of years. So, there’s definitely been a lot of hard lessons learned, here. But honestly, what I would do is, I would encourage anybody who wants to get into Spinnaker to not just hang around the Slack channel, because the Slack channel does have a tendency to get overrun with, you know, just kinda people posting their stack traces and just saying, “Hey, has anybody ever seen this before?”

    And, you know, I’ve gotten a lot more success by going through the commits, looking at the people who actually authored the code, and then reaching out to them directly with more than, “Hey, can you explain this to me?” but rather, you know, “Hey, I see what you did here and I see what you did here. I’m running into problems with these lines. Do you have a different approach, or is there something that I can be doing differently?”

    And the other thing that I just can’t overstate is the value of being a member of one of the special interest groups.

    Ashley: Hmm, interesting.

    Lane: So—yeah. So, I’m a member of the Kubernetes V2 special interest group that’s lead by Eric Semene and Ethan Rogers from Armory, Eric’s from Google. And it has, you know, it has been just an absolute wealth of information and, you know, honestly, I don’t know if we ever would’ve gotten nearly as far as we had without those two people.

    So, you know, yeah, I would just say that the community is really friendly and we always welcome new members who are ready to learn.

    Ashley: You know, I think both of those are great suggestions, and I really appreciate that you made those. Because, one, things like the Slack channels on projects, open source, those can be a bit intimidating. Sometimes they’re not approachable, because there’s just so much noise so much happening on it and people reaching out for help like you’ve talked about, “Here’s a stack trace I’m trying to figure out.”

    But also, your recommendation of reaching out to the code authors, you know what, it’s—people like to help each other. And, you know, it might seem like, “Hey, the people who wrote this aren’t gonna have time to bother with me”—they love to hear from people that are using their stuff to talk about it and help them out, but also hear about how they’re using it, and of course, they’re always looking for ideas and feedback and stuff like that. It sounds like you’ve had that kind of an experience.

    Lane: Yeah, absolutely. I mean, and you know, the big thing is just, you know, as a coder, it’s easy to tell the people who are coming to you and just kinda wanting you to fix it for them and the people that are coming to you that have really tried to tackle it themselves. And, you know, I can’t speak highly enough about the latter rather than the former. I mean, you know, just give it your best shot and when you get stuck, reach out to somebody and it’s, you know, it can be immensely valuable.

    Ashley: It’s kinda like going to a foreign country. Wherever you are, if you’re American going somewhere else or vice versa, everybody appreciates you trying to speak the native language, and at some point, when they see you struggle enough and how far you can go, they’re glad to help you and, you know, speak in your language.

    So, same kind of thing with helping people with code. If you’re gonna just fob your problem off onto the developer of it—not appreciated so much. But they appreciate that you tried to take it as far as you could. As you mentioned, go look at the code and figure out what’s going on.

    Lane: Yeah.

    Ashley: You’ll get immense respect, you know, even if you aren’t a coder at that level, at that level of software developer, it’ll mean a lot to the developers.

    Lane: Absolutely, absolutely.

    Ashley: Well, hey, I think you’re gonna have a fascinating talk, and I love that, you know, this idea of trying to figure out what’s happening in production and some of the techniques that you’ve come up with and developed and have experienced and the fact that you’re sharing those with others. I’m curious, do you contribute any code to the Spinnaker project in any areas? Are you primarily a practitioner user of it?

    Lane: So, I have. I’ve got a few small RPRs that have been pushed through, but really, the bulk of my commits right now have actually been to the Spinnaker website, the documentation. So, you know, I don’t know Java or Groovy or Kotlin, maybe, as well. Some of the patterns they use are a little bit—I come from a .net world, so they’re a little bit foreign to me.

    But yeah, I’ve definitely written a number of different documentation pages and I found that that’s a great way to kinda get in and I’ve even got some PRs that are coming ups soon that aren’t doc related. So, yeah, hopefully, you’ll see my name more.

    Ashley: Excellent. Well, you know what, documentation is important, too. There’s some folks doing talks at Spinnaker Summit that, you know, kinda ran into that point where, “Hey, there’s documentation for doing it one way, but not under this set of configurations or software.” So, that’s a contribution, too, so congratulations for being a part of the community and for sharing, also, your experience at the Summit.

    So, Chuck, appreciate you being on the podcast today.

    Lane: Oh, thank you. It’s been an absolute privilege. Thank you so much.

    Ashley: Absolutely. Fantastic. I wish you all the best with your talk, I’m sure it’ll be great, and hopefully folks listening to this podcast will draw some more interest and bring folks to listen to you.

    So, I’d like to thank our guest today, Chuck Lane. He’s lead software engineer at Salesforce.com, so you can imagine the environment he’s working in. There’s some super good lessons that Chuck’s bringing to the table. He’s gonna be talking at Spinnaker Summit 2019, which is in San Diego, November 15th through the 19th. His talk is debugging and profiling Spinnaker applications live, and his talk is on Saturday, November 16th at 3:45 p.m.

    I’d also to thank you—you, our listeners—for joining us today. This is Mitch Ashley with staging-devopsy.kinsta.cloud. Have a great day and be careful out there.

    — Mitchell Ashley

  • DevOps Chats: Spinnaker Extensibility, With Armory

    DevOps Chats: Spinnaker Extensibility, With Armory

    Spinnaker Summit 2019 Preview: Most open source software isn’t one size fits all, and Spinnaker is no different. While it captures the most general use cases of software delivery, most organizations find that it lacks some features they need to automate their entire delivery pipeline, such as integrations with internal tooling or compliance systems.

    Out of the box (figuratively), Spinnaker provides Run Job and Webhook stages that teams can use to build custom integrations to help fill any void and enable adoption by a wide range of users.

    Ethan Rogers, staff software engineer with Armory, joins us on DevOps Chat to share a preview of his Spinnaker Summit 2019 talk on Spinnaker extensibility. Ethan shares how Run Job and Webhook stages are used to build custom stages to capture any number of use cases with no code and simple configuration. He also covers situations when it’s necessary to write code.

    Ethan’s talk, “Making Spinnaker Your Own,” is on Saturday, Nov. 16, 3:45 PM, at Spinnaker Summit 2019 in San Diego. Separately, Ethan appears on the “State of the Kubernetes V2 Provider…” panel Friday, Nov. 15 at 2:30 PM.

    As usual, the streaming audio is immediately below, followed by the transcript of our conversation.

    Transcript

    Mitch Ashley: Hi, everyone. This is Mitch Ashley with staging-devopsy.kinsta.cloud and you’re listening to another DevOps Chat podcast. Today, I’m joined by Ethan Rogers. Ethan is a staff software engineer and also a speaker at Armory. He’s gonna be talking at the upcoming Spinnaker Summit 2019 in San Diego. Now, Ethan’s gonna be talking about making Spinnaker your own, that’s his talk, and then he’s also gonna be on a panel about the state of Kubernetes V2 Provider, plus improved rollout strategies for deploying Kubernetes using Spinnaker.

    So, let’s get started. Ethan, welcome to DevOps Chat.

    Ethan Rogers: Thank you so much for having me. I’m really excited to be here and get to talk about Spinnaker.

    Ashley: Wow. Absolutely. Great topic. I know you’re passionate about it, so fantastic to have you. Would you just introduce yourself? Tell us a little bit about you, what kinda work that you do and a little bit about Armory, the company that you work for?

    Rogers: Yes. So, like Mitch said, my name is Ethan Rogers, I’m a staff software engineer at Armory. I’ve been actually working on the Spinnaker project for about the last three years. I’ve been at Armory for a year and a half, and then a year and a half before that, I was working for my previous company where we were, you know, I was part of a team for delivery and developer experience. And, you know, we were using Kubernetes for this new project that we were working on and looking for a way to deploy and got involved in Spinnaker super early on. So, I’ve kind of been a part of the community ever since then and it’s been a really great experience so far.

    Yeah, so, that is actually kind of how I found Armory. Closer to the end of my time at my last job, I met the founders, DROdio, Ben and Isaac, and they had just—eh, within probably six to nine months started Armory. And we kinda hit it off. I had a really good relationship with the founders and it was kind of an opportunity to go keep working on Spinnaker full time, which is something I was interested in.

    So, Armory, just a little bit about us—we’re bringing Spinnaker to the enterprise. So, what we’re doing is taking open source Spinnaker and, you know, on top of offering stuff like support and training, we also offer layers on top of open source Spinnaker, various different extensions and features that make it more aligned with what the enterprise is looking for today.

    Ashley: Mm-hmm. Very good. Well, that explains, I know there’s yourself, but there’s also some other folks that are speaking at the conference.

    Now, your topic, you gave a nice introduction about how you got involved with Spinnaker—how long ago was that, when you started using it?

    Rogers: Yeah, it was about three years ago, kind of before the—like, right when the first version of the Kubernetes provider landed. I think, at the time, I didn’t realize how early we were. It kind of explains a lot of the PRs I had to submit when I first got involved.

    Ashley: I’ll bet, I’ll bet. Do you write any code, submit any code or are you primarily on the usage end of Spinnaker?

    Rogers: Yeah, so, I’m actually a developer by training. I got my degree in computer science, and transitioned pretty early on in my career into the ops side of things. So, I’m pretty fluent in both. I actually do write a lot of code. These days a lot less, just because of the role I’m in. At Armory, I do a lot more higher level stuff like BitTorrent, but I still try to get a few PRs submitted every week. I lead a couple of initiatives in the Kubernetes provider itself, so I do a little bit of everything.

    Ashley: Well, I won’t ask you which is the dark side, ops or developing—no, I’m just kidding. It’s DevOps, it’s all together now, right?

    Rogers: [Laughter] Yep.

    Ashley: Okay. So, let’s jump into what you’re gonna be talking about. And, of course, we’re not gonna cover it all here, but I know that you’re focusing on kinda how open source, you usually don’t just use it out of the box, if you wanna use that term. I know it’s free and downloaded.

    Rogers: Mm-hmm.

    Ashley: But that you have to do some customizations to kinda get it working in your environment, and we’re talking about Spinnaker here. What’s your experience working with customers? What’s their expectation when they start to use Spinnaker and what do you have to kinda help them get started with?

    Rogers: Yeah. So, the interesting thing about Spinnaker open source is that it’s actually, like, a very small part of what Spinnaker is at Netflix. For anyone listening who doesn’t know, Spinnaker was a project that was born out of Netflix. They built it, it was the next iteration in deployment technology from a project called Asgard.

    Ashley: Mm-hmm.

    Rogers: And one of the driving principles of Spinnaker was making it extendible. And the primary reason for that was because Netflix has a lot of kind of internal workflows and things that they like to—that make their development experience better, that don’t necessarily apply to the broader community.

    And so, when you start using open source Spinnaker kind of out of the box, there are a lot of things that I think most companies want, right? The same reason that Netflix has those, the things that don’t apply to the community but work for them, every other company has something like that.

    Ashley: Mm-hmm.

    Rogers: So, when I was at my previous company, we had this workflow around our Docker images where we would add labels to our Docker images and then we wanted those labels to show up in the Kubernetes manifest where those images were deployed.

    So, that was something that wasn’t available in Spinnaker out of the box, but we were able to use some of the extension methods to actually go grab that information and inject that stuff, which is kind of the basis of the talk I’m giving around Run Job.

    Ashley: Okay, great. Well, let’s explain that, then, because I know you’re gonna be talking about Run Job and about Webhook Stages. What is Run Job, what is Webhook?

    Rogers: Sure, yeah. So, just kind of, at a high level, Run Job and Webhook are ways that you can extend Spinnaker to kind of make it—to do the things that it doesn’t do out of the box that you need it to.

    Ashley: Mm-hmm.

    Rogers: There are more complicated ways to extend Spinnaker that involve actually taking the code and writing Java code, but we found that the majority of users or operators, I should probably say, are not super familiar with Java. So, that’s a pretty—it’s fraught with friction, I would say.

    Ashley: Yeah, it’s a whole other ball of wax, right?

    Rogers: Yeah, exactly. Yeah. So, the Run Job stage and the Webhook stage give kind of a lower friction way of extending Spinnaker. So, the Webhook stage is a stage within Spinnaker that lets you call out to some API. There’s the, you know, I select it as a stage in my pipeline via drop down and I configure it there. There’s also what we call preconfigured Webhooks or custom Webhooks that allow you to actually configure this as part of Spinnaker’s configuration, and then it shows up like it’s a native stage.

    Ashley: Mm-hmm.

    Rogers: So, under the hood, it’s still calling out to APIs. It’s as if you’ve configured it in line, but it shows up and, to the end user, actually looks like it’s part of Spinnaker itself.

    Ashley: Very nice. So, what are some of the most common uses of that to extend integrations into other things? Is it internal third party tools, internal custom software people who built, is there, you know, typically—the typical use cases for this?

    Rogers: I think that we usually end up seeing people integrating more with internal systems than, say, hosted services like GitHub or other—I don’t know, GitHub’s probably the easiest example. New Relic is another one that we see.

    But I think where, in most cases, Spinnaker has integrations for the public stuff out of the box, but obviously, if the tool’s internal to your company, you don’t have that. So, that is usually where we see it. So, I think there was one company that I had talked to that was, they had this internal integration testing suite or load testing suite that they wanted to call out to, and they used the Webhook stage to make that call and kick off those tests.

    Ashley: Mm-hmm. Great, great. Now, I understand you can do a lot of this without writing code, is that correct?

    Rogers: Yeah. Yeah.

    Ashley: So, what does that entail? What kind of a—what’s the configuration process look like to tie, let’s use that example that you gave of someone that wrote their own test framework of some kind that they wanna call out to in a Webhook stage?

    Rogers: Yeah. So, I guess I should probably back up and clarify my answer a little bit.

    Ashley: Okay.

    Rogers: So, at some level, you will need to write code. The question is whether or not you actually have to write that into Spinnaker itself.

    Ashley: Okay.

    Rogers: One of the, I guess we can call it native extensions, which is essentially using Spinnaker as a library and writing Java and doing some spring magic to make it happen.

    Ashley: Mm-hmm.

    Rogers: So, if you were to use a Webhook or a Run Job stage, for that matter, you would have to write a little code. But the benefit of that, if it’s a Webhook the service, whatever service that you’re trying to connect to can be written in whatever language that you want, right? So, if it’s Go, if it’s Python, if it’s Ruby, if it’s—you know, whatever you wanna do, the service is running in the API as the contract that you expect.

    Ashley: Mm-hmm.

    Rogers: And so, with the Webhook stage, all you’re doing is making a call. You know, you could argue that writing a cURL command is writing code, but I think a lot of people are familiar enough with that, that it’s a pretty low overhead.

    Ashley: Mm-hmm.

    Rogers: So, if you can make a call in your terminal, you can make a call from Spinnaker with the Webhook stage.

    The Run Job stage is actually a little more interesting, because what it allows you to do is run a container as if it were a script that you might be running. So, a really common pattern for doing custom logic in CI systems specifically is, I have a container, the container has all my dependencies and it has a script that I need to run. And we just run it like a one shot, you know—once this script finishes, if it fails, the pipeline fails. If it succeeds, the pipeline continues. And then, like, the world is really open, the world is your oyster at that point, because you can write the script in whatever you are most comfortable writing, and then you can run that within Spinnaker.

    One of the more interesting things that we did with the Run Job stage is that you can actually output—you can output these special markers in your logs and Spinnaker will go grab those and inject them back into the pipeline so that you can use them in downstream stages. So, maybe your job is going off and collecting some information for you. You can actually use that in, like, say, your deploy stages or a downstream Run Job stage or Webhook stage as, like, arguments you pass down.

    Ashley: It’s almost used like the log is like an asynchronous message bus, to make it more complicated than it is, but a super simple version of that.

    Rogers: Yeah, yeah.

    Ashley: Some down stage version where you could look at what happened before it and take appropriate action.

    Rogers: Yeah. Yeah, and that’s really—you know, we heard, the motivation kind of behind that feature is, we heard a lot of users, like, you know, they wanna be able to extend Spinnaker, but they don’t wanna write Java code, you know? And containers, for the Run Job stage specifically, containers are obviously super prolific in our industry today. I’ve built my career off of Kubernetes and containers, so.

    Ashley: Mm-hmm.

    Rogers: Containers are kind of like the perfect medium for being able to do this, because you can contain everything that you need—[Laughter] you can contain everything that you need in this container, and then, because Spinnaker can run containers in Kubernetes, it’s kind of like a natural fit, right?

    Ashley: Mm-hmm.

    Rogers: We can kind of take advantage of the fact that it doesn’t matter what your code is, the interface that we need is a container and then the logs, and we can get all of that information out.

    Ashley: I’m curious—this is totally a side question, but did you learn anything in college about containers and that kind of thing in software architectures, or is this something that you really started in after you graduated?

    Rogers: I wish I had learned something in college about this. So, I actually, I graduated college in 2013, and got my start writing PHP and doing a lot more software development. But in the back of my head, I’d always been like—I am very interested in servers.

    Ashley: Mm-hmm.

    Rogers: So, writing the code was interesting to me, but it always seemed like the people with all the power and all the really, people that I looked up to, they were the guys, like, in a shell, deploying software, checking, like, “Is this server overloaded?” It was very much not DevOps, but it was something that really interested me.

    Ashley: Sure, yeah.

    Rogers: And I—you know, I kind of, I spent probably three years as a software developer without having done that. But once I transitioned kind of into this CI/CD space, that really gave me an opportunity to focus more on ops, while doing some development and, you know, we picked up Kubernetes at my last job, and that’s where it all kind of kicked off.

    So, I think we started on Kubernetes 1.4, and that’s really where I learned everything I know about containers and orchestration platforms and stuff like that.

    Ashley: Well, it puts you in a really great position, because you can speak to both sides, both audiences, if you will, you know, developers and operations folks.

    Rogers: Yeah.

    Ashley: And, you know, my experiences before kinda containers and all of this, all the DevOps environment, some of the best developers were engineers who really knew how to write code, but they knew the full stack down to the operating system and the network, right? They kinda knew enough of all of it. Maybe there were really good parts of it, too, besides writing code.

    Rogers: Mm-hmm.

    Ashley: There’s a lot of, I don’t know, power, if you will, in understanding that full stack and how it works.

    Rogers: Yeah. I found that the more you understand about the stack, the better positioned you are to take advantage of it, right, and to actually write better code. You can go back and probably look at the birth of DevOps and see this, but engineers that spend the majority of their time writing code and then throwing it over the wall are more likely to write less performant code, right? Code that is not aware—you know, we wanna write really abstract code, but at some level, you have to be aware of what type of environment you’re running in.

    So, if you’re running on a machine that has 128 cores and hundreds of Gigs of RAM, it doesn’t really matter how inefficient your algorithm is, because it’s gonna run fast. You know, it all depends on scale, but … I’ve found that engineers that understand kind of the cloud and the environments they’re running in and what the actual, you know, where are your network boundaries and what are your actual constraints are writing better code, they’re shipping more frequently, they’re iterating—like I said, they’re iterating more quickly and that all of that is super important.

    I think the other thing that I’ve found is that it’s a growing process, right? I kind of like to look at it as—well, at least in my personal experience, I’ve found that I have a hard time grasping, kind of, the next level of what I’m doing until I’ve fully understood the previous level. And I think that’s really true for a lot of people, right? Until you’ve got a solid grasp of programming, it’s kind of hard for you to imagine what it would be like to understand ops, right? To understand how your program is actually running on servers.

    Ashley: Yeah. I would describe it as kinda context. It helps to know when you understand the environment of what you’re operating in and what you’re doing, what you’re building is going to operate in. And vice versa, right? If you’re in operations or security now, with DevSecOps, understanding now how software is being built. So, it’s the great thing about the accessibility of resources around the cloud and contemporary software architectures and all of that.

    Rogers: Yeah.

    Ashley: Well, we’re running out of time. I wish we had more time to invest in this, but the great thing is that you have your talk coming up, so that’s where you’re really gonna get into more details about that. And Ethan, I thank you for being on the podcast.

    Rogers: Of course, man. Thanks for having me.

    Ashley: Absolutely. Ethan Rogers, staff software engineer and also speaker from Armory, he’s speaking at Spinnaker Summit 2019, which is on November 15th through the 19th in San Diego. Now, he has a topic that we were just talking about, “Making Spinnaker Your Own,” that’s on Saturday, November 16th, 3:45. He’s also on the panel about the state of Kubernetes V2 provider. That is on Friday, the day before—the 15th at 2:30.

    So, thanks to you, all of our listeners, for joining us for this episode. This is Mitch Ashley with staging-devopsy.kinsta.cloud, and you’ve listened to another DevOps Chat podcast. Be careful out there.

    — Mitchell Ashley

  • DevOps Chats: Database for Cloud-Native Apps with MemSQL

    DevOps Chats: Database for Cloud-Native Apps with MemSQL

    Data and database technologies are experiencing disruption. The volume of data collected is exploding, creating new potential for disruptive services and new companies. Data is distributed across locations in the cloud and the data center. Also, cloud technologies bring a plethora of database types, distribution options, analytics capabilities and AI/ML processing. Any company not leveraging data to the benefit of its customers and business are just asking to be the next victim of disruption.

    Nikita Shamgunov, MemSQL co-CEO/co-founder, joins DevOps Chat to discuss how these disruptive factors changed how containerized and cloud-native applications approach data collection and storage. Speed, high volume data collection, distributed design and scale are all challenges Nikita sees as vital for cloud-native applications of today and the future.

    As usual, the streaming audio is immediately below, followed by the transcript of our conversation.

    Transcript

    Mitch Ashley: Hi, everyone, this is Mitch Ashley with staging-devopsy.kinsta.cloud, and you’re listening to another DevOps Chat podcast. Today, I’m joined by Nikita Shamgunov, co-CEO and co-founder of MemSQL. The topic today is really looking at the state of database and the market and technology as we think about database now. Nikita, welcome to DevOps Chat.

    Nikita Shamgunov: Thank you, thank you—excited to be here.

    Ashley: It’s a pleasure to have you on. Thanks for joining us. Would you start by introducing yourself, tell us a little bit about what you do and a little bit about MemSQL?

    Shamgunov: Absolutely. My name is Nikita Shamgunov, I’m the co-CEO and co-founder of a database company, MemSQL. Prior to MemSQL, I worked at Facebook, and prior to that at Microsoft on the SQL server kernel. So, I’m a kernel engineer by training. And before that, I graduated with the Ph.D. program from St. Petersburg, Russia.

    Ashley: Excellent, excellent. Well, let’s jump right into it. We were talking a little bit before we started the podcast recording. You know, databases have been around for a long time, and you know, the technologies created decades ago are still in production in many areas, but there’s a lot of changes that are happening, too. Why don’t you talk a little bit about some of the market forces that are happening to change the database environment?

    Shamgunov: Absolutely. So, there’s several market forces and industry forces that are, I would say, piercing through the market and technology. And the first one is that the data volumes are growing, right?

    Ashley: Mm-hmm.

    Shamgunov: And a lot of companies—and you don’t have to go far for examples, you know, companies like Uber and Netflix and Amazon and Google drive most of their value from data, right? That somebody in the ’80s would describe any of those companies as a glorified database application—but, of course, they are a lot more than just that.

    And with the data volumes and variety of that data growing, what separates winners from losers is how well people can capture and act on data.

    Ashley: Mm-hmm.

    Shamgunov: Right? A great example is, you know, when you’re hailing an Uber, you know, you pop it in your phone and then you know exactly where the cars are, how fast they’re gonna arrive, how long it’s gonna take you to the destination.

    And that’s really, really powerful. We’re all used to this right now, but 10 years ago, that was not a reality and not a possibility. So, that is the force number one, right? People see who are those companies that are advancing in the market and they wanna take on the same opportunity in their state, in their category.

    They also are afraid of being disrupted by the next Uber, the next Facebook, the next Amazon—particularly Amazon.

    Ashley: Exactly.

    Shamgunov: So, the second market force that is going through the market is architectural. And, because the data volumes are growing and the Moore’s Law is not working any more, so we cannot just rely on better, cheaper hardware, we have to build distributed systems.

    And, in the database world, distributed systems, for the most part, have been exotic until quite recently. And we’re starting to see distributed database systems—and the reality is, it’s just very hard to build a distributed kinda system of record transactional system and make it production ready, because the amount of technology that goes in there is enormous, and different, kinda other database products that are successful in the market have been products for 30 plus years. You know, thinking about Oracle, thinking about SQL Server.

    Ashley: Mm-hmm.

    Shamgunov: And when this massive architectural change comes in where it’s a business reality that you need to re-architect those systems, then it becomes very, very hard for the vendors to actually deliver those. So, that’s the second market force. Transition [Crosstalk]—

    Ashley: Let me ask you a little bit about that, Nikita, if I could, about the—

    Shamgunov: —towards distributed systems—go ahead.

    Ashley: The distributed architectures, do you see that happening largely because of the displacement between on-prem data centers and private clouds and public clouds? Is it more pushing content to the edge and doing caching? Is it pushing content toward a geographical location? What’s forcing the distribution of databases?

    Shamgunov: This is a wonderful question, first and foremost, because I was about to say that the third mega market force that is happening right now is moved to the cloud.

    Ashley: Mm-hmm.

    Shamgunov: Right, and cloud, if you make it parallel with electricity, right, the majority of the world consumes electricity in the clouds, you know? You just plug your stuff into a socket and here comes your electricity.

    Ashley: Mm-hmm.

    Shamgunov: On premises, generators still exist at important places—manufacturing, health care, hospitals—but it’s a tiny percentage of the global electricity consumption.

    Similarly, it just makes too much sense to consume IT in the cloud, and within IT, you know, there’s applications, analytics, databases, machine learning—you know, all the whole spectrum of services that typically a cloud provider now offers. Now, if we buy into this worldview—and certainly, we see what happened with electricity, and this is the same thing that’s happening in IT—then you can split the same into different workloads. And once you zoom in on the database workloads, you will realize that, yes, you wanna take workloads that currently run in data centers and shift them into the cloud.

    Cloud is different, you know? The most important thing in the cloud is that software just runs as a service. But also, it creates an opportunity to rebuild every single piece of IT for the cloud. And that’s already happening. In our space, it’s already happened in data warehousing, but it also happened with identity because of Okta. It happened with cloud storage, with services like S3 on Amazon or Dropbox as a higher level service. So, basically, as infrastructure IT is shifting towards the cloud, we are rebuilding it for the cloud. And now, it just so happens that, through that rebuild, you want to cater to the next generation workloads. That’s where distributed systems comes in. And also, you’re architecting for the cloud because the cost equation in the cloud is different.

    So, being scalable and elastic allows you to provide a better cost equation, as opposed to legacy architectures. So, if you take an Oracle and—yeah, you can run it in the cloud, in EDM, no problem. Well, actually, it’s gonna be way too slow and way too expensive as opposed to the system that’s architected specifically for the cloud that scales both for cost, right, so you can only—you will only consume as many resources as you need from the storage and compute standpoint. Storage compute network bandwidth—whatever you need to run the workload. But it also is gonna be a lot more efficient, because running a database as EDM, everybody will tell you this is not such a great idea, but if it’s a database service architected for the cloud, then you can build for and around—around limitations and for all the services and resources that are available for the cloud.

    Ashley: What I’d like to ask you about that is, it seems like there’s a parallel between lift and shift for applications and lift and shift for database. It isn’t architected for the cloud if you’re just taking the application into a cloud environment, but you’ve gotta re-architect, rebuild, take advantage of things like database as a service, for example. Is that correct?

    Shamgunov: Absolutely. Alright, and then in the application space, what’s different than the cloud is that certain applications can be more popular than others. And if you just put those apps as you had on premises in VMs—so, very likely, you both overprovision and underprovision resources.

    For less popular applications you are wasting the whole VM to run an app, and for a very popular application you don’t have enough network bandwidth or CPU to run this app at scale. And that’s why Kubernetes really changed the game, and if you build against Kubernetes, you can pre-release scaled, stateless applications. And then, if you run it in the cloud, of course, the database that you use to power those applications, you consume as a service.

    Ashley: Say a little bit more about that, because we’ve talked on other webinars and podcasts about the statefulness of data versus the more dynamic nature of building software and rebuilding software that may not match up with the state of the database. How do you handle that in a Kubernetes environment?

    Shamgunov: So, there are two pieces to this. The first one is, you run your application in the Kubernetes environment and you consume data through a database as a service. So, that is a blueprint of a cloud architecture today. Let’s say you run your Kubernetes on Amazon, your application is written in whatever language, but let’s say Node.js, and you fire up DynamoDB, using an Amazon example, as a service, you connect to that database from the application and there it is, there’s your app.

    Now, in the world of relational databases—which, by the way, is a much larger slice of the market than object databases—

    Ashley: Sure, yeah.

    Shamgunov: And for a very good reason. You can run it in a similar fashion. You know, in this case, you can launch MemSQL as a service, you know, run  your application inside a Kubernetes cluster.

    The other thing that you can do is, you can bring MemSQL with you into the Kubernetes environment. And that gives you control and cloud portability. Now, not only do you deploy your application into Kubernetes and the application scales depending on the demands and the popularity of the app, but also the database that stores the data and powers the app can be deployed in the same Kubernetes cluster, and it will do the same thing. It will scale or contract, depending on the needs of your application.

    So, what it gives you, it gives you cloud portability. So, you can pack your bags and go from AWS to Google, for example, and the other thing it allows you to do is to take it and bring it down, right, from cloud to on premises. And today, the reason to do so is usually twofold. It’s either security and compliance or it’s cost, right? If the workloads are pretty static and you cannot drive a good amount of efficiency through the elasticity of your both application and data infrastructure, then on premises may still be cheaper, and then that creates an incentive for you to bring it down on the ground.

    Ashley: How important is it to have kind of a single tool to do ETL versus separating those as you move not only within one cloud but across clouds?

    Shamgunov: Well, I think, at the end of the day, we are talking about lock in, and we’re talking about developer productivity. Those are completely different. Now, lock in is something that a lot of people in our space experience with Oracle as a vendor.

    Ashley: Sure.

    Shamgunov: Where they go through what they describe as extortion cycles. Every time an Oracle renewal comes in, then you end up negotiating with Oracle, and if you wanna reduce the amount of the database spent, it becomes a very difficult conversation. You know, Oracle runs an audit, which they have the right to, so they inspect and find all the applications that consume Oracle. If you reduce the amount of consumption, Oracle jacks up the maintenance fee or jacks up the next license fee.

    So, it’s really, really hard to get out of spending a lot of money on this vendor. We hear that time and time again in major, major enterprises. I had a conversation with the CIO of a major bank, and the Oracle spend there is $1 billion over several years.

    Ashley: Not surprising, yeah.

    Shamgunov: And this is absolutely insane, if you think about it, right? The reason for that is that Oracle has been a great partner for delivering their services, you know, cost aside. You know, banking runs—in their bank, it’s a consumer bank, and it runs on Oracle.

    But at the same time, competitors catch up to the functionality, right? MemSQL comes with a very strong system of record capabilities and, to this day, that has been the stronghold of Oracle, and you know, with a much  more attractive price, with a much more flexible deployment model and the ability to run in the cloud as a database as a service where we guarantee SLAs enough time. And with an integration with AI and ML, we see that, slowly but surely, that equation is shifting. And, you know, people seriously ask questions, “Why would we spend so much money on this technology and can we not do this? We would save hundreds and hundreds of millions of dollars and drive shareholder value, here. In addition to that, we’ll have our development teams move a lot faster,” right?

    So, back to your question about across clouds and developer productivity, right? The first one is lock in, and I gave you an example of an Oracle lock in. But now, the CIOs don’t want to fall into the same trap by getting married to a single cloud provider. And, you know, thank God, we have multiple cloud providers available on the market, of which three in the United States are very, very strong, and you know, for the most part, similar in capabilities—GTP, Azure, and AWS.

    So, it’s very typical for a key enterprise to choose two or three cloud vendors. So, that prevents the lock in. However, in order to truly avoid lock in, you’ve gotta choose service providers that allow to run their software the same way on all the clouds.

    Ashley: Exactly.

    Shamgunov: And ideally, they deliver their services as a path service, you know, like a database as a service, and that service is available on every cloud. And if that’s the case, then, you know, you take a dependency on a vendor, but you are not taking a dependency on the mega vendor, which is one of the clouds. So, that is something that is top of mind of every CIO that I know.

    Ashley: You make a good point about Oracle, kinda going back to the lock in of the 2000s and maybe ‘90s. You know, for VLDB, that was really the best game in town, arguably. You could also Microsoft, maybe.

    Shamgunov: Mm-hmm.

    Ashley: But that isn’t true today. You’ve got many options when you go to cloud providers and folks like yourselves, I would imagine, live in all of those cloud environments or multiple of them. So, you can also serve customers that, of course, are gonna have a multi-cloud strategy.

    Shamgunov: Absolutely, you know, I couldn’t have said it better.

    Ashley: Okay, great. [Laughter] Lock that one in—great. What’s the number one challenge as folks move out of—let’s pick Oracle again—move out of that traditional VLDB environment and move into the cloud. What’s the paradigm shift that you have to make to really think about how you can leverage the flexibility of the cloud?

    Shamgunov: There is a three step strategy if you wanna get out of Oracle. The biggest stickiness of any database technology is all the applications that are built against the database and they’re using and sharing that data across the applications. And then different database technologies can either be app compatible—you know, an example of that is MySQL and AWS Aurora, right?

    Ashley: Mm-hmm.

    Shamgunov: Aurora is a new database, but the compute piece of it is actually MySQL code, so they’re app compatible. If an app works against MySQL, it’s gonna work against Aurora.

    So, that’s why, in the low end of the market, in which both MySQL and Aurora play, this can be a viable strategy. And it’s a very good one, right, because it’s so easy to just point your app at a different database. And it’s very hard for a different database product to achieve application level compatibility between the two databases.

    The second one is workload compatibility. So, all our workloads that are running on Oracle can be moved to another database, assuming there is a good enough reason. You know, assuming the database you’re moving the workload to has all the same set of features and provides strong system of record guarantees and so on and so forth.

    So, if you are workload compatible, then you can move from one database into another, but you need to put in some work, right? You need to augment and potentially rewrite certain pieces of the application.

    Ashley: Mm-hmm.

    Shamgunov: And, in our space, there’s a massive piece of the market that is called tier one workloads. Those tier one workloads typically run on Oracle, and they have either very, very strong system of record requirements or they have performance requirements, or they have availability requirements, or they have very strong concurrency requirements.

    And because of our architecture, we are practically the only game in town that can be the destination for moving workloads from Oracle, and especially systems like Oracle—you know, heavy Oracle systems running on enterprise storage or systems like Oracle Rack or Oracle Exadata over to running on MemSQL. And we have countless examples of how we quote-unquote liberated our customers from Oracle.

    So—and this is, frankly, our biggest opportunity. So you will have to put some work, but if the pain that you’re experiencing is so high that you’re willing to put that work, we are a fantastic destination for moving your Oracle workloads.

    Ashley: Well, excellent. You’ve been a great guest and a fascinating conversation. I wish we had more time to chat more. Maybe we can do that on another podcast.

    Well, thank you, and with my audience, I’d also like to thank you, Nikita Shamgunov, for joining us, co-CEO and co-founder of MemSQL.

    Shamgunov: Yeah, very happy to be here. Thank you so much.

    Ashley: You bet. And thank you to our listeners. This is Mitch Ashley with staging-devopsy.kinsta.cloud and you’ve listened to another DevOps Chat podcast. Be careful out there.

    — Mitchell Ashley