Author: Uri Cohen

  • Advanced Automation – Getting Your Systems to Work for You

    Advanced Automation – Getting Your Systems to Work for You

    In my previous post I discussed how to take your DevOps to the next level by taking it beyond infrastructure automation, to the automation of your deployments and code pushes, through patches and updates.

    And then I promised to make it interesting…so here is the next stage – actually using the extracted data to get your systems to work for you.

    Monitoring Like IT’S YOUR BUSINESS

    Let’s start by discussing what it means to monitor your application, and what kind of data you can extract from it.

    At the most basic level, you want to know whether your application is available for your users. It sounds very basic and simple, but setting it up properly is actually not an easy task. It is intended to give you an answer to the most important question to your business: can my users use the application? Think of it as a big red or green light on your dashboard. Although the answer may seem obvious and simple, but it is not at all. There are all kinds of parameters to take into account when answering this question, for example, what if the system is apparently up and running but the response time is very slow, or what if only parts of the system are actually running. At the end of day, you want to have that kind of an alert “traffic light” that tells you whether your system is functioning – yes/no, up/down.

    System, Application, and Business Metrics

    The green-red traffic light is an aggregation of multiple levels of monitoring. The most basic one is system related. It contains indicators like process availability, CPU and memory utilization, etc.

    The next level of monitoring is application related and is composed of application level KPIs (key performance indicators). KPIs can be anything from the average response time for a user for a certain request, to the number of concurrent database connections. In other words, anything that’s specific to the application and its architecture.

    And then above these, comes the highest level of metrics, which are business metrics. Ultimately these are what’s really important. If your business metrics are ok, then you can assume everything is running fine. Sample business metrics include the rate of failure to register to your website, how many users did a certain operation in the website or how many users executed a specific transaction. If we want to be really specific, then take Google hangout for example, a business metric of theirs would be how many people in a certain timeframe joined or started a hangout.

    Logs to the Rescue

    Logs are often used here as means to collect your metrics and alert about erroneous conditions. Logs usually contain a lot of useful information, but you need to be able to extract it and make sense of it. This means gathering all of the logs emitted from all of your servers, parsing them and looking for specific patterns that will help you generate the KPIs over time. For example, you look for a log message that says “User X signed in”, and by counting these messages you can produce a business level metric for how many users register over time. If you add some more data to the log message, such as the user’s location, or operating system, or anything else for that matter, you can slice and dice data and get deeper insights into how your application is behaving. When logs are emitted in a structured and consistent way (e.g. in a JSON format), it becomes very easy to analyze them and produce the relevant KPIs. There are many toolsthat can help with that.

    From Simple Automation to Orchestration

    That’s where orchestration comes into the picture. At the most basic level, orchestration is a higher form of automation, which helps you setup all the pieces that are related to your application, starting from the infrastructure (VMs, networks, block storage volumes, security groups, etc.), to the platforms your app runs on (database, web server, etc.), and all the way up to the application modules and code. This entire setup is often referred to as a topology. The role of an orchestration framework is to materialize a certain topology. More advanced orchestrators go beyond materializing the topology, and change it to meet the current workloads and needs of the application.

    As you probably gathered by now, monitoring and log gathering are an essential part of any running app, so they should be an integral part of the orchestrated application topology. Since the orchestration process is topology aware, it can wire and configure monitoring for your application components very easily, which is one of its greatest benefits. Going through these process without a global view of the topology can often be time consuming and error prone. Moreover, as the topology changes , you need to reconfigure and rewire your monitoring tools, and a good orchestrator will that for you as well.

    Next Up: Reactive and Proactive Orchestration, The Devops Holy Grail

    Hopefully by now you see the value of orchestration as a higher form of automation. In the next post I’ll dive deeper into the post-setup phase, and discuss how orchestration tools can react to events and monitoring data and adjust the application’s runtime topology to best fit the current workloads.

  • From Simple Automation to DevOps Like a Boss

    From Simple Automation to DevOps Like a Boss

    Automation as We Know It Today

    If you haven’t been hiding in a cave in the last year or two, you’ve probably heard the terms DevOps and infrastructure automation more than once… But even today, infrastructure automation is mostly focused on setup and deployment of complex systems. For example, if you’d like to deploy your application to the cloud, you would likely automate the steps of provisioning the cloud resources, installing the right components on top of these, and then orchestrating the startup of your components – or better known as – cloud orchestration. Take even the simplest application that has a web server and database. After installing and configuring everything, you’d first need to ensure that the database is started, and only then the web server. You’d also need to propagate specific runtime information from the database to the web server, such as the database’s host and port. This stage, for the most part, is where most automation processes focus on today.

    DevOps: Stage Two

    Automation, however, can do much more for you. Let’s imagine, following your initial automation of the setup and deployment of our your application, that you want to make changes to your infrastructure or application. For example if you have a patch you’d like to install to the database or operating system, or if you want to update a piece of your code on your cluster of web servers. Both of these concerns involve pretty complex processes that you need to automate. For an infrastructure upgrade, you would typically roll out the upgrade server by server, take down one server at a time, update it, then restart it. Even though it doesn’t happen very often, there’s still a lot of benefit in automating this process and minimizing the chances of human errors. Pushing new code is a bit different. For starters, you do it much more often, even a few times a day. Also, (no offence), it is likely that the stability and quality of your code is lower than that of an approved patch to the OS or database. So the risk of something going wrong if you don’t automate is even higher with this process. That’s why there are common automation strategies that help you deal with this risk. Two such strategies are the Canary Instance strategy and the blue/green strategy. The former derives its name from the mining days, when a canary bird was sent out into the mine before the miners, to check if the air was toxic or not. So in this fashion, a “canary instance” is chosen, the new code is pushed to that instance alone, and then a number of sanity tests are performed. If these pass, the code is then pushed to a few more instances. If all goes well again, the code is then sent to the rest of the infrastructure. In the latter strategy (A/B), the code is sent to only a part of the servers. In an even more sophisticated approach, the code is only enabled to a portion of the users, for example, a certain geography or time of day. All of this ties into how to setup and manage an ongoing system. But although this process, as well the infrastructure upgrade, can get quite complex, especially when dealing with large scale, live systems, in most organization code pushes and upgrades are still not automated. But there are more than a few companies that are blazing the trail for the rest of us.

    The Plot Thickens – Monitoring Like a Boss

    So once you finally saw the light, put in your sweat and tears and managed to automate your setup, upgrades and code pushes. What’s next? This where it really becomes interesting, and where you can really take your business to the next level, with the monitoring of the entire system, and eventually, what you do with the aggregated data.

    More to come…

    In my next post I’ll discuss the different levels of monitoring and logging to get the to promised land of proactive DevOps.  And then…take that even further.  So stay tuned.

  • Changing Organizational Culture – A Sweaty Use Case

    Changing Organizational Culture – A Sweaty Use Case

    When we think about changing organizational culture our immediate reaction is usually – yeah right, go fight city hall – and we take the Homer Simpson route, and don’t even bother trying.

    Over the course of the last year, I managed to achieve what felt like an impossible change within my organization, and I took some valuable lessons from the experience.

    It dawned on me afterwards that there were a lot of parallels in this process that we also underwent when trying to instill a DevOps culture in our organization, and I thought I’d share what I learned.

    The Epiphany

     I love riding my bikes – mostly mountain biking, and I usually am only able to do this over the weekend.  It’s something I really love to do, but in addition to my full time day job – Head of Product at my company, I’m a full time dad – a proud father of three.  All of this doesn’t leave me a lot of time to do the things I love.

    On a good day I spend more than an hour stuck in traffic, and on bad days more than an hour and a half.  This is basically dead wasted time.

    So I had an epiphany one day. Why shouldn’t I take advantage of this time to my favorite thing?  I could easily bike to work – and use this wasted time to my benefit, doing the thing I love most.  The problem lies in the way I look after an hour bike ride.  Pretty sweaty.  Here’s a pic.

    And, I know what you’re thinking – the answer is – no it wasn’t raining that day, in case you were wondering.

    The challenge though, is that I work with people, so I needed to find a way to clean up after my rides.  But, alas alack, we don’t have a shower in the office.  The plot thickens.

    So you might be wondering at this point, how this all relates to DevOps.

    DevOps is about pushing change in organizations – and not necessarily just technical ones. There are cultural organizational practices that often times need to be changed before you can even start to think about technical changes.

    So I believe this story represents a real parallel case study on the organizational changes often needed, to create a DevOps culture within an organization.  These processes many times involve a change in hardware – in this case, the shower itself.  As well as require a change in culture, actually motivating the management to realize the importance of it, and have people start choosing to do athletics before work (and preferably also ride to work).

    The Method

    And then I ran into Jesse’s Robbins’ rules. For those who don’t know, Jesse is one of the founding fathers of DevOps and web operations, and happens to also be the founder of Chef (OpsCode at the time). In one of the recent CloudExpo conferences in Europe, Jesse gave an awesome keynote about how to hack organizational culture to drive change. So working on the premise of “Jesse’s Rules” – I got to work:

    • #1 – Start small – Build trust

    Until the new shower was built, I’d shower daily with a HelioPressure shower  (you can find it on Amazon for less than $100), and I’d shower every day on Nati Shalom’s [LINK TO DEVOPS BLOG] balcony in the office.  Demonstrating I was serious and dedicated to my biking to work. Yes, it definitely raised a few eyebrows, but then again, better to shower on a balcony than walk around smelling like a sweaty rug all day…

    • #2 – Create Champions

    Doing crazy stuff yourself is the easy part. Finding like-minded individuals that will do the same, well that’s a different story. But after working hard at convincing others, I managed to find at least two people that would occasionally use my improvised spa, and this really helped me drive change in my company.

    • #3 – Use Metrics to Build Confidence

    For me there were actually a few metrics that contributed to the driver for change.  The first was that, I right off the bat, lost quite a few pounds – which was an incentive for a number of people to jump on the bandwagon – or rather bike – in this context.  The next metric was hard research that shows that people that are in shape tend to be absent from work 80% less often than employees that are not in shape. That helped me in convincing my CEO that it’s actually an investment and not just thoughtless spending…

    • #4 – Celebrate Success

    After committing to our makeshift shower for months, budget was officially allocated, and a shower was built for all the employees looking to participate in sport activity on their way to work.  Needless to say that I’m not the only one riding my bikes to work today 🙂

    Lessons Learned

    So the lessons I learned throughout the process, and my key take aways as I see them are:

    •  #1 You’ll run into a lot of skeptics along the way. IGNORE THEM.

    Many people will attempt to discourage you from your goals along the way – don’t listen to them. If I had a dollar for the amount of times I heard “it’s never gonna happen”, I wouldn’t need company budget to build a shower.  But I stuck with the plan, and it ultimately paid off.

    •  #2 Change takes time

    Rome wasn’t built in a day.  It’s a cliche – but cliches are born from common experience.  We’d been discussing building a shower for 2-3 years at the least.  It’s true that until I didn’t demonstrate that I was truly dedicated, rule #1, it wasn’t taken seriously. It took time for the company to come around, and digest the idea, until it gradually happened.   So again, if you stick to your guns, and have faith – a change ‘gone come.

    •  #3 Take it one small step at a time

    Basically start small, and be creative. Until you can achieve the big results, celebrate the small achievements too.  Adding champions to your cause, increasing awareness, possibly a meeting to discuss the changes you propose.

    •  #4 Be prepared to improvise

    I owe a lot to my colleague that came up with the pressure shower idea.  I think until I actually demonstrated that there was a really dire need for such a change, it wasn’t really considered seriously.

    •  #5 Find a management sponsor

    Luckily our CEO is athletic in orientation, and is an avid bike rider as well, and he was committed to help making this happen.  Many times it won’t necessarily be the CEO, but find a management champion that believes in your cause, and is willing to help you bring about the change.

     Bottom Line

    No matter what change you’re trying to push, these lessons and rules have proven to hold true for me in any circumstances, and I often times use this method to push significant changes in work and in life.

    Oh, and here’s a pic of the new shower.

    If you’re ever in Israel, and need a place to work, and want to bike there…you’re welcome.