Tag: bugs

  • Applause Report Surfaces Functional Testing Issues

    Applause Report Surfaces Functional Testing Issues

    An analysis of more than 340,000 bugs collected by Applause, a provider of a testing platform designed to integrate with a range of DevOps platforms, found functional bugs accounted for 68% of issues. The research gathered data from 13,000 mobile devices and 1,000 unique desktops running 500 versions of operating systems and found that the majority of issues could be traced to functional bugs compared to visual (17%), content (9%), crashes (4%) and lags and latency (2%) issues that, collectively, only add up to 32%, the report found.

    The report also found screen readers comprised 66% of all accessibility bugs compared to keyboard navigation issues and insufficient color contrast which accounted for only 12%. In terms of localization, poor and missing translations account for more than two-thirds (67%) of bugs, the report found. Based on ongoing feedback data that Applause also continuously collects from customers, nearly half of organizations (47%) identified currency and number formatting as the most valuable bugs to identify when it came to localization.

    Overall, organizations ranked the discovery of crashes (75%), functional bugs (61%) and lag and latency issues (53%) as exceptionally valuable, according to Applause.

    Luke Damien, chief growth officer for Applause, said when it comes to testing, in general, there is still not enough focus on user and customer journeys. That lack of focus is becoming a larger issue as more organizations invest in digital business transformation initiatives that are driven largely by mobile applications, he noted. Payments are especially problematic because they often rely on application programming interfaces (APIs) exposed by a third party in a way that results in suboptimal application experiences that have a direct impact on revenues, he added.

    The challenge is that it’s not entirely clear how far left responsibility for application testing is shifting. In some cases, developers are assuming responsibility. In other cases, a dedicated testing team is still responsible. Developers, of course, will test applications as they build them but application experiences on a local machine may not always be replicated in a production environment. The issue is that most end users today are not especially forgiving when a mobile application fails to meet expectations. Months of development effort can be wasted simply because a function was not tested in a production environment.

    Less clear is to what degree testing will be automated in the future. Machine learning algorithms are making it possible to automate testing of both functions and user interfaces. It’s not likely the need for humans to be involved in testing will be eliminated any time soon, but many of the routine tedious tasks that conspire to limit the rate at which applications can be tested should be sharply reduced in the months and years ahead.

    In the meantime, an organization’s entire brand reputation is now tied to the quality of the application experience it enables on a mobile device. Revenue targets can now be easily missed if users of a mobile device choose one application over another simply because a function was broken—even for a short amount of time. As such, testing has never been more critical given organizations’ dependency on software.

  • Rookout Extends Reach of Debugging Platform to On-Prem

    Rookout Extends Reach of Debugging Platform to On-Prem

    Rookout, a provider of a software-as-a-service (SaaS) platform designed to simplify the debugging of applications, has extended the reach of its tools into on-premises IT environments.

    Company CTO Liran Haimovitch said its new Data On-Prem capability extends the reach of a bytecode manipulation capability developed by Rookout to insert a snapshot capture of a specific line of code to an on-premises IT environment. This eliminates the need for DevOps teams to upload code and associated data to the Rookout platform manually.

    The Data On-Prem capability is made available via a software development kit (SDK) provided by Rookout, he said.

    Legacy approaches to debugging code are essentially broken because the processes employed to fix code are too cumbersome, said Haimovitch. Rookout is designed to collect full-stack data without requiring code to stop running or impact code execution. That approach allows developers to debug live code on the Rookout platform and then reinsert the code into their applications.

    Haimovitch said Rookout is making a case for transforming how application debugging is addressed within a DevOps process. Rather than relying on tools that need to be tightly integrated and maintained within a continuous integration/continuous deployment (CI/CD) pipeline, it will be easier for developers to debug live code on the Rookout platform, said Haimovitch.

    That approach will enable DevOps teams to go well beyond simple observability to embrace true understandability of how applications work down to individual lines of code, he added.

    Debugging applications has always been problematic. In an ideal world, developers would have a much better handle on what bugs need to be fixed earlier rather than later. By making it easier to debug live code, the number of issues that developers could address before an application is deployed in a production environment should increase. It may never be possible to address every issue before an application is deployed; however, the number of critical issues theoretically should decline. Many code issues go unaddressed simply because the effort required to debug that code is too great. Many developers too often put off fixing those issues until the next update usually because they are rushing to meet a delivery deadline.

    At the very least, developers should repeat the same mistakes less often as they work to debug live code.

    Regardless of how the debugging issue is addressed, issues within code in the age of microservices are starting to have a bigger impact. An issue within a line of code in a single microservice can have a cascading impact on all other elements of an application that depend on that microservice.

    Debugging application code is never going to be enjoyable. The challenge is finding a way to make the entire process less painful. There may come a day soon when machine and deep learning algorithms help automate much of the debugging process. In the meantime, DevOps teams might want to re-evaluate existing approaches to debugging code that clearly leave much to be desired.

  • The Code Doesn’t Lie, and Other Operations Mantras

    The Code Doesn’t Lie, and Other Operations Mantras

    As engineers, we spend a lot of time talking about things such as release processes, QA environments and deployments. But at the end of the day, most software systems fail because the software itself is faulty. The code doesn’t lie—you can almost always find the solution to your problem within the code itself. Of course, looking to the code for answers is a time-consuming process if you’ve deployed too much in one batch. After all, sometimes the tiniest mistake can derail a whole string of code, and having to sort through a haystack of code looking for a tiny needle is not fun.

    There are a few corollaries to this mantra that “the code doesn’t lie”:

    • “Know your data like you know yourself.” Sometimes the quality of the data itself can cause a system to behave differently.
    • “A feature is a bug in a tuxedo,” which is largely for you product folks out there.
    • “Where there is one bug, there are many.” When you find a problem and it leads to a systemic set of problems, it may seem terrible at first, but the reality is, you have the opportunity to fix a lot in a short amount of time.

    I had a chemistry teacher who used to say, “A mole is a mole is a mole,” referring to Avogadro’s number. My version is: “A bug is a bug is a bug.” A feature that behaves incorrectly (even if built to spec) is a bug. A badly performing system is a bug. A transient failure is a bug. A race condition is a bug. No matter what they look like, it’s our job to squash the bugs.

    Branch Readiness is a Joke

    A very talented engineering director once made the comment, “Branch readiness is a joke,” about the branch/development model. As 23 different teams were preparing their branches for integration onto the main trunk, code changes were flying around with very little testing/certification. We even had code that would not compile.

    When all 23 branches hit the main release branch, all hell broke loose. And this is where the actual functional, integration and performance testing began. The key problem with this model is that integration and certification happened late in the game. The other problem was that we did not abide by our own criteria of what constituted branch readiness; instead, we were constantly softening the restrictions until bad code made its way into the system.

    We moved to a trunk-based model of development where check-ins are made directly to the trunk, and all code must compile and be pretested. Continuous integration and testing are run against the trunk all day and all night. We can ship from the trunk at any moment.

    Now you may be asking yourself, What does this have to do with operations? The answer is: everything. We can control what changes get to the site. We can certify what changes get to the site. We are assured at every moment of a minimum quality in the changes that do get to the site. These smaller changes give us the ability to move bits onto the site more frequently. We are not forced into large-scale, monolithic deployments, which inherently create more risk to the site itself, often with little ability to roll back.

    Learning from ‘Branch Readiness is a Joke’

    Just because the code doesn’t lie does not mean that you can always understand all of its implications. The key is limiting the amount of code being changed in any individual release to an amount an engineer or automated system can understand. By avoiding large, monolithic releases containing thousands of changes and instead testing, verifying and considering each change individually, we can understand what the code is telling us. We can release to the trunk with confidence that what we build and deploy is going to work.

    The One-Character Change

    We live in a world of bits and bytes (eight bits per byte). One byte is the equivalent of one character. In some cases, a one-character change is benign. For example, changing the number of members of a site on an informational page from 134M to 135M (4 to 5) will not cause any harm to the site.

    But sometimes, a one-character change can be disastrous. Consider a DNS change that has a one-character mistake for www.yourcompany.com. Get it wrong and you are off the air. It is very important to understand the impact of a change going awry. If it can cause a large impact, then we need to ensure we understand the change and have a clear plan to roll it back if we need to.

    Here’s a great example of a very small change going sideways very fast: We once scheduled a maintenance to test routing traffic through a European point-of-presence. Unfortunately, the time-to-live parameter was set to a large number (hours instead of minutes) on the DNS entry. The result was that the change, once implemented, couldn’t be fully undone for hours. It was a one-character change with a bad outcome.

    Learning from the One-Character Change

    The code doesn’t lie, but that only matters if we are paying attention. Put simply, not all changes are equal. By taking the time to think about what a change is intended to accomplish and the various ways it could go wrong, we will always be better off than if we didn’t do due diligence. The one-character change can be benign, or it can be catastrophic. It’s important to figure out the possible ramifications of any changes (no matter how small they may seem) before proceeding.

    Operations tends to be the first group to get called after hours when something goes wrong. As a result, we have a vested interest in quality of code shipping to our site. By working to improve the code quality before it hits production, you can significantly reduce the number of problems you encounter.

    Start with where the code gets committed into source control. For each change made ask the simple questions: Does the code still compile? Does your application still build correctly? Does it still work as intended? By automating the process of asking these questions for every change, we gain the ability to deploy new versions of our application based on any commit.

    Once we have good code quality making it into source control, we can begin making further improvements through the use of a canary process and start focusing on things such as performance. Remember, the code doesn’t lie. If we can understand each change, then we can start preventing problems before they begin.

    This post is part of the series “Every Day Is Monday in Operations.” Throughout this series we discuss our challenges, share our war stories and walk through the learning we’ve gained as Operations leaders. You can read the introduction and find links to the rest of the series here.

    About the Author / David Henke

    David Henke has more than 35 years of experience working in technology, including senior engineering leadership positions at LinkedIn, Yahoo!, and AltaVista Company. He’s also been a founder at two different software companies, both of which were acquired. Currently, David serves in a variety of board and advisory positions with organizations like NerdWallet and UC Santa Barbara.

  • Communication: Who is to blame for misunderstandings?

    Communication: Who is to blame for misunderstandings?

    Communication is everywhere and it is far more than only the spoken word and its origin meaning. Communication is one of the most complex and also complicated topics in human relationships: most of the time very thrilling, sometimes frightening but always important!

    As we all grew up with communication, you might think, we should be used to communicating very well. But daily life shows us that misunderstanding is the companion of communication. We all know from our own experience how hard it is to clear up such misunderstandings and how much precious time is wasted on them!

    The more people with different personalities are trying to communicate with each other, the higher the probability of misunderstandings.

    To take a closer look at all those cross-functional teams like devops or agile scrum teams you will recognize at once that there are lots of individual issues because of misunderstandings in every team. Each single issue on its own is disrupting daily work, something we all know from personal experience. Now try to imagine how much of a team´s efficiency is taken away by the sum of all issues caused by bad communication and misunderstandings! It is unimaginable!

    Communication is the link between all stakeholders: it determines the atmosphere in a team, the relationship between all parts, the understanding as well as the knowledge and information transfer. It influences the project times, the employees´ identification with the company or their work, individual responsibility and how to deal with mistakes. This list could go on and on.

    Time in IT projects is one of the limiting factors. To save time and avoid mistakes, new tools are be developed again and again. Companies invest money in these tools; sometimes tools are integrated in the teams, sometimes the teams are arranged around the tools or systems. But the very human part – the communication – which is important for every team, no matter which tools or which systems are used, is seldom the focus of IT education! You can changes methodologies, systems and tools, but if you do not get beyond bad communication, you won´t reach the full potentials of all those efforts. No matter if you have already chosen your tools or systems or if you are still searching for the right way for your team, you should always think about the communication within your team. Instead of always asking what the ROI of communication courses will be, you should better ask what the ROI of all those tools and systems could be with optimized communication!

    Before optimizing the communication you first have to find the bugs in the individual and in the team communication and eliminate them. Due to the inherent complexity, the support of an expert is sometimes useful. Just as not all people using computers are IT experts, not all people are experts in communication – even though they are communicating the whole time, sometimes even by saying nothing. Some are more gifted than others, some are able to learn by studying on their own and others need a little more time and support. Changing communication mostly means to change personal habits and this is always an extremely challenging requirement. Particularly for “brain-workers” like developers or other people from IT, it is very helpful to have some background information in combination with practical applications given by experts. First little steps of debugging the communication will soon be honored by success.

    “First do it right, than do it fast!”, is a developers wisdom. Concerning the communication this wisdom could be changed into: “First debug your communication, then optimize it!”

    Coming back to the question in the headline, it is not important to blame anybody for a bad communication. Everybody is responsible for it and instead of searching for the guilty one, you should start working on it. Be brave, start debugging you communication!

  • GitHub, Bug bounties and DevOps

    GitHub, Bug bounties and DevOps

    A year after starting up its bug bounty program GitHub is showing how complementary bug bounties can be to DevOps practices. Essentially crowdsourcing vulnerability hunting by systematically paying independent researchers prizes for flaws they find, bug bounties are in a sense an extension of the DevOps ethos. They offer a means of continuously delivering application security assurance.

    “The advantage of having a community that is just working 24/7 worldwide is it does tend to be a more continuous approach to vulnerability hunting as opposed to the traditional method of hiring out a consultancy or third party to spot check at some point in time,” explains Shawn Davenport, vice president of security for GitHub.

    And according to GitHub’s security leaders, their program has become one of the most effective and cost-conscious tools they have for extending internal appsec resources. With approximately four application security experts on staff, GitHub has a more than respectable team in place for a company of its size, with an employee total of just over 250. But with so many deploys a day, the team needs air support to ensure they’re thoroughly inspecting the code base.

    “We ship software 100 times a day, we update the site constantly and the security team isn’t looking at every one of those changes,” says Ben Toews, application security engineer at GitHub and the mastermind behind its bug bounty program. “The way we fit security into our software development life cycle is that when someone is getting ready to ship a feature, they’ll reach out to the security team and ask for a review if they think it’s necessary and we’ll get a review in that point in time, but then a lot of smaller changes don’t necessarily get our attention. And then there’s a lot of legacy code we’ve never looked at.”

    Toews says the bug bounty fills the holes a number of ways. He estimates that for any given feature on the site, there’s probably been several hundred researchers who have tested them, “so the odds of someone looking at new code pretty quickly after a change is pretty high.” In its first year, GitHub’s bug bounty program fielded nearly 2,000 reports from bug hunters, 869 of which warranted further review.

    As GitHub rounds the corner into its second year of offering bug bounties, it wants to increase those odds of good coverage with some tweaks to its program. In addition to doubling its max bounty prize to $10,000, it is also experimenting with giving bounty hunters preview access to new features before they ship.

    “We think by doing that we can get even better coverage on features before potentially putting the community at risk,” says Toews, who admits it does complicate rapid release schedules. “We’ve been trying recently to be more methodical about how we ship larger features. If we’re developing a new feature, you will have that feature hidden from the general population and you can develop that and ship it many times a day and iterate on it quickly, but then we do have a little bit of a process around making that feature published and exposing it to the rest of the world.”

    Of course, the efficacy of a bug bounty program depends on organizations finding a good way to marry bounty vulnerability remediation with the deployment process. It’s not good enough to find the vulnerability—you also need to be able to fix it quickly. Fortunately, the continuous delivery model makes this easier than in traditional organizations, says Toews, who reports that independent researchers are often “astounded” that his team can sometimes have problems fixed within hours of receiving a report, considering that they’re used to other organizations taking months to remediate problems.

    Toews says the secret to success is rapid triage of vulnerabilities. If it is simple, he and his team of application security engineers will fix it on the spot. If it is more complex, they’ll loop in the development team. This is where it is important to have buy-in from developers and senior leadership about bug bounties and application security in general, he says.

    “Our developers care a lot about security and they know that our leadership cares a lot about security, so if there’s a problem, everyone is going to jump in and try to fix it,” he says.