Tag: Fluentd

  • Switching From FluentD to Vector Log Aggregation Tool

    Switching From FluentD to Vector Log Aggregation Tool

    Log files are extremely important to the data analysis process as they contain essential information about usage patterns, activities and operations within an operating system, application, server or device. This data is relevant to a number of use cases across an organization from resource management, application troubleshooting, regulation compliance and SIEM and business analytics and marketing insights. To manage logs created by these use cases and make use of this wealth of data, log aggregation tools enable organizations to systematically collect and standardize log files. However, choosing the right tool can be quite challenging.

    This blog will detail and compare the popular open source fluentD and Vector tools for log aggregation.

    fluentD Configuration and Efficiency Calculation

    When using orchestration tools like Kubernetes to deploy containers or other API resources, there is a need for a log aggregator to store the pod or node logs in a cloud platform. For a particular requirement, fluentD was used as a log aggregator tool to push K8s pod logs to cloud storage buckets with a sample configuration as shown below:

    <match kubernetes.**>

    @type <cloud platform name>

    project <project name in cloud platform>

    keyfile <credential json to access the cloud storage> bucket <cloud storage bucket name> object_key_format <name for the file to be used>

    path <file prefix/path where the file have to be stored>

    <buffer tag,time>

    @type file

    path /var/log/fluent/gcs timekey 1m timekey_wait 30 timekey_use_utc true flush_thread_count 16 flush_at_shutdown true flush_mode interval flush_interval 1 chunk_limit_size 10MB retry_max_interval 30 retry_wait 60

    </buffer> <format>

    @type json </format>

    </match>

    Using this system, fluentD was pushing only 47.62% of total logs to cloud storage. Since there was a loss of more than 50%, changes were made to the configuration. In most of the changes, the efficiency was somewhere between 40% and 50%, with a maximum efficiency achieved at an average 67% for an entire day. Below are some of the changes made along with the percentage of logs that were pushed to cloud storage:

    <buffer tag,time> @type file

    path /var/log/fluent/gcs timekey 1m timekey_wait 30 timekey_use_utc true flush_thread_count 16 flush_at_shutdown true retry_max_interval 60 retry_wait 30

    </buffer> Efficiency:- 46.32%

    <buffer tag,time> @type file

    path /var/log/fluent/gcs timekey 1m timekey_wait 30 timekey_use_utc true flush_thread_count 16 flush_at_shutdown true

    </buffer> Efficiency:- 49.89%

    <buffer tag,time>

    @type file

    path /var/log/fluent/gcs timekey 10m timekey_wait 0 timekey_use_utc true flush_at_shutdown true

    </buffer> Efficiency:- 37%

    <buffer tag,time> @type file

    path /var/log/fluent/gcs timekey 30 timekey_wait 0 timekey_use_utc true flush_thread_count 15 flush_at_shutdown true

    </buffer> Efficiency:- 60.88%

    <buffer tag,time> @type file

    path /var/log/fluent/gcs timekey 1 timekey_wait 0 timekey_use_utc true flush_thread_count 16 flush_at_shutdown true flush_mode immediate

    </buffer> Efficiency:- 66.77%

    Vector Deployment, Configuration and Resultant Efficiency

    To improve this further, the open source Vector tool by Datadog was also considered. This tool was suitable for K8s setup with a similar configuration as fluentD and was installed in the nodes.

    A Helm command was used to clone the official repository in its VMs; the configuration was changed as described below and installed it as an agent. Vector comes in two working modes: Agent and aggregator. While agent is the plain mode that pushes logs/events from source to destination, aggregator is used to transform and ship data collected by other agents (in this case, Vector).

    The installation of this tool requires a Helm repository in the local machine to fetch the source code. Hence the below commands were run in a sequential pattern before installing Vector in a K8s cluster:

    helm repo add vector https://helm.vector.dev (Adding vector repo to helm list)

    helm repo update (Updating the helm repos)

    helm fetch –untar vector/vector . (command to clone the repository to local machine)

    Configuration:-

    data_dir: /vector-data-dir

    sources:

    <Custom source id>:

    type: kubernetes_logs (because we are using kubernetes as our source)

    exclude_paths_glob_patterns: <Array of directories which has be excluded when collecting the logs from the nodes> (Optional)

    sinks:

    <Custom sink id>:

    type: <Destination cloud storage>

    inputs: <Array of source id’s from where log has to be pushed>

    bucket: <Bucket name of cloud storage>

    key_prefix: <Path inside the bucket where the logs has to collected> (Optional) encoding:

    codec: <Encoding of the log file> (Optional) Command to install the Vector:-

    helm install vector . –namespace vector

    After deploying Vector in the development environment and testing it, the efficiency was ~100% with negligible loss. The switch was then made to Vector and deployed in the production environment. Vector can ship up 100,000 events or logs/sec, which is a very high throughput rate compared to other tools for log aggregation performance. Vector was able to achieve 99.98-100% efficiency even in the Kubernetes production cluster.

    To learn more about how DataOps can enable highly performant data pipelines with real-time logging and monitoring, watch this video. 

  • Why DevOps are Keen on Open Source Logging

    Why DevOps are Keen on Open Source Logging

    With the advent of DevOps, we are now moving toward more collaborative workflow models, where the integration of development and operations teams make it easier to quickly move development projects into production, and with fewer roadblocks. However, within DevOps, many of our activities are not as clear-cut as they were in previous siloed methods. While previous methods made it easier to keep track of activities and processes, these systems were not particularly efficient.

    In the DevOps model, moving software from development to production is much faster. But, we need to be able to identify actions that occur in our increasingly complicated applications. We now have many chefs working at the same time, and visibility suffers.

    Logging helps us better understand applications throughout each stage in the process so that we can identify problems or larger issues that might arise already during the development process. This is particularly true within CI/CD environments where items are regularly pushed into production. We need methods for identifying not only if there are critical errors, but also the ability to attribute the errors to specific versions throughout the cycle.

    To deal with the increasing complexity and decreasing visibility, DevOps identifies complex issues using logs to keep track of services, APIs, containers, infrastructure issues, network activity and also security related logs.

    Log files provide a tremendous amount of useful information about what goes on beyond the CI/CD pipeline and also identify issues that directly impact customers. Good log analysis tools make it possible for DevOps teams to get much closer to the end user. And open source is the preferred starting point for many organizations.

    Why Open Source?

    Lack of Lock-in 

    Many DevOps teams make use of a wide-range of tools and find that they can work much better without vendor lock-in, which can be caused when using proprietary stacks exclusively. This makes it far easier to make any needed changes as the need arises, and not be reliant on individual vendors. This also helps avoid many risks, such as vendors disappearing, ending support, changing pricing models, etc.

    Widely Used

    Within CI/CD, there are many commonly used open source tools, which makes it easier for new members to be on-boarded. There is already a widely accepted suite of tools that almost all DevOps teams use. It’s hard to imagine working without Jenkins, Docker, Kubernetes, Git, etc.

    Regular Skill Development

    Open source work ensures that DevOps teams are actively learning and directly interacting with software and systems. It becomes a way of ensuring that skills and abilities of teams are continually staying updated. This can translate into better troubleshooting and problem solving skills. When working with closed systems, there’s considerably less opportunity for knowledge advancement, as closed systems tend to encourage rote work with existing software rather than skill development.

    Quality

    Fears about subpar open source software are largely unfounded, particularly if the projects are widely used and have a large community. There are far more developers working on a project in the open source community than in proprietary systems. Closed systems can lag behind, and there may actually be bad code but that is hidden due to its closed-source nature. The ability to make modifications to create functionality beyond that which already exists in closed source systems, makes open source extremely appealing for many DevOps personnel.

    Some Drawbacks

    It’s important to note there are a few drawbacks when attempting to work entirely within open source environments. Many resources used are not open by nature–for example leading clouds are decidedly not open source in nature. Individual users (outside these companies) have much less control over how they actually operate. However, if we look at the toolkits that we described above, we can think of open source tools as needed resources for working in and around these environments.

    Another drawback to consider is the actual costs of managing open source software versus a fully-managed service such as Coralogix, for example. Many times, teams will start with the self-managed approach and move to a managed service when faced with scaling and complexity issues as the organization grows. 

    Types of Features to Look for in Open Source Logging Models

    When setting up to work with open source logging tools, it’s important to understand there is a range of different types of models available, which can work better or worse depending upon needs and capacities.

    • OpenAPI: provides a useful interface into many different systems. It is language agnostic and can generate clear documentation of methods, parameters and models. In most cases it will work within RESTful interfaces, but will also work with other protocols such as SOAP.
    • Open Standards: Even if a tool is not open source, typically these should follow a set of standards to make sure that different pieces can talk to each other. OpenAPI is one example, but not the only one.
    • Federated Model: This is a broader type of open source model which is like many default open tools, but allows a much wider set of input and may allow individual users to provide input and maintain some control over local development areas. This means that aggregation, processing and control remain local. However there’s still a central organization which will be able to collect summarized or complete code. The advantage of a federated model is it increases flexibility for individual teams, while still contributing largely to the project as a whole. 

    Elasticsearch, Logstash, Kibana (ELK)

    One of the most popular open source logging stacks is known as ELK. It is particularly powerful and effective for the purposes of those wishing to create a clear aggregation and understanding of their log files, with the ability to find both problems and solutions quickly. 

    ELK is a combination of three tools.

    • Elasticsearch: a popular NoSQL Lucene search engine.
    • Logstash: a pipeline system which will take data from logs, ingest it and transform it into a data store (such as Elasticsearch), so that it can be searched and analyzed.
    • Kibana: a tool which makes it possible to create clear visualizations for Elasticsearch to appear in a human-readable format. 

    What ELK provides is a single place for all of your data, where it can be stored, searched and analyzed in a method. It is useful for DevOps teams wishing to make use of the vast amounts of data stored in log files. 

    Upfront expenses for open source tools like ELK are minimal; the software itself is free. However, this does come with a caveat. Professional management of ELK tools is required, otherwise costs can grow quickly out of control. 

    Generally, for low volumes of data, one can manage ELK for as little as a few hundred dollars a month on a platform such as AWS. However, without careful management, the cost can grow exponentially for large amounts of data.   

    Work required for ELK involves:

    • Setup, which may take some time but that’s typically not a problem for many ops teams—and is a project many will enjoy.
    • Maintenance of elastic search clusters can be quite a bit of work—several hours per week simply solving issues or resolving downtime. This can vary based on the size and growth of clusters. The more activity you have occurring, the more time you will need to spend maintaining your logs.

    FluentD

    FluentD is an open source alternative to Logstash which can collect, parse and transform data for further analysis. It has features which may be particularly helpful if you are running your logs in a Kubernetes environment.

    One of the drawbacks of Logstash is that it is written in JRuby, which means it requires a Java runtime for each implementation. If you are running many different microservices in Kubernetes, this can result in a large amount of your memory being taken up.

    Because FluentD is developed in CRuby, it doesn’t require as much memory to use with each new pod creation. This can help solve some memory allocation issues which could cause your logs to slow down your applications.

    FluentD also has a wide range of plugins, which can cover pretty much any use case that you come up with. Another advantage is that its tag-based routing makes it slightly easier for developers to work with than Logstash. 

    What to Log?

    With continuous integration, there are many things that merit monitoring. These can include everything from checking the frequency of merge conflicts in Git, or whether pulls are preceded by pushes prior to any other pulls.

    Are you receiving a high number of compiler warnings? You can identify code that has a strong likelihood of breaking builds, change your unit testing procedures and implement better automated checking in Jenkins.

    Much of this data is worth visualizing in Kibana for analysis. Deciding what specific activities to log can be a bit involved. Areas that are a good idea to log include:

    • Diagnostics: For example, understanding the causes of repeated errors.
    • Auditing: Track how well errors in code have been fixed, and whether they succeed or fail a second time.
    • Profiling: For example, get an idea of how long it takes for certain parts of code to execute, for the purpose of identifying whether this has any impact on customer experience, and identify areas for improvement.
    • Statistics: Code performance, such as execution time or memory consumption.

    Conclusion

    Gaining complete visibility or your infrastructure will go a long way to make sure that your operation runs smoothly and that your CI/CD pipelines function at optimum efficiency. Open source tools remain an attractive option for DevOps because of its advantages and knowing how to get around drawbacks.

    Finding the right tools is, of course, only the first step, but if you’ve picked good architectures and have implemented them properly, you are already moving in the right direction to a more streamlined operation.

  • Identifying and Abstracting Business Intelligence from Kubernetes Workloads

    Identifying and Abstracting Business Intelligence from Kubernetes Workloads

    In the age of big data, businesses are inundated with data points. While they know there’s an abundance of valuable resources available to them, making sense of this data to derive actionable insights is often still a challenge.

    CNCF industry survey data shows workload complexity and monitoring remain top challenges for enterprises in terms of using and deploying Kubernetes. Many understand there are valuable resources within these environments, but struggle to best identify and extract meaning from machine data. While most of this data is readily available, it just takes the right tools to gather and view intelligent insights.

    The Benefits of Open-Source Data Collection

    Perusing the CNCF website, there are a wealth of tools built around Kubernetes to enable not only monitoring but networking, storage and security. There is even a Kubernetes-specific package manager, Helm, to make the deployment and management of these resources easy and consistent.

    There are a couple of key benefits to taking advantage of these open-source collectors:

    1. They stay up to date. Each of these tools benefits from deep community support. As new versions of Kubernetes are released, the extensive use of each of these tools ensures they are quickly updated.
    2. They integrate with everything. Prometheus, for example, has an impressive list of integrations and exporters.  Regardless of your unique stack, it is likely that there is support for what you might want to export data from. The importance of these integrations cannot be overstated, as they enable the flexibility needed to grow and evolve a Kubernetes deployment over time.

    Ensuring Complete and Organized Metrics Collection

    Prometheus–endorsed by the CNCF–is the de facto tool of choice for metrics monitoring with extensive support for anything you might use to collect metrics data.

    Prometheus works by pulling data from all of the components and jobs running in Kubernetes since every component of Kubernetes exposes its metrics in a Prometheus format. The processes running behind those components then serve up the metrics on an HTTP URL. For example, the Kubernetes API Server serves its metrics on https://$API_HOST:443/metrics.

    Prometheus is particularly good at auto-discovering the jobs and services currently running in a Kubernetes cluster. As pods are added, removed or restarted, the Kubernetes Service construct keeps track of what pods exist for a given service. This auto-discovery capability is one of the primary reasons for Prometheus’ popularity, ensuring that all new and existing components are monitored.

    The Importance of Cluster Level Logging and Event Collection

    Kubernetes does not define a single standard approach to log collection, but the most common method is called cluster level logging. Cluster level logging deploys a node level logging agent to each node which then funnels data to a separate backend for storage and analysis of logs. The primary benefit of this solution is that if a pod dies, the logs detailing what happened are retained. Implementing node level logging, without funneling data to a logging backend, will not retain log data if pods die or are evicted, while cluster level logging ensures that data is captured and retailed. A common tool for implementing cluster level logging is Fluentd—or Fluentbit, a lightweight version of Fluentd—which acts as the node level logging agent funneling data to a logging backend.

    Events provide insight into decisions being made by the cluster and unexpected events that occur in Kubernetes. Events are stored by the API server on the master node and collected using the same method as log collection—via a node level logging agent like Fluentd.

    Establishing a Continuous Intelligence Dashboard

    Finally, collectors for logs, metrics, events and security can be easily deployed using Helm—an open source Kubernetes package manager. Helm can significantly simplify the setup process, reducing hundreds of lines of configuration to one. These collection plugins can be used on any Kubernetes cluster, whether one from a managed service such Amazon Elastic Kubernetes Service (EKS) or a cluster you are running entirely on your own.

    According to Sumo Logic’s recent Continuous Intelligence Report, enterprise adoption and deployments of multi-cloud grew 50% year-over-year, and 80% of users across multi-cloud environments are now utilizing Kubernetes architectures. As Kubernetes continues to become mainstream, it’s essential that businesses understand how these workloads operate and how to best extract valuable insights to inform business decisions.

    — Katie Lane