Tag: big data

  • Data APIs: Realizing the Future of Data Warehousing

    Data APIs: Realizing the Future of Data Warehousing

    Data is the currency of business in today’s digital economy, with organizations collecting mountains of data from their customers, products and services. Enterprises are increasingly turning to data warehouses to store this valuable enterprise data and to make it useful and actionable. Data APIs–both GraphQL and REST–are emerging as a tool to power frictionless and high-quality integrations between applications/services and data in enterprise data warehouses. But what exactly is driving this shift toward using APIs on top of data warehouses? In this article, I will talk about the drivers behind the rise of data APIs on data warehouses, challenges with current approaches to building, operating and scaling data APIs and how organizations can address those challenges.

    The Rise of Data APIs

    Data APIs enable organizations to create powerful integrations between applications without having to go through lengthy development cycles or write complex code. As a result, businesses can deliver improved user experiences faster than ever before by quickly connecting disparate applications and services together using intuitive REST or GraphQL endpoints. Because these endpoints are standardized across multiple systems (e.g., REST or GraphQL), developers don’t have to learn new languages or coding techniques every time they want to connect two different applications together. This makes it much easier for developers to work with multiple systems simultaneously while still ensuring that all applications remain connected via one interface–the API endpoint itself.

    Challenges With Current Approaches

    Despite all these advantages provided by Data APIs, there still remain a few challenges that organizations need to consider when building out their own solutions. For example, many existing approaches require manual setup processes, which can be extremely time-consuming and tedious for developers working with multiple systems at once. In addition, many existing solutions only allow for limited query functionality, which means that more complex queries may not be possible without additional coding or manual setup processes being performed. Finally, some existing solutions require an extensive amount of organizational resources in order to build out a full-fledged solution—time that could be better spent on more important tasks like product development or customer service initiatives.

    How Organizations Can Address These Challenges

    Fortunately, there are now tools available that make it easy for organizations to quickly build out powerful data API solutions with minimal effort required from their developers or IT staff members. These tools provide pre-built templates that can be used as a starting point for any integration project while also allowing users to customize their own solution based on specific needs and requirements without having to write any code themselves. Additionally, these tools allow users access to advanced query functions such as joins across multiple tables so they can get the most value out of their databases without compromising either speed or security either. Finally, these tools also include powerful monitoring capabilities so users can ensure that their data APIs remain running smoothly at all times, no matter how much traffic they receive.

    Benefits of Using an API for Integrating Cloud Data Warehouses

    When it comes to integrating cloud data warehouses into your application or service, there are several advantages that come with using an API as opposed to traditional methods such as manual coding or ETL tools:

    • Increased Flexibility: By leveraging an API, you are able to access data from multiple sources quickly and easily without having to manually code every integration yourself. This also allows your application or service more freedom when it comes to accessing additional data sets in the future; all you need is an existing API connection in place.
    • Improved Security: Since APIs are designed specifically for secure communication between different systems, they offer enhanced levels of protection against malicious attacks such as DDoS attacks or unauthorized access attempts. You can also add additional layers of security by implementing authentication protocols like OAuth or JWT tokens on top of your existing API connection.
    • Enhanced Performance: APIs can be designed with specific performance optimizations in mind; this means that you can create custom solutions tailored towards the exact needs of your application or service which will ultimately result in improved end-user experiences and higher engagement rates from customers who interact with your product or service regularly.

    Data warehouses have become an essential part of modern digital businesses due to the sheer amount of valuable enterprise information they store. However, traditional approaches can often lack the flexibility or scalability needed in order for businesses to stay competitive in today’s ever-changing market environment. This is why smart organizations are using data APIs on top of them for quick integration among disparate applications/services along with advanced query capabilities.

    By leveraging tools designed specifically for this purpose, companies can easily overcome current challenges associated with building out their own solutions while still getting maximum bang for their buck when it comes down to successful implementation and operation over the long term.

  • New-Age Tech and Culture Driving Data Democratization

    New-Age Tech and Culture Driving Data Democratization

    Data democratization has become critical for organizations to leverage data’s true value and fully realize its benefits. Making data available across the organization helps companies to serve customers better and make data-driven decisions that align with their goals and objectives. It essentially means enabling access to a company’s data resources to employees, bound by reasonable limitations on legal confidentiality and security.

    As a means to remove data silos in organizations, the indispensability of data democratization is clear. It allows more widespread access to and use of data within an organization. This can lead to more informed decision-making, increased collaboration and innovation and more efficient use of resources.

    Breaking down data silos can help reduce the risk of data breaches and improve data security. While the jury may be out on whether data democratization is beneficial in preventing security breaches, silos between networks or security systems make it difficult to identify large-scale attacks as they tend to represent disjointed security protocols. There is less likelihood of robust security protocols between non-communicating teams.

    New-Age Tech as an Enabler of Data Democracy

    New-age technology such as cloud computing, blockchain, and advanced analytics are aiding data democratization. For example, 5G enables more significant volumes of data to be moved faster and more securely in the form of big data. Blockchain provides higher-quality data and improves the operation of cyber and physical systems. In addition, IoT allows for exchanging data between devices, where processes can be established between devices without intermediaries.

    A hybrid cloud adoption survey conducted in March 2022 of IT professionals primarily based in North America and Europe revealed 93% of the IT industry would adopt a hybrid of cloud and on-premises solutions or migrate fully to the cloud within five years, using it to store and manage data. Gartner estimated that 85% of enterprises would embrace a cloud-first principle by 2025 and that it will be essential to execute digital strategies in business.

    This suggests that new-age technologies play a significant role in breaking down data silos and enabling organization-wide access to data. By leveraging these technologies, organizations can easily store, process and share data, leading to better decision-making, increased collaboration and improved business outcomes.

    Data creation, development, and utilization directly relate to Industry 4.0 applications. In practice, data democratization is seamlessly transferable to Industry 4.0 applications. The question arises: Do industrial companies have the capabilities today to apply the concept of data democratization? And to what extent?

    As new-age tech drives the emergence of new business models, companies that will lead a cultural shift as well as a technology evolution will be successful. This is possible via a business-centric data strategy, along with the democratization of analytical tools and platforms to enable stakeholders to help build better insights into their business processes.

    Culture as a Prerequisite for Data Democracy

    A Google Cloud and Harvard Business Review survey of industry leaders revealed 97% believe organization-wide access to data and analytics is critical to the success of their business. However, only 60% of respondents believed that their organizations effectively distribute that access today.

    What are the prerequisites for data democracy? First, companies should inculcate a culture change to create a pull from inside the organization rather than first developing the technology. Second, decision-making in companies is based on experience and intuition than on data; they should compulsorily begin to use data as a baseline for business decisions. Third, leaders should encourage a data-driven culture to grow from within.

    Formalizing and sharing knowledge can remove roadblocks to the democratization of data. Therefore, one form of action is to set specific incentives to foster a mindset of knowledge sharing. For instance, existing practices to incentivize participation in the suggestion or lessons-learned programs can be used as a reference.

    Many companies have procedures to involve employees in the continuous improvement process. But while employees can use the existing IT systems to perform their day-to-day work, they need more capabilities to cover their individual information needs. Thus, as a way forward, capability building within is critical. Another strategy is decentralizing decision-making; if employees can gain insights independently, they should be able to act on these insights independently.

    Organizational culture is vital to data democratization, as it determines the attitudes, behaviors and values that guide how data is shared and used. On the other hand, a culture that values secrecy or hierarchical control over data can lead to resistance to data democratization efforts and perpetuate data silos. A culture that lacks trust and transparency can also lead to reluctance to share data, which can impede collaboration and hinder data-driven decision-making.

    Data Democratization – The Way Ahead

    Organizations need to assess and work on their tech stack and internal culture to promote a data-friendly culture supporting data democratization. To foster a data-friendly culture, organizations should encourage open communication and collaboration among employees and provide training and resources for employees to access, understand, and analyze data.

    Open communication, collaboration, innovation and trust are essential attributes for breaking data silos. A corporate culture that values data-driven decision-making will lead to more widespread access to and use of data and support the implementation of new-age technologies. Finally, there is a need to establish clear data governance policies and procedures that support data democratization and build a continuous improvement process by encouraging experimentation and learning from failures.

    Image Source: Joshua Sortino, Unsplash

  • Study: Real-Time Data Has A Transformative Effect

    Study: Real-Time Data Has A Transformative Effect

    We all know that data is the heart of the modern digital business. And real-time data is arguably the next evolution—having up-to-date data can empower energy savings and optimizations to reduce cost. Real-time data processing also aids fraud detection and improves health care responsiveness.

    Especially in DevOps, access to real-time data is necessary to reduce incident response time and inform ongoing SRE objectives. Leveraging performance data can also empower future decision-making and open new revenue opportunities. In essence, real-time data can have a transformational effect on the operations of a digital business.

    DataStax and ClearPath Technologies recently released the 2022 State of the Data Race Report. The study found that the majority of data leaders are using real-time data to create additional value. The report also suggested that working with cutting-edge real-time data technologies increases developer experience, which is good to maintain amid ongoing talent shortages.

    Below, I’ll review the key findings from the study to infer how organizations can effectively use these technologies and how they can mitigate roadblocks that arise throughout the process.

    The State of the Data Race

    Simply put, customers expect real-time experiences. And it’s not only the end consumers—developers are also increasingly accustomed to building in this modality. This is evidenced by the recent rise of asynchronous, event-driven architectures. The survey depicted how fundamental real-time data is to modern software, finding that 78% of respondents agreed that it is a “must-have,” not just “nice-to-have.”

    The investment appears to be paying off. A full 71% of all respondents said they can directly tie revenue growth to real-time data. Not only does it seem to have a transformative impact on revenue growth, but it enhances productivity, too—66% of real-time data-focused companies agreed that developer productivity has improved. Naturally, the report found that developers are the ones most likely to be working with real-time data.

    Although modern databases and real-time data processing tools are more readily available these days (and open sourced), the maturity of data handling is not homogenous. Based on the classifications of this report, “data leaders” made up only 10% of respondents.

    Roadblocks to Value

    The road to data maturity has a few obstacles that could limit adoption. First is data complexity—33% said this is one of the biggest barriers to leveraging real-time data. This was closely followed by controlling data costs (32%), data accessibility (30%) and available skillsets (28%).

    Most notably, a skills gap continues to be a commonly cited hindrance throughout industry reports and across technology niches. Interestingly, a lack of in-house talent rises slightly higher in organizations with a high real-time data maturity. Those surveyed in this category said that their top barrier is finding the appropriate skills in their business unit.

    Cybersecurity is also a top concern—51% say security challenges are a big concern when it comes to the use of real-time data. As the report surmised, the business-critical nature of real-time data raises its worth as a commodity and area for attack. Furthermore, amid new data sharing regulations, IT departments are now under more pressure than ever to uphold data security and avoid the leakage of sensitive personal information.

    Unlocking Real-Time Data

    This area has a lot of potential; 86% of developers at organizations with a strategic focus on real-time data said that the “technology is more exciting than ever,” the report found. Real-time data is beginning to be applied at the edge, where AI/ML can be trained to give a more responsive advantage to modern applications. Yet, organizations must prioritize its adoption and mitigate the roadblocks mentioned above to unlock the benefits of real-time data.

    Prioritizing real-time data will take more than tools—it will take vision. According to the study, those who created more value with real-time data initiatives were found to:

    • Have clear product owners
    • Have business unit accountability of data
    • And have department representatives working together in a cross-functional team.

    As you can see, these foundational goalposts have less to do with technology and more with lucid leadership and collaboration. That being said, an investment in the appropriate tooling is undoubtedly required to achieve real-time data.

    For example, this tooling might encompass NoSQL databases optimized for real-time or Apache Cassandra for high throughput. The transition to real-time will likely require streaming and messaging technologies as well. Less obvious, however, is an API layer to sit between the applications and the database. This element could greatly help open data accessibility in a standardized fashion.

    The State of the Data Race: Final Thoughts

    Working with real-time data can be a differentiator and have a transformative impact. As a result, most developers and leaders see real-time data as an exciting area. Yet, organizations must reduce data siloes to open up data access to realize its benefits. It will also take collaboration between data scientists and business divisions to construct meaningful outcomes.

    The State of the Data Race report surveyed 500 leaders across various industries. Above, we’ve summarized the key points from the survey. Readers are encouraged to pick up a copy for more nuanced data points and further commentary.

  • Where DataOps and Opportunities Converge

    Where DataOps and Opportunities Converge

    In today’s data age, getting data analytics right is more essential than ever. A robust data analytics implementation enables businesses to hit key performance metrics, build data and AI-driven customer experiences (think ‘personalize my feed’) and capture operational issues before they spiral out of control. The list of competitive advantages goes on, but the bottom line is that many organizations successfully compete based on how effectively their data-driven insights inform their decision-making.

    Unfortunately, implementing an effective data analytics platform is challenging due to orchestration (DAG alert!), modeling (more DAGs!), cost control (Who left this instance running all weekend?!) and fast-moving data landscapes (data mesh, data fabric, data lakehouses …). Enter DataOps. Recognizing modern data challenges, organizations are adopting DataOps to help them handle enterprise-level datasets, improve data quality, build more trust in their data and exercise greater control over their data storage processes.

    What is DataOps?

    DataOps is an integrated and agile process-oriented methodology that helps businesses develop and deliver effective analytics deployments. It aims to improve the management of data throughout the organization. 

    While there are multiple definitions of DataOps, below are common attributes that encompass the concept while going beyond data engineering. Here’s how we define it:

    We broadly define DataOps as a culmination of processes (e.g., data ingestion), practices (e.g., automation of data processes), frameworks (e.g., enabling technologies like AI) and technologies (e.g., a data pipeline tool) that help organizations to plan, build and manage distributed and complex data architectures. DataOps includes management, communication, integration and development of data analytics solutions, such as dashboards, reports, machine learning models and self-service analytics.

    Why DataOps?

    DataOps is attractive because it eliminates the silos between data, software development and DevOps teams. The very promise of DataOps encourages line-of-business stakeholders to coordinate with data analysts, data scientists and data engineers. Via traditional agile and DevOps methodologies, DataOps ensures that data management aligns with business goals. Consider an organization endeavoring to increase the conversion rate of their sales leads. In this example, DataOps can make a difference by creating an infrastructure that provides real-time insights to the marketing team, which can help the team to convert more leads. Additionally, an Agile methodology can be employed for data governance, where you can use iterative development to develop a data warehouse. Lastly, it can help data science teams use continuous integration and continuous delivery (CI/CD) to build environments for the analysis and deployment of models.

    DataOps can Handle High Data Volume and Flexibility

    The amount of data created today is mind-boggling and will only increase. It is reported that 79 zettabytes of data were generated in 2021 and that number is estimated to reach 180 zettabytes by 2025. In addition to the increasing volume of data, organizations today need to be able to process it in a wide range of formats (e.g., graphs, tables, images) and with varying frequencies. For example, some reports might be required daily, while others are needed weekly, monthly or on demand. DataOps can handle these different types of data and tackle varying big data challenges. Add in the internet of things (IoT), such as wearable health monitors, connected appliances and smart home security systems, and that introduces another variable for organizations that also have to tackle the complexities of heterogeneous data as well. 

    OK, so, how can we make this a reality? First, to manage the incoming data from different sources, DataOps can use data analytics pipelines to consolidate data into a data warehouse or any other storage medium and perform complex data transformations to provide analytics via graphs and charts.

    Second, DataOps can use statistical process control (SPC)—a lean manufacturing method—to improve data quality. This includes testing data coming from data pipelines, verifying its status as valid and complete, and meeting the defined statistical limits. This enforces the continuous testing of data from sources to users by running tests to monitor inputs and outputs and ensure business logic remains consistent. In case something goes wrong, SPC notifies data teams with automated alerts. This saves them time as they don’t have to manually check data throughout the data life cycle.

    DataOps can Automate Repetitive and Menial Tasks

    Around 18% of a data engineer’s time is spent on troubleshooting. DataOps enables automation to help data professionals save time and focus on more valuable high-priority tasks.

    Consider one of the most common tasks in the data management life cycle: Data cleaning. Some data professionals have to manually modify and remove data that is incomplete, duplicate, incorrect or flawed in any number of ways. This process is repetitive and doesn’t require any critical thinking. You can automate it by either setting customized scripts or installing a built-in data cleaning software tool.

    Additional processes that can be automated via DataOps include:

    • Simplifying data maintenance tasks like tuning a data warehouse
    • Streamlining data preparation tasks with a tool like KNIME
    • Improving data validation to identify flags and typos, such as types and range

    Building Your Own DataOps Architecture 

    To develop your own DataOps architecture, you need a reliable set of tools that can help you improve your data flows, especially when it comes to crucial aspects of DataOps, like data ingestion, data pipelines, data integration and the use of AI in analytics. There are a number of companies that provide a DataOps platform for real-time data integration and streaming that ensures the continuous flow of data with intelligent data pipelines that span public and private clouds. Looking to increase the likelihood of success of data and analytics initiatives? Take a closer look at DataOps and harness the power of your data.

  • Time-Series Database Basics

    Time-Series Database Basics

    There’s been a lot of buzz about time-series data for the past few years. No matter what you’re monitoring—financial data, server performance, social media activity or something else entirely—time-series data can give you valuable insights into trends and patterns.

    But what exactly is a time-series database, and why do you need one? This article will look closely at time-series databases and how they can help you make sense of your data.

    What is a Time-Series Database?

    A time-series database (TSDB) is a database management system optimized for storing and querying data that changes over time. Time-series data is often used in monitoring applications where it’s important to quickly retrieve information about a system’s current state and trends and patterns over time.

    In most definitions, time-series data, often referred to as time-stamped data, is a sequence of values indexed in time order. Time stamping refers to data collected at various times, where each value is time stamped. These data points are usually gathered from the same source and are used to measure progress over time.

    Important Notes:

    You can use relational or NoSQL databases to crunch time-series data, but purpose-built time-series databases are tailored to exploit the unique features of time-series data. This implies that time-series databases ingest at a faster rate, query more quickly and compress data more efficiently. Furthermore, time-series databases include special analytical capabilities and management features that are not found in most relational or NoSQL databases.

    Some Advantages

    This list is not exhaustive, but here are some of the advantages that you might get from using a time-series database:

    1.) Time-series databases can handle high-velocity data very well.

    2.) Time-series databases are purpose-built for storing time-series data, making them more efficient in storage and querying.

    3.) Time-series databases often have built-in analytics and management features designed specifically for time-series data.

    4.) Many time-series databases are open source, which means they’re free to use.

    Some Disadvantages

    There are also some disadvantages that you should be aware of before using a time-series database:

    1.) Time-series databases are often more complex to set up and manage than relational or NoSQL databases.

    2.) There are relatively few time-series databases to choose from, so you might not have as much flexibility in choosing a platform that meets your specific needs.

    3.) Time-series data can be very large, so you’ll need enough storage capacity to accommodate your data.

    What Are the Characteristics of Time-Series Data?

    There are several characteristics of time-series data that make it unique and require special handling. From a database standpoint, the most essential ones are as follows:

    1. Timestamp: Every data point has its timestamp. The timestamp is crucial for calculating or analyzing the information.
    2. Structure: Metrics from devices or monitoring are almost always structured, unlike those produced by Internet applications. They have predetermined data types and lengths, and the structure will not alter until the device’s firmware is updated.
    3. Stream-like: Data sources, such as audio or video programs generate data at a set or constant rate. These data streams are completely unrelated to one another.
    4. Stable flow: Time-series data traffic is constant over time, and it may be calculated and predicted if enough data sources are used within a certain sampling period.
    5. Immutability: The source of this data is a time-series data store. Each data point is created only once and never updated or corrected. Time-series data, like log data, is typically append-only.

    How Is Time-Series Data Used?

    Many of you may be wondering why time-series data is so important. The answer is that it allows us to track changes over time, which can be incredibly valuable for monitoring and troubleshooting purposes.

    For example, let’s say you’re a system administrator tasked with keeping an eye on server performance. By monitoring various metrics—such as CPU usage, memory usage and network traffic—you can quickly identify when a server is starting to experience issues.

    Time-series data can also be used for predictive purposes. For example, if you notice that CPU usage spikes at certain times of day, you can use that information to plan for future capacity needs.

    Of course, time-series data isn’t just limited to server performance. It can be used for any application where it’s important to track changes over time. This includes everything from weather and financial data to social media and website analytics.

    Let’s Get Specific

    Here’s what you should know about time-series data: It’s typically used to seek insights into operations, create alerts based on real-time analysis and forecast future trends.

    The following characteristics are found in time-series data applications:

    1. High write-read ratio: Twitter and LinkedIn are internet applications with single articles that millions of people read, but raw time-series data is primarily scanned and evaluated by apps and algorithms.
    2. Retention policy: In general, time-series data is not kept permanently. Organizations have a retention policy that specifies when and how their data is destroyed.
    3. Real-time analytics and computing: Time-series data must be calculated in real-time to identify out-of-the-ordinary activity and sound alarms based on the acquired information or aggregate findings.
    4. Query scope: Time-series data is generally requested over a period or a set of data sources, and filters are used to prevent all historical information from being requested. Furthermore, all or a portion of the data sources with a filter condition are always aggregated.
    5. Trends: Single data points are typically not significant in time-series data. The emphasis is on how data evolves, such as fluctuations in the previous hour or day.

    Popular time-series solutions like TDengine or InfluxDB, for example, create more efficient processing of time-series data and greater performance than general databases by taking advantage of these qualities.

    Do Time-Series Databases Need Specialized Databases?

    We live in a modern world where data is constantly being generated at an unprecedented rate. To keep up with the demand, businesses need to be able to store and process large amounts of data quickly and efficiently. This is where specialized time-series databases come in.

    Everything is online, including meters, automobiles, lifts, assembly lines and even bicycles. And these items are sending out a never-ending stream of numbers and events that IoT and the cloud have unleashed an explosion in time-series data generation.

    With this in mind, it’s not surprising that several specialized time-series databases have been created to deal with this type of data influx.

    Time-series data sets are enormous and pose a significant problem for general database management systems, such as relational and NoSQL databases. Non-specialized databases have trouble with the following elements of time-series data:

    1. Data ingestion rate: In many time-series data scenarios, millions of data points are generated every second and must be ingested in real-time. Relational databases are not built to handle this quantity of data, and while NoSQL databases may be scaled to do so, the amount of resources required soon becomes prohibitive.
    2. Query latency: It is often the case that a time-series system must scan a massive number of data points to get an aggregate result, which might lead to sluggish performance. For example, it would take days for a basic database to compute the average response time of all Amazon.com clicks, by which time the overall conclusion may be incorrect.
    3. Storage cost: Data generated by internet-connected devices and apps is nonstop 24/7, with a single day producing as much as a terabyte of data. Because relational and NoSQL databases cannot compress this data effectively, storage costs can quickly skyrocket.

    These problems are generally linked to efficiency in processing big data sets, although there are some places where general databases frequently do not meet even the fundamental demands of time-series applications:

    1. Data life cycle management: Time-series data is generally removed in bulk, not one data point at a time, as it ages out.
    2. Roll-up: In most cases, time-series data is collected and rolled up over a set period before being stored in the new table. Raw data and rolled-up data can have distinct life cycles and retention policies.
    3. Special analytic functions: Time-series applications need more specialized features than general databases, such as time-weighted average, moving average, cumulative sum, rate of change, elapsed time for a specific state and the delta between two consecutive data points.
    4. Interpolation: The database management system must be able to interpolate data based on the adjacent data points and rules to regularize data sets when applications or algorithms require.
    5. Continuous query: Therefore, if you hear a user or customer say they want to know when their data is updated, it’s reasonable to assume they want this functionality.
    6. Session and state windows: Aggregations and analytical procedures may be performed on a session or state window, not just time—for example, consider one that only calculates average power consumption when a machine is switched on.

    With general databases, developers must write custom code to implement features specific to their data set. Different data workloads require different database solutions; one size does not fit all. For time-series data, no matter the size of your data set, a purpose-built time-series database is the best tool for the job.

    Why Are Time-Series Databases Becoming Popular?

    Two years ago, time-series databases were the most popular database type in the business, owing to its expanding number of use cases. It is especially useful if you’re executing sophisticated transactions like advertising, e-commerce, supply chain management and so on.

    However, because of the rapid expansion of the IoT, they are becoming even more popular. As more devices become internet-connected and send data—time-series, of course—to the cloud, many industries are interested in purpose-built time-series databases.

    Time-series data is proving to be an important tool, not just for decision-making and optimization but also for industrial applications. Finally, IT infrastructure has been growing rapidly, with everything from servers, containers, network devices, apps, and microservices being monitored to create massive quantities of time-series data.

    Older time-series databases, on the other hand, are often closed systems employing antiquated structures that do not scale to handle the increasing amount of data. A million time-series data points used to seem like a lot, but now millions and even billions of data points are commonplace.

    Finally, integrating old (previous) time-series solutions with popular data analysis tools like artificial intelligence and machine learning platforms is difficult, if not impossible. These legacy systems cannot be transferred to the cloud without a lot of work and their licensing terms are no longer adequate for current apps.

    Final Thoughts

    The expanding market and the constraints of previous time-series databases are making room for a new generation. These new time-series databases are built from the ground up to take advantage of the cloud, big data and AI/ML technologies.

    If your business or application deals with time-series data, you need a purpose-built time-series database. The advantages of using such a database are too great to ignore, and the disadvantages of using an older, less capable database are becoming more and more apparent.

  • Deciphering the Observability Market

    Deciphering the Observability Market

    Observability has become one of the most overused buzzwords in IT and cybersecurity. Today, the term is used by vendors to refer to everything from application performance to network monitoring, cybersecurity and data and analytics.

    While the term’s ubiquity has created confusion for everyone from end users to journalists, startups in the space have also attracted over $2 billion dollars in venture capital investment over the past two years. This momentum prompted TechCrunch to ask if this area, specifically data observability, was effectively recession-proof.

    This observation triggers two questions. First, what is observability? And second, what are the different kinds or variants?

    What is Observability?

    Observability, as a broad practice or capability, originated in the 1960s as part of industrial control theory. The idea is that by watching the output of a system, you can figure out what’s happening inside the black box. This sounds a lot like monitoring, a popular practice for IT and security teams. However, observability takes a different approach than monitoring’s alert-based methods; it allows you to ask questions about a system that dig deeper than the pre-defined thresholds from your monitoring system. 

    Observability requires collecting massive amounts of data from systems, networks and applications to feed its discovery process. As systems and applications become more complex, figuring out what went wrong and why becomes much more challenging. At a high level, these are the kinds of challenges many startups in this area are addressing. Each takes a different approach and solves a different part of the observability problem.

    Types of Observability

    The most frequently mentioned type is data observability, which is concerned with the health and quality of data passing through data pipelines for analytical use cases like feeding a data warehouse. Vendors like Acceldata, Bigeye or Monte Carlo Data typically target their products to data engineers tasked with building and operating analytical data pipelines. 

    Another group of companies addresses applications. These firms, like Honeycomb and Observe.ai, collect data from applications to help site reliability engineers (SREs) understand performance issues and aid them in debugging and troubleshooting. 

    There are also companies specializing in machine learning observability, and these are concerned with the performance and drift of models in production. These companies, like Arize and WhyLabs, target data scientists. 

    Finally, there are companies in the observability data space which is comprised of logs, events, metrics and traces essential for all other forms of observability and monitoring to work. Companies in this space, like Cribl, use specially-designed pipelines to connect the sources and destinations of observability and security data. These pipelines allow companies to route data to multiple destinations, enrich data in flight and reduce data volumes before ingestion.

    As you can see, there are dozens of companies in the space. Each category of vendors addresses a unique challenge in today’s sprawling IT and security landscape.

    Beware Observability-Washing

    As the hype around this topic grows, companies hoping to differentiate themselves from their competitors may suddenly rebrand as observability companies. This is already happening. Legacy monitoring and alerting companies are calling themselves observability companies, hoping to attract the same kind of attention as others in the space that actually do have differentiated products and positioning.

  • Filter the Firehose

    Filter the Firehose

    We are tired. Information overload is a problem in the modern world. We hear instantly about events we never would have known about otherwise, or that we would have learned about months after the fact. Today, moments after an event, we have thousands of “professionals” analyzing it for us, a millions-strong army of amateurs telling us they know everything about it and a legion of bots telling us what to think about the event. No matter what the event is. Often “the event” isn’t even newsworthy, and yet the process feeds on itself.

    We in IT have it worse. We have a flood of security and application events constantly flooding in and so much data that we call it a ‘data lake.’ More like a data ocean at some organizations. We’re tired. They call it by various names—alert fatigue, data overload, event flood, “the firehose.”

    No matter what it’s called, it’s a big problem. And if it is not effectively managed, we will make mistakes. Particularly in security—but also across IT—we cannot afford to allow the flood to distract us or burn us out.

    Look into AI. The amount of data, alerts, information, metrics, logs, traces, etc. that we are accumulating—and the rate at which we are doing so—is growing faster than humans can reasonably keep up with—even if there were no staffing constraints. Don’t hesitate any longer; don’t assume you or a vendor NOC are managing this well enough with only eyeballs. AI solutions have matured enough to reliably sift through at least the top layer(s) of noise, reducing the burden without increasing risk. In fact, given the increasing possibility of human error as volumes increase, probably with less risk.

    Ask vendors about it; talk to your security vendors about their offerings. A huge number of security vendors have NOCs now that they would love to have you subscribe to. Many of them are highly automated and use AI. Data vendors are behind a bit in this regard (for a lot of reasons) but they are catching up. Not too long ago, the vast majority of work done to normalize disparate datasets was done via trial and error. AI and scripts written from massive experience are lightening that load, at least.

    So check it out. Find ways to make things manageable again. We all know things are slipping through the cracks as the volume of items reported and the number of locations they are reported from increases. Don’t wait until you have another emergency; look into the options and get something running.

    For those of you already using AI/ML, unless you implemented it very recently, you’ll probably want to validate what you’re using, at what level and how much the AI/ML solution is helping, and tweak or replace. This space is moving too fast for “set it and forget it” today. Maybe in a few years, we can.

    With the rare exception of something I use every day and think you might benefit from, I adhere to a policy of not recommending products here, so I can’t point you in the right direction other than to say that your vendors (data, SIEM and security) can offer solutions and/or suggestions. That, and “You shouldn’t wait” are my suggestions.

    And keep rocking it. Part of that data/log/alerting growth is all the apps and infrastructure you have helped put in and that are critical to the organization. So protect it with every tool available—and use that free time to work more miracles.

  • Quality Is a Top Challenge for Data-Driven Projects

    Quality Is a Top Challenge for Data-Driven Projects

    Software systems continue to produce more and more data. And making use of it has proven benefits — so much so that many analysts have, over the years, referred to data as the new oil. As a result, the majority of organizations are expending effort into refining their data — in fact, a recent study found that 84% of organizations have either already deployed or have data-driven projects on their roadmaps.

    Corporations like Facebook and Google are the poster children for business models that harvest and monetize end-user data. However, this is only one aspect; valuable data is being produced by internal systems as well, which, if leveraged correctly, can provide insight into software observability and increase process automation for DevOps teams and developers. For example, time-stamped application logs are necessary for informing SLOs to maintain reliability standards. There is an ongoing parallel investment into AI/ML deployed with cloud-native tools to act upon production data to drive further business growth. As such, 63% of IT decision-makers say new revenue opportunities have emerged due to data and analytics.

    The 2022 Data & Analytics Study, conducted by Foundry (formerly IDG Communications), explored how data-driven initiatives continue to be an important investment area for executive leadership. Below, I’ll review the key results from the survey to consider how organizations can continue to make intelligent decisions now and into the near future.

    The State of Data-Driven Investments

    We undoubtedly live in a data-driven world, and investment in related projects continues to rise. More than half (55%) of IT decision-makers plan to increase their data-focused investments — the report found the average spend to be $12.3 million in the coming year. This figure rises to an average of $23 million for financial services, which makes sense given the nature of modern banking.

    Naturally, the use of analytics platforms is increasing in tandem with the amount of data generated. In terms of specific analytics tools, 50% of respondents said they used business intelligence platforms while 47% used relational databases and 19% planned to invest in them in the next one to two years.

    We’re also noticing an increasing use of cloud-based solutions. For example, 27% of an organization’s data analytics workloads now run in the cloud. The report also found an uptick in cloud-based enterprise-scale data warehouse technologies, an area that shows signs of increasing in the coming year.

    Common Objectives

    So, what are the most common driving factors behind these types of projects? Well, automating internal business processes is the top goal for data-driven projects—half of IT leaders described this as a primary objective. This is closely followed by other ambitions such as improving customer insights (46%), aiding customer support (43%) and automating IT operations (43%).

    In terms of type, transactional data tends to be the most useful—54% of companies are using transactional data in their data-driven projects. This includes consumer purchasing behaviors such as sales, returns and credits. The next-most-common type is machine-generated data, which includes information from logs, sensors, telemetry, networks, security systems and/or IoT devices. The third-most-common type is customer profile information.

    Data has a powerful impact in a business context as a means to refine existing digital products. Other respondents added that a data-driven approach to DevOps helps increase visibility and drives continuous improvement.

    Data Quality: Highest Ranked Challenge

    Undoubtedly, particular challenges remain. The greatest hurdle is retaining data quality—41% of organizations reported dealing with this issue. Quality may be poor because data is unstructured or incompatible with other sources. Other widespread challenges include governance issues and integrating data from multiple sources.

    For those undertaking data-driven projects, 44% said they lacked appropriate skillsets, such as analytics training, data management, security, business intelligence and integration expertise. Companies also often faced funding and talent-related quandaries when beginning data-driven projects. For example, 26% of small-to-medium-sized businesses lacked the necessary funding to take on data-driven projects, according to the research.

    Another challenging area (that I’ve covered previously) is optimizing how data is collected and stored. As cloud storage fees rise, organizations will likely want to refine retention and resolution to avoid creating unnecessarily large, expensive data lakes.

    Rising Deployment of AI/ML

    AI/ML is already well-established as a way to advance the use of data. More than half (54%) of companies used predictive analytics or planned to incorporate it into their systems in the next 12 months. Just under one-third (31%) also either used or planned to use anomaly detection in the future. These areas along with other functionalities (such as natural language processing and predictive analytics) can be used to aid applications such as process automation, decision support, customer analysis, virtual agents and others.

    The industry is at an exciting moment for leveraging data in DevOps. Data is becoming more and more accessible to enable DevOps with a full-stack picture of the application life cycle. To direct future engineering investment, data-driven decisions will rely on things like performances and usage habits. Thus, it’s an interesting time to consider how you might leverage data for greater process automation and software development fluidity.

    The 2022 Data & Analytics Report surveyed 872 IT decision-makers (ITDMs) from around the globe working in various industries. The respondents’ average company size was about 12,000 employees. To view the report for more insights, you can pick it up here.