AlertD this week emerged from stealth to launch a DevOps platform that leverages generative artificial intelligence (AI) to provide deeper insights into Amazon Web Services (AWS) environments by automatically generating […]
Using AI to Improve Operational Efficiency
https://drive.google.com/file/d/1KCSqX4cK66EAD8MoucZQl0WPyDDBMS4Z/view?usp=sharingThe sheer volume and complexity of data available for businesses to observe, store and use is overwhelming. Knowing how to apply AI in practical ways to respond promptly to operational […]
How SREs Benefit From Feature Flags
When you think of who uses feature flags, your mind most likely goes to software developers. In general, feature flags are closely associated with software engineering. But site reliability engineers […]
Top Nine Skills for SREs to Master
It’s easy to talk at a high level about what site reliability engineers do: They ensure that IT systems achieve availability and performance requirements. But which skills, exactly, do SREs […]
Service Level Objectives – Techstrong TV
Service Level Objectives (SLOs) are being more widely discussed as they provide a clear path to delivering exceptional results by defining clear reliability goals. Kit provides updates on the community, […]
Reflections: A DevOps Christmas Carol
“Marley was dead, to begin with.” So begins my DevOps Christmas Carol. In this reflection, I am visited by the ghosts of DevOps past, present and future. I hope that, […]
Kentik Open Sources Five Projects to Advance Network Observability
Kentik today unfurled a developer hub, dubbed Kentik Labs, through which it is making five open source projects available as part of an effort to advance observability across network operations […]
How to Build an Options-Based Observability Strategy
The global pandemic forced companies to accelerate their digital transformations, pushing them toward new technologies and architectures as they struggled to adapt to a changing world. Companies wanted fast development […]
Fixing Risk Sharing With Observability
Incentives are mismatched among SREs, SecOps, and application developers. These mismatches create challenges around how and what information is shared across siloed teams. This asymmetrical information creates a moral hazard […]
Incident Resolution for Remote Teams
People working in IT support and incident management right now are faced with unusual difficulties supporting large remote workforces and managing unpredictable workloads. On Reddit, system admins and other IT […]











