There’s a good reason Datadog is one of the most popular monitoring solutions available. The power of the platform is summed up in the tagline, “See inside any stack, any app, at any scale, anywhere” and explained in this chart: “Datadog brings together end-to-end traces, metrics, and logs to make your applications, infrastructure, and third-party… Continue reading Using Datadog For Observability? Speed up Troubleshooting with Zebrium
Category: AI/ML
Using Datadog For Observability? Speed up Troubleshooting with Zebrium
Application monitoring is experiencing a sea change. You can feel it as vendors rush to include the phrase “root cause” in their marketing boilerplate. Common solutions enhance telemetry collection and streamline workflows, but that’s not enough anymore. Autonomous troubleshooting is becoming a critical (but largely absent) capability for meeting SLOs, while at the same time,… Continue reading Observability: It’s Time to Automate the Observer
Observability: It’s Time to Automate the Observer
Native machine learning for ElasticSearch was first introduced as an Elastic Stack (ELK Stack) feature in 2017. It came from Elastic’s acquisition of Prelert, and was designed for anomaly detection in time series metrics data. The Elastic ML technology has since evolved to include anomaly detection for log data. So why is a new approach… Continue reading Elasticsearch Machine Learning -An Improved Approach Using Correlated Anomaly Detection To Find Root Cause
Elasticsearch Machine Learning -An Improved Approach Using Correlated Anomaly Detection To Find Root Cause
When a new/unknown software problem occurs, chances are an SRE or developer will start by analyzing and searching through logs for root cause – a slow and painful process. So it’s no wonder using machine learning (ML) for log analysis is getting a lot of attention. But machine learning (ML) with logs is hard. Here’s… Continue reading Log Analysis with Machine Learning: An Automated Approach to Analyzing Logs Using ML/AI
Log Analysis with Machine Learning: An Automated Approach to Analyzing Logs Using ML/AI
If you are a New Relic user, you’re likely using New Relic to monitor your environment, detect problems, and troubleshoot them when they occur. But let’s consider exactly what that entails and describe a way to make this entire process much quicker. Imagine that the dashboards used to monitor your application suddenly show a “blip”.… Continue reading Using New Relic For Observability? Speed up Troubleshooting with Zebrium
Using New Relic For Observability? Speed up Troubleshooting with Zebrium
The Elastic Stack (often called ELK) is one of the most popular observability platforms in use today. It lets you collect metrics, traces, and logs and visualize them in one Kibana dashboard. You can set alerts for outliers, drill down into your dashboards and search through your logs. But there are limitations. What happens when… Continue reading Using the Elastic Stack (ELK) For Observability? Here’s How to Speed Up Troubleshooting
Using the Elastic Stack (ELK) For Observability? Here’s How to Speed Up Troubleshooting
A few weeks ago, Larry Lancaster, wrote about a new beta feature leveraging the GPT-3 language model – Using GPT-3 for plain language incident root cause from logs. To recap – Zebrium’s unsupervised ML identifies the root cause of incidents and generates concise reports (typically between 5-20 log events) identifying the first event in the… Continue reading Real World Examples of GPT-3 Plain Language Root Cause Summaries
Real World Examples of GPT-3 Plain Language Root Cause Summaries
We believe the future of monitoring, especially for platforms like Kubernetes, is truly autonomous. Cloud-native applications are increasingly distributed, evolving faster, and failing in new ways, making it harder to monitor, troubleshoot and resolve incidents. Traditional approaches such as dashboards, carefully tuned alert rules, and searches through logs are reactive and time intensive, hurting productivity,… Continue reading Anomaly Detection as a Foundation of Autonomous Monitoring
Anomaly Detection as a Foundation of Autonomous Monitoring
This project is a favorite of mine and so I wanted to share a glimpse of what we’ve been up to with OpenAI’s amazing GPT-3 language model. Today I’ll be sharing a couple of straightforward results. There are more advanced avenues we’re exploring for our use of GPT-3, such as fine-tuning (custom pre-training for specific… Continue reading Using GPT-3 for plain language incident root cause from logs
Using GPT-3 for plain language incident root cause from logs
Entertainer Jim Stafford had a hit song in 1974 entitled “Spiders and Snakes.” I was thinking that would be a good song to put on a playlist for the ScienceLogic crew that takes the road trip to the Department of Defense Test Integration Center (TIC) in Fort Huachuca, located in the desert of southeast Arizona.… Continue reading ScienceLogic’s DoDIN APL Certification Journey: Watch out for Spiders & Snakes



