<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Monitoring on GeekyRyan</title><link>https://rnemeth90.github.io/tags/monitoring/</link><description>Recent content in Monitoring on GeekyRyan</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Thu, 05 Sep 2024 09:00:00 +0000</lastBuildDate><atom:link href="https://rnemeth90.github.io/tags/monitoring/index.xml" rel="self" type="application/rss+xml"/><item><title>Building an Operational Strategy for Understanding and Responding to Alerts</title><link>https://rnemeth90.github.io/projects/2024-09-05-operational-alerting-strategy/</link><pubDate>Thu, 05 Sep 2024 09:00:00 +0000</pubDate><guid>https://rnemeth90.github.io/projects/2024-09-05-operational-alerting-strategy/</guid><description>&lt;p&gt;You can have all the alerts configured you want, but if the team doesn&amp;rsquo;t agree on what they mean, who&amp;rsquo;s responsible for them, and what a good response looks like, they&amp;rsquo;re just noise. This was about building that operational strategy - making sure every alert had a clear owner, a documented playbook, and an agreed severity/escalation path.&lt;/p&gt;&#10;&lt;p&gt;It&amp;rsquo;s the human side of alerting, not the technical side. Once you&amp;rsquo;ve got the right metrics feeding into alerts, you need the operational discipline to make sure those alerts actually get owned and responded to consistently, whoever&amp;rsquo;s on call.&lt;/p&gt;</description></item><item><title>Closing Gaps in Infrastructure and Application Monitoring</title><link>https://rnemeth90.github.io/projects/2024-03-15-closing-gaps-in-infrastructure-monitoring/</link><pubDate>Fri, 15 Mar 2024 09:00:00 +0000</pubDate><guid>https://rnemeth90.github.io/projects/2024-03-15-closing-gaps-in-infrastructure-monitoring/</guid><description>&lt;p&gt;Infrastructure evolves fast, and monitoring coverage always lags behind. Full review of what we were actually monitoring to find and close the gaps in alerting and visibility.&lt;/p&gt;&#10;&lt;p&gt;Problem was simple: as our footprint grew, metrics coverage didn&amp;rsquo;t keep up. Blind spots on CPU, memory, disk, network throughput on various resources. Assessed everything, implemented missing alerting, wrote response procedures.&lt;/p&gt;&#10;&lt;p&gt;Pairing every new alert with &amp;ldquo;here&amp;rsquo;s what you do when this fires&amp;rdquo; was as important as the alert itself. An alert with no playbook is just noise.&lt;/p&gt;</description></item><item><title>Improving Release Environment Stability Through Proactive Monitoring</title><link>https://rnemeth90.github.io/projects/2023-11-30-improving-release-stability-and-visibility/</link><pubDate>Thu, 30 Nov 2023 09:00:00 +0000</pubDate><guid>https://rnemeth90.github.io/projects/2023-11-30-improving-release-stability-and-visibility/</guid><description>&lt;p&gt;Release environments get less monitoring love than production, which is backwards - instability there slows down every team trying to validate their work.&lt;/p&gt;&#10;&lt;p&gt;Hit a 95% stability target in the release/daily environments by monitoring proactively instead of waiting for complaints. Built visibility into the full change set landing in these environments, so when something broke we could correlate it against what actually changed instead of guessing.&lt;/p&gt;&#10;&lt;p&gt;Dashboard showing deployment metrics over time - failure rates, time-to-recovery, what kinds of changes cause instability. Once the visibility existed, we caught issues before they became &amp;ldquo;developer needs to file a ticket&amp;rdquo; problems. Real win was reducing how much environment instability was dragging on everyone&amp;rsquo;s velocity.&lt;/p&gt;</description></item></channel></rss>