Topic
Incident Management
My writing, talks, podcasts, and projects about Incident Management.
Selected work
LLMs Broke the SRE Runbook. Now What?
AI-generated code is outpacing traditional runbooks. How SRE teams are adapting their incident response playbooks for the LLM era.
From Vibes to Outages: Riding the AI Code Wave
AI-assisted coding is exploding — but acceleration doesn't mean reliability. Real examples of hard-to-trace LLM bugs, hallucinated dependencies, and operational fallout for lean SRE teams.
Is AI-assisted coding an incident magnet?
AI-generated code ships faster but introduces subtle bugs that are harder to trace. What engineering leaders need to know about the reliability trade-offs.
AI Fails Like We Do
Sylvain Kalache went looking for the new ways AI breaks software. Every one turned out to be a habit we've had for years.
Nir Soudry — Head of R&D at 7AI
AI vs. AI — how autonomous attackers are scaling phishing and malware, how defenders can use agents to cut alert noise, and where accountability belongs when AI investigates and remediates incidents.
More Code, More Incidents? Staying Reliable When AI Writes the Code
Speaking on the Incident Fest Main Stage about staying reliable as AI-assisted development increases the rate at which software is created.
Eran Kampf — VP of Engineering at Twingate
Why Twingate stopped shipping features to rebuild reliability — from active-active multi-region architecture and smaller blast radiuses to preserving human ownership as agentic coding accelerates delivery.
AIOps Summit
Meta-hosted summit in Menlo Park focused on applying AI, LLMs, and agents to software incidents and response.
Hamed Silatani — Co-Founder and CEO at Uptime Labs
Every pilot trains for engine failure, but most engineers face their first major incident unprepared — why communication, coordination, and decision-making need deliberate simulation practice.
The Science of On-Call Burnout: Why "How Are You Doing?" Always Gets "Fine"
Why traditional burnout detection fails in engineering teams. Connecting Maslach's burnout research and the Copenhagen Burnout Inventory to observable on-call metrics, and demonstrating that observed data outperforms subjective self-reporting.
Gandhi Kumar — Principal Incident Commander at Twilio
The golden hour — why the first 15 minutes of an incident decide everything, from mitigation strategy to customer trust, and why human psychology matters more than tools.
Cliff Snyder — Senior SRE at Multimedia LLC
From 600 to 6,000 — an 18-month transformation replacing legacy tooling with a purpose-built incident management platform, scaling operational responsibility from 600 SREs to 6,000 engineers across the org.
Reliability Rebels Podcast — Guest Appearance
Reliability Rebels Guest appearance to discuss reliability engineering, AI, and modern incident response.
How AI is Changing Incident Response (SRE Deep Dive)
From faster detection to smarter debugging — where AI delivers practical value in incident response, its limitations in production environments, and how teams can integrate AI while maintaining reliability standards.
From Vibes to Outages: When AI Writes the Code You Debug
LLM-assisted development introduces hard-to-trace bugs, AI-generated tests that mirror flawed logic, and hallucinated dependencies. Exploring operational consequences and actionable strategies from AI-powered incident tools to incident vibing.
Will Wilson — CEO at Antithesis
The incident you never had — deterministic simulation testing, why conventional testing misses bugs that cause real outages, and how simulation-based approaches improve software reliability.
Google SRE NYC Tech Talk — On-Call Health
The science behind responder overload and introducing On-Call Health, an open-source tool for detecting incident responder fatigue. At Google NYC Pier 57.
Stephen Townshend — SRE Team Lead at isDown
Burnout doesn't ask permission — the physiology of burnout from an incident responder's perspective, warning signs, recovery, and the systemic changes the industry must make.
OpenClaw Demo Night w/ Rootly AI, Convex, Sentry & DigitalOcean
An evening of AI demos and networking in Toronto — presenting Rootly AI alongside teams from Sentry, Red Brick Labs, Convex, and DigitalOcean.
Dileshni Jayasinghe — VP of Technology at commonsku
Democratizing reliability — empowering non-engineers with operational power, incident management as a muscle, and AI-powered postmortems.
MCPs and the Next Wave of Reliability w/ Rootly AI, WorkOS, Block, Microsoft & Groq
Panel at the AWS GenAI Loft in San Francisco on MCPs, incident automation, observability, and generative AI tooling for reliability.
AI Security Demo Night w/ Rootly AI, Okta, Panther, Tailscale & More
Live demos at Okta HQ in San Francisco — presenting Rootly AI alongside cybersecurity startups tackling identity, endpoint security, and threat response.
Exploring AI's Role in Incident Management
Interview with Alan Shimel at PlatformCon NYC 2025 on how AI applies to incident management and reliability engineering — triage, root cause analysis, and why AI enhances rather than replaces engineers.
Rootly Roundtable: From Weak Signals to Confident Fixes
Invite-only roundtable on filtering alert noise, enriching alerts with automated ownership context, and streamlining bug validation workflows.
Rootly Roundtable: Centralized vs. Distributed Incident Teams
Invite-only roundtable comparing centralized specialist responders with distributed on-call models — which strategy works for your org.
Incident Vibing: The Self-Healing System
Tracing the arc from ingesting logs at LinkedIn/SlideShare to LLM-driven RCA today. How fine-tuning, MCP, and incident vibing are reshaping SRE.
Vibe Coding Is Here — But Are You Ready for Incident Vibing?
If developers are vibe coding, SREs are now incident vibing. What happens when AI-generated code meets production reality.
Dan Slimmon — Incident Response Trainer
Training teams for incident response — chaos engineering, game days, and building organizational muscle memory.
Rootly Roundtable: The State of AI in Incident Management
Invite-only roundtable examining AI's role in incident response — separating practical applications from hype with industry leaders.
Rootly-MCP-server
An MCP server for Rootly — enabling AI agents to interact with incident management workflows via the Model Context Protocol.
Sean Goedecke — Staff Software Engineer at GitHub
Reliability at GitHub scale — incident management, deployment practices, and engineering culture.
SRE-skills-bench
A benchmark suite for evaluating AI agents on real-world SRE tasks — incident diagnosis, runbook execution, and infrastructure troubleshooting.
On-Call-Health
Detects early warning signs of on-call engineer burnout by pulling data from Rootly, PagerDuty, Linear, GitHub, and Slack.
AI-Driven Incident Resolution — Hype or Reality
A grounded look at what LLMs can and can't do in an incident response workflow today, drawn from real production experiments at Rootly AI Labs.
How DevOps Failed 60K Users
A postmortem on a DevOps failure that affected 60,000 users — lessons in monitoring gaps, deployment practices, and incident response.