Back to all work

Topic

Incident Management

My writing, talks, podcasts, and projects about Incident Management.

Selected work

Thumbnail for LLMs Broke the SRE Runbook. Now What?
Article
The New Stack
The New Stack

LLMs Broke the SRE Runbook. Now What?

AI-generated code is outpacing traditional runbooks. How SRE teams are adapting their incident response playbooks for the LLM era.

Thumbnail for From Vibes to Outages: Riding the AI Code Wave
Talk
SREcon EMEA 2025
SREcon EMEA 2025

From Vibes to Outages: Riding the AI Code Wave

AI-assisted coding is exploding — but acceleration doesn't mean reliability. Real examples of hard-to-trace LLM bugs, hallucinated dependencies, and operational fallout for lean SRE teams.

Thumbnail for Is AI-assisted coding an incident magnet?
Article
LeadDev
LeadDev

Is AI-assisted coding an incident magnet?

AI-generated code ships faster but introduces subtle bugs that are harder to trace. What engineering leaders need to know about the reliability trade-offs.

Thumbnail for AI Fails Like We Do
Podcast
Engineering Futures
Engineering Futures

AI Fails Like We Do

Sylvain Kalache went looking for the new ways AI breaks software. Every one turned out to be a habit we've had for years.

Thumbnail for Nir Soudry — Head of R&D at 7AI
Podcast
Humans of Reliability
Humans of Reliability

Nir Soudry — Head of R&D at 7AI

AI vs. AI — how autonomous attackers are scaling phishing and malware, how defenders can use agents to cut alert noise, and where accountability belongs when AI investigates and remediates incidents.

Thumbnail for More Code, More Incidents? Staying Reliable When AI Writes the Code
Talk
Incident Fest
Incident Fest 2026

More Code, More Incidents? Staying Reliable When AI Writes the Code

Speaking on the Incident Fest Main Stage about staying reliable as AI-assisted development increases the rate at which software is created.

Thumbnail for Eran Kampf — VP of Engineering at Twingate
Podcast
Humans of Reliability
Humans of Reliability

Eran Kampf — VP of Engineering at Twingate

Why Twingate stopped shipping features to rebuild reliability — from active-active multi-region architecture and smaller blast radiuses to preserving human ownership as agentic coding accelerates delivery.

Thumbnail for AIOps Summit
Talk
AIOps Summit
AIOps Summit 2026

AIOps Summit

Meta-hosted summit in Menlo Park focused on applying AI, LLMs, and agents to software incidents and response.

Thumbnail for Hamed Silatani — Co-Founder and CEO at Uptime Labs
Podcast
Humans of Reliability
Humans of Reliability

Hamed Silatani — Co-Founder and CEO at Uptime Labs

Every pilot trains for engine failure, but most engineers face their first major incident unprepared — why communication, coordination, and decision-making need deliberate simulation practice.

Thumbnail for The Science of On-Call Burnout: Why "How Are You Doing?" Always Gets "Fine"
Talk
DevOpsDays Austin
DevOpsDays Austin 2026

The Science of On-Call Burnout: Why "How Are You Doing?" Always Gets "Fine"

Why traditional burnout detection fails in engineering teams. Connecting Maslach's burnout research and the Copenhagen Burnout Inventory to observable on-call metrics, and demonstrating that observed data outperforms subjective self-reporting.

Thumbnail for Gandhi Kumar — Principal Incident Commander at Twilio
Podcast
Humans of Reliability
Humans of Reliability

Gandhi Kumar — Principal Incident Commander at Twilio

The golden hour — why the first 15 minutes of an incident decide everything, from mitigation strategy to customer trust, and why human psychology matters more than tools.

Thumbnail for Cliff Snyder — Senior SRE at Multimedia LLC
Podcast
Humans of Reliability
Humans of Reliability

Cliff Snyder — Senior SRE at Multimedia LLC

From 600 to 6,000 — an 18-month transformation replacing legacy tooling with a purpose-built incident management platform, scaling operational responsibility from 600 SREs to 6,000 engineers across the org.

Thumbnail for Reliability Rebels Podcast — Guest Appearance
Podcast
Reliability Rebels
Reliability Rebels

Reliability Rebels Podcast — Guest Appearance

Reliability Rebels Guest appearance to discuss reliability engineering, AI, and modern incident response.

Thumbnail for How AI is Changing Incident Response (SRE Deep Dive)
Podcast
The Rollback Café
The Rollback Café

How AI is Changing Incident Response (SRE Deep Dive)

From faster detection to smarter debugging — where AI delivers practical value in incident response, its limitations in production environments, and how teams can integrate AI while maintaining reliability standards.

Thumbnail for From Vibes to Outages: When AI Writes the Code You Debug
Talk
Conf42 Site Reliability Engineering 2026

From Vibes to Outages: When AI Writes the Code You Debug

LLM-assisted development introduces hard-to-trace bugs, AI-generated tests that mirror flawed logic, and hallucinated dependencies. Exploring operational consequences and actionable strategies from AI-powered incident tools to incident vibing.

Thumbnail for Will Wilson — CEO at Antithesis
Podcast
Humans of Reliability
Humans of Reliability

Will Wilson — CEO at Antithesis

The incident you never had — deterministic simulation testing, why conventional testing misses bugs that cause real outages, and how simulation-based approaches improve software reliability.

Thumbnail for Google SRE NYC Tech Talk — On-Call Health
Talk
Google SRE NYC
Google SRE NYC Tech Talks

Google SRE NYC Tech Talk — On-Call Health

The science behind responder overload and introducing On-Call Health, an open-source tool for detecting incident responder fatigue. At Google NYC Pier 57.

Thumbnail for Stephen Townshend — SRE Team Lead at isDown
Podcast
Humans of Reliability
Humans of Reliability

Stephen Townshend — SRE Team Lead at isDown

Burnout doesn't ask permission — the physiology of burnout from an incident responder's perspective, warning signs, recovery, and the systemic changes the industry must make.

Thumbnail for OpenClaw Demo Night w/ Rootly AI, Convex, Sentry & DigitalOcean
Talk
OpenClaw Demo Night

OpenClaw Demo Night w/ Rootly AI, Convex, Sentry & DigitalOcean

An evening of AI demos and networking in Toronto — presenting Rootly AI alongside teams from Sentry, Red Brick Labs, Convex, and DigitalOcean.

Thumbnail for Dileshni Jayasinghe — VP of Technology at commonsku
Podcast
Humans of Reliability
Humans of Reliability

Dileshni Jayasinghe — VP of Technology at commonsku

Democratizing reliability — empowering non-engineers with operational power, incident management as a muscle, and AI-powered postmortems.

Thumbnail for MCPs and the Next Wave of Reliability w/ Rootly AI, WorkOS, Block, Microsoft & Groq
Panel Moderation
Rootly AI Labs
Rootly AI Labs

MCPs and the Next Wave of Reliability w/ Rootly AI, WorkOS, Block, Microsoft & Groq

Panel at the AWS GenAI Loft in San Francisco on MCPs, incident automation, observability, and generative AI tooling for reliability.

Thumbnail for AI Security Demo Night w/ Rootly AI, Okta, Panther, Tailscale & More
Talk
AI Security Demo Night

AI Security Demo Night w/ Rootly AI, Okta, Panther, Tailscale & More

Live demos at Okta HQ in San Francisco — presenting Rootly AI alongside cybersecurity startups tackling identity, endpoint security, and threat response.

Thumbnail for Exploring AI's Role in Incident Management
Podcast
Techstrong TV
Techstrong TV

Exploring AI's Role in Incident Management

Interview with Alan Shimel at PlatformCon NYC 2025 on how AI applies to incident management and reliability engineering — triage, root cause analysis, and why AI enhances rather than replaces engineers.

Thumbnail for Rootly Roundtable: From Weak Signals to Confident Fixes
Panel Moderation
Rootly
Rootly

Rootly Roundtable: From Weak Signals to Confident Fixes

Invite-only roundtable on filtering alert noise, enriching alerts with automated ownership context, and streamlining bug validation workflows.

Thumbnail for Rootly Roundtable: Centralized vs. Distributed Incident Teams
Panel Moderation
Rootly
Rootly

Rootly Roundtable: Centralized vs. Distributed Incident Teams

Invite-only roundtable comparing centralized specialist responders with distributed on-call models — which strategy works for your org.

Thumbnail for Incident Vibing: The Self-Healing System
Podcast
Adventures in DevOps
Adventures in DevOps

Incident Vibing: The Self-Healing System

Tracing the arc from ingesting logs at LinkedIn/SlideShare to LLM-driven RCA today. How fine-tuning, MCP, and incident vibing are reshaping SRE.

Thumbnail for Vibe Coding Is Here — But Are You Ready for Incident Vibing?
Article
The New Stack
The New Stack

Vibe Coding Is Here — But Are You Ready for Incident Vibing?

If developers are vibe coding, SREs are now incident vibing. What happens when AI-generated code meets production reality.

Thumbnail for Dan Slimmon — Incident Response Trainer
Podcast
Humans of Reliability
Humans of Reliability

Dan Slimmon — Incident Response Trainer

Training teams for incident response — chaos engineering, game days, and building organizational muscle memory.

Thumbnail for Rootly Roundtable: The State of AI in Incident Management
Panel Moderation
Rootly
Rootly

Rootly Roundtable: The State of AI in Incident Management

Invite-only roundtable examining AI's role in incident response — separating practical applications from hype with industry leaders.

Thumbnail for Rootly-MCP-server
Project
GitHub
GitHub

Rootly-MCP-server

An MCP server for Rootly — enabling AI agents to interact with incident management workflows via the Model Context Protocol.

39
Thumbnail for Sean Goedecke — Staff Software Engineer at GitHub
Podcast
Humans of Reliability
Humans of Reliability

Sean Goedecke — Staff Software Engineer at GitHub

Reliability at GitHub scale — incident management, deployment practices, and engineering culture.

Thumbnail for SRE-skills-bench
Project
GitHub
GitHub

SRE-skills-bench

A benchmark suite for evaluating AI agents on real-world SRE tasks — incident diagnosis, runbook execution, and infrastructure troubleshooting.

12
Thumbnail for On-Call-Health
Project
GitHub
GitHub

On-Call-Health

Detects early warning signs of on-call engineer burnout by pulling data from Rootly, PagerDuty, Linear, GitHub, and Slack.

32
Thumbnail for AI-Driven Incident Resolution — Hype or Reality
Talk
DevOpsDays Austin
DevOpsDays Austin

AI-Driven Incident Resolution — Hype or Reality

A grounded look at what LLMs can and can't do in an incident response workflow today, drawn from real production experiments at Rootly AI Labs.

Thumbnail for How DevOps Failed 60K Users
Article
Linux.com
Linux.com

How DevOps Failed 60K Users

A postmortem on a DevOps failure that affected 60,000 users — lessons in monitoring gaps, deployment practices, and incident response.