
AI handles incidents, engineers lose touch with their systems
AI-assisted incident response can lower MTTR while leaving engineers less prepared for the complex incidents automation cannot solve.
Topic
My writing, talks, podcasts, and projects about Incident Management.

AI-assisted incident response can lower MTTR while leaving engineers less prepared for the complex incidents automation cannot solve.
AI-generated code is outpacing traditional runbooks. How SRE teams are adapting their incident response playbooks for the LLM era.
AI-assisted coding is exploding — but acceleration doesn't mean reliability. Real examples of hard-to-trace LLM bugs, hallucinated dependencies, and operational fallout for lean SRE teams.
AI-generated code ships faster but introduces subtle bugs that are harder to trace. What engineering leaders need to know about the reliability trade-offs.
Sylvain Kalache went looking for the new ways AI breaks software. Every one turned out to be a habit we've had for years.
AI vs. AI — how autonomous attackers are scaling phishing and malware, how defenders can use agents to cut alert noise, and where accountability belongs when AI investigates and remediates incidents.
Speaking on the Incident Fest Main Stage about staying reliable as AI-assisted development increases the rate at which software is created.
Why Twingate stopped shipping features to rebuild reliability — from active-active multi-region architecture and smaller blast radiuses to preserving human ownership as agentic coding accelerates delivery.
Meta-hosted summit in Menlo Park focused on applying AI, LLMs, and agents to software incidents and response.
Every pilot trains for engine failure, but most engineers face their first major incident unprepared — why communication, coordination, and decision-making need deliberate simulation practice.
Why traditional burnout detection fails in engineering teams. Connecting Maslach's burnout research and the Copenhagen Burnout Inventory to observable on-call metrics, and demonstrating that observed data outperforms subjective self-reporting.
The golden hour — why the first 15 minutes of an incident decide everything, from mitigation strategy to customer trust, and why human psychology matters more than tools.
From 600 to 6,000 — an 18-month transformation replacing legacy tooling with a purpose-built incident management platform, scaling operational responsibility from 600 SREs to 6,000 engineers across the org.
Reliability Rebels Guest appearance to discuss reliability engineering, AI, and modern incident response.
From faster detection to smarter debugging — where AI delivers practical value in incident response, its limitations in production environments, and how teams can integrate AI while maintaining reliability standards.
LLM-assisted development introduces hard-to-trace bugs, AI-generated tests that mirror flawed logic, and hallucinated dependencies. Exploring operational consequences and actionable strategies from AI-powered incident tools to incident vibing.
The incident you never had — deterministic simulation testing, why conventional testing misses bugs that cause real outages, and how simulation-based approaches improve software reliability.
The science behind responder overload and introducing On-Call Health, an open-source tool for detecting incident responder fatigue. At Google NYC Pier 57.
Burnout doesn't ask permission — the physiology of burnout from an incident responder's perspective, warning signs, recovery, and the systemic changes the industry must make.
An evening of AI demos and networking in Toronto — presenting Rootly AI alongside teams from Sentry, Red Brick Labs, Convex, and DigitalOcean.
Democratizing reliability — empowering non-engineers with operational power, incident management as a muscle, and AI-powered postmortems.
Panel at the AWS GenAI Loft in San Francisco on MCPs, incident automation, observability, and generative AI tooling for reliability.
Live demos at Okta HQ in San Francisco — presenting Rootly AI alongside cybersecurity startups tackling identity, endpoint security, and threat response.
Interview with Alan Shimel at PlatformCon NYC 2025 on how AI applies to incident management and reliability engineering — triage, root cause analysis, and why AI enhances rather than replaces engineers.
Invite-only roundtable on filtering alert noise, enriching alerts with automated ownership context, and streamlining bug validation workflows.
Invite-only roundtable comparing centralized specialist responders with distributed on-call models — which strategy works for your org.
Tracing the arc from ingesting logs at LinkedIn/SlideShare to LLM-driven RCA today. How fine-tuning, MCP, and incident vibing are reshaping SRE.
If developers are vibe coding, SREs are now incident vibing. What happens when AI-generated code meets production reality.
Training teams for incident response — chaos engineering, game days, and building organizational muscle memory.
Invite-only roundtable examining AI's role in incident response — separating practical applications from hype with industry leaders.
An MCP server for Rootly — enabling AI agents to interact with incident management workflows via the Model Context Protocol.
Reliability at GitHub scale — incident management, deployment practices, and engineering culture.
A benchmark suite for evaluating AI agents on real-world SRE tasks — incident diagnosis, runbook execution, and infrastructure troubleshooting.
Detects early warning signs of on-call engineer burnout by pulling data from Rootly, PagerDuty, Linear, GitHub, and Slack.
A grounded look at what LLMs can and can't do in an incident response workflow today, drawn from real production experiments at Rootly AI Labs.
A postmortem on a DevOps failure that affected 60,000 users — lessons in monitoring gaps, deployment practices, and incident response.