Topic
SRE
My writing, talks, podcasts, and projects about SRE.
Selected work
Eduardo Ordax — Principal GTM GenAI at AWS
The reality of GenAI in production — why organizational culture is the biggest blocker, scaling non-deterministic LLM systems, and what separates AI winners from experimenters.
LLMs Broke the SRE Runbook. Now What?
AI-generated code is outpacing traditional runbooks. How SRE teams are adapting their incident response playbooks for the LLM era.
From Vibes to Outages: Riding the AI Code Wave
AI-assisted coding is exploding — but acceleration doesn't mean reliability. Real examples of hard-to-trace LLM bugs, hallucinated dependencies, and operational fallout for lean SRE teams.
Will LLMs and Vibe Coding Fuel a Developer Renaissance?
Exploring whether AI-assisted coding tools will democratize software development or create new categories of hard-to-debug production issues.
Is AI-assisted coding an incident magnet?
AI-generated code ships faster but introduces subtle bugs that are harder to trace. What engineering leaders need to know about the reliability trade-offs.
AI Fails Like We Do
Sylvain Kalache went looking for the new ways AI breaks software. Every one turned out to be a habit we've had for years.
More Code, More Incidents? Staying Reliable When AI Writes the Code
Speaking on the Incident Fest Main Stage about staying reliable as AI-assisted development increases the rate at which software is created.
Eran Kampf — VP of Engineering at Twingate
Why Twingate stopped shipping features to rebuild reliability — from active-active multi-region architecture and smaller blast radiuses to preserving human ownership as agentic coding accelerates delivery.
AIOps Summit
Meta-hosted summit in Menlo Park focused on applying AI, LLMs, and agents to software incidents and response.
Alexey Grigorev — Founder of DataTalks.Club
How a chain of reasonable-sounding decisions led an AI coding agent to run terraform destroy against a live production database — and the guardrails that separate moving fast from losing everything.

AI Made Developers 25% More Productive. It Also Tripled Our Incident Rate.
At AI DevSummit New York: how AI-assisted coding is reshaping software delivery — the productivity gains, the surge in incident rates, and what reliability teams need to do about it.
Hamed Silatani — Co-Founder and CEO at Uptime Labs
Every pilot trains for engine failure, but most engineers face their first major incident unprepared — why communication, coordination, and decision-making need deliberate simulation practice.
The Science of On-Call Burnout: Why "How Are You Doing?" Always Gets "Fine"
Why traditional burnout detection fails in engineering teams. Connecting Maslach's burnout research and the Copenhagen Burnout Inventory to observable on-call metrics, and demonstrating that observed data outperforms subjective self-reporting.
Gandhi Kumar — Principal Incident Commander at Twilio
The golden hour — why the first 15 minutes of an incident decide everything, from mitigation strategy to customer trust, and why human psychology matters more than tools.
Cliff Snyder — Senior SRE at Multimedia LLC
From 600 to 6,000 — an 18-month transformation replacing legacy tooling with a purpose-built incident management platform, scaling operational responsibility from 600 SREs to 6,000 engineers across the org.
Ganesh Datta — Co-Founder & CTO at Cortex
AI didn't change the game, it just exposed your bottlenecks — how platform engineering and SRE teams solve identical human problems through influence rather than authority, and why AI amplifies existing bottlenecks instead of transforming operations.
Reliability Rebels Podcast — Guest Appearance
Reliability Rebels Guest appearance to discuss reliability engineering, AI, and modern incident response.
Dana Lawson — CTO at Netlify
Fear, identity, and flaky tests — why SRE resistance to AI agents stems from identity and control concerns rather than the technology itself, and practical strategies for adopting AI-driven reliability tools starting with low-risk tasks.
The Software Leaders Uncensored Podcast
A candid conversation on software leadership, reliability engineering, and the evolving role of AI in modern engineering teams.
How AI is Changing Incident Response (SRE Deep Dive)
From faster detection to smarter debugging — where AI delivers practical value in incident response, its limitations in production environments, and how teams can integrate AI while maintaining reliability standards.
From Vibes to Outages: When AI Writes the Code You Debug
LLM-assisted development introduces hard-to-trace bugs, AI-generated tests that mirror flawed logic, and hallucinated dependencies. Exploring operational consequences and actionable strategies from AI-powered incident tools to incident vibing.
Will Wilson — CEO at Antithesis
The incident you never had — deterministic simulation testing, why conventional testing misses bugs that cause real outages, and how simulation-based approaches improve software reliability.
Google SRE NYC Tech Talk — On-Call Health
The science behind responder overload and introducing On-Call Health, an open-source tool for detecting incident responder fatigue. At Google NYC Pier 57.
Stephen Townshend — SRE Team Lead at isDown
Burnout doesn't ask permission — the physiology of burnout from an incident responder's perspective, warning signs, recovery, and the systemic changes the industry must make.
Swizec Teller — Bestselling Author
Code is cheap, reliability isn't — owning production in the AI era, the hidden complexity of SRE work, and why human ownership remains essential.
Dileshni Jayasinghe — VP of Technology at commonsku
Democratizing reliability — empowering non-engineers with operational power, incident management as a muscle, and AI-powered postmortems.
Tomás Hernando Koffman — Co-founder at Not Diamond
99%+ accuracy on a moving target — model deprecation, reliability with LLMs, and treating prompts as architectural components.
Julien Simon — VP and Chief Evangelist
A conversation about developer advocacy, AI evangelism, and building reliable systems at scale.
Conor Bronsdon — Head of Content at Galileo
Developer awareness, content strategy, and the human side of reliability engineering.
Tea, Pipelines, and Retries: A Practical Guide to MLOps
Discussion at SREcon EMEA 2025 on how AI is transforming the software development lifecycle — CI/CD pipelines, deployments, scaling, monitoring, incident management, reliability tooling, and emerging disciplines like LLMOps.
AI Meets Reliability
Panel exploring AI-driven automation and observability with leaders from NVIDIA, OpenAI, Baseten, Replit, and Weights & Biases on scaling operations and reducing MTTR.
Rob Zuber — CTO at CircleCI
The end of good code, AI throughput, and what reliability means at CI/CD scale.
MCPs and the Next Wave of Reliability w/ Rootly AI, WorkOS, Block, Microsoft & Groq
Panel at the AWS GenAI Loft in San Francisco on MCPs, incident automation, observability, and generative AI tooling for reliability.
Shery Brauner — SVP of Razor Group
Leading engineering at scale in e-commerce aggregation and the operational challenges that come with it.
Frontiers of AI: Building with Rootly AI, Zscaler, CircleCI, Fireworks AI & Google DeepMind
Panel at Google HQ with 300+ attendees exploring real-world Gemini models, reinforcement learning, next-gen agent systems, and AI reliability.
Brian Shaw — SVP of Infrastructure and Core Banking
Building and operating core banking infrastructure with a focus on uptime and regulatory compliance.
Exploring AI's Role in Incident Management
Interview with Alan Shimel at PlatformCon NYC 2025 on how AI applies to incident management and reliability engineering — triage, root cause analysis, and why AI enhances rather than replaces engineers.
Rootly Roundtable: From Weak Signals to Confident Fixes
Invite-only roundtable on filtering alert noise, enriching alerts with automated ownership context, and streamlining bug validation workflows.
David Owczarek — Former Engineering Director
Lessons from leading engineering teams through organizational change and system migrations.
Ryan Lockard — VP of Platform Engineering at CVS Health
Platform engineering at healthcare scale — compliance, reliability, and developer experience.
Cosmo Wolfe — Head of Technology at Metronome
Building reliable billing infrastructure and the unique challenges of usage-based pricing systems.
Chase Roberts — COO at Northflank
Operating a developer platform company and the reliability challenges of running infrastructure for others.
Rootly Roundtable: Centralized vs. Distributed Incident Teams
Invite-only roundtable comparing centralized specialist responders with distributed on-call models — which strategy works for your org.
Justin Reock — Deputy CTO at DX
Developer experience, engineering metrics, and what makes teams productive and reliable.
The Future of AI-Driven Reliability
Panel with leaders from a16z, Y Combinator, and Google Cloud on how MCP servers and agent-to-agent communication are revolutionizing developer tools. Demos from Anthropic, Sentry, Postman, and Browserbase.
Incident Vibing: The Self-Healing System
Tracing the arc from ingesting logs at LinkedIn/SlideShare to LLM-driven RCA today. How fine-tuning, MCP, and incident vibing are reshaping SRE.
Marino Wijay — Staff Solutions Architect at Kong
API gateway architecture, service mesh patterns, and reliability at the network edge.
Mark Quigley — Head of Platform Engineering
Building internal developer platforms and the organizational dynamics of platform teams.
Navigating DORA: A Guide to the EU's Digital Operational Resilience Act
Breaking down the EU's Digital Operational Resilience Act — what it means for financial services and tech companies operating in Europe.
Vibe Coding Is Here — But Are You Ready for Incident Vibing?
If developers are vibe coding, SREs are now incident vibing. What happens when AI-generated code meets production reality.
Dan Slimmon — Incident Response Trainer
Training teams for incident response — chaos engineering, game days, and building organizational muscle memory.
Rootly Roundtable: The State of AI in Incident Management
Invite-only roundtable examining AI's role in incident response — separating practical applications from hype with industry leaders.
Mariano Cocirio — Staff Software Engineer at Vercel
Frontend infrastructure reliability, edge computing, and the operational side of Vercel's platform.
Adriana Villela — Principal DevRel at Dynatrace
Observability culture, OpenTelemetry adoption, and the human side of monitoring and debugging.
Sean Goedecke — Staff Software Engineer at GitHub
Reliability at GitHub scale — incident management, deployment practices, and engineering culture.
SRE-skills-bench
A benchmark suite for evaluating AI agents on real-world SRE tasks — incident diagnosis, runbook execution, and infrastructure troubleshooting.
On-Call-Health
Detects early warning signs of on-call engineer burnout by pulling data from Rootly, PagerDuty, Linear, GitHub, and Slack.
AI-Driven Incident Resolution — Hype or Reality
A grounded look at what LLMs can and can't do in an incident response workflow today, drawn from real production experiments at Rootly AI Labs.
Automated Management of a Distributed Computing System
US Patent (US9674031B2) — A machine learning-powered self-healing infrastructure that monitors distributed systems, detects anomalies, and automatically applies remedies based on historical incident data. Co-designed at LinkedIn.
First 5 Commands When I Connect on a Linux Server
The essential first five commands every sysadmin runs when SSH-ing into a Linux server — from uptime to disk usage.
How DevOps Failed 60K Users
A postmortem on a DevOps failure that affected 60,000 users — lessons in monitoring gaps, deployment practices, and incident response.
DevOps Students Learn the Value of Uptime With 3 a.m. Calls
Teaching DevOps through real-world on-call rotations — how students learn the value of monitoring, alerting, and system reliability the hard way.