Back to all work

Topic

SRE

My writing, talks, podcasts, and projects about SRE.

Selected work

Thumbnail for Eduardo Ordax — Principal GTM GenAI at AWS
Podcast
Humans of Reliability
Humans of Reliability

Eduardo Ordax — Principal GTM GenAI at AWS

The reality of GenAI in production — why organizational culture is the biggest blocker, scaling non-deterministic LLM systems, and what separates AI winners from experimenters.

Thumbnail for LLMs Broke the SRE Runbook. Now What?
Article
The New Stack
The New Stack

LLMs Broke the SRE Runbook. Now What?

AI-generated code is outpacing traditional runbooks. How SRE teams are adapting their incident response playbooks for the LLM era.

Thumbnail for From Vibes to Outages: Riding the AI Code Wave
Talk
SREcon EMEA 2025
SREcon EMEA 2025

From Vibes to Outages: Riding the AI Code Wave

AI-assisted coding is exploding — but acceleration doesn't mean reliability. Real examples of hard-to-trace LLM bugs, hallucinated dependencies, and operational fallout for lean SRE teams.

Thumbnail for Will LLMs and Vibe Coding Fuel a Developer Renaissance?
Article
The New Stack
The New Stack

Will LLMs and Vibe Coding Fuel a Developer Renaissance?

Exploring whether AI-assisted coding tools will democratize software development or create new categories of hard-to-debug production issues.

Thumbnail for Is AI-assisted coding an incident magnet?
Article
LeadDev
LeadDev

Is AI-assisted coding an incident magnet?

AI-generated code ships faster but introduces subtle bugs that are harder to trace. What engineering leaders need to know about the reliability trade-offs.

Thumbnail for AI Fails Like We Do
Podcast
Engineering Futures
Engineering Futures

AI Fails Like We Do

Sylvain Kalache went looking for the new ways AI breaks software. Every one turned out to be a habit we've had for years.

Thumbnail for More Code, More Incidents? Staying Reliable When AI Writes the Code
Talk
Incident Fest
Incident Fest 2026

More Code, More Incidents? Staying Reliable When AI Writes the Code

Speaking on the Incident Fest Main Stage about staying reliable as AI-assisted development increases the rate at which software is created.

Thumbnail for Eran Kampf — VP of Engineering at Twingate
Podcast
Humans of Reliability
Humans of Reliability

Eran Kampf — VP of Engineering at Twingate

Why Twingate stopped shipping features to rebuild reliability — from active-active multi-region architecture and smaller blast radiuses to preserving human ownership as agentic coding accelerates delivery.

Thumbnail for AIOps Summit
Talk
AIOps Summit
AIOps Summit 2026

AIOps Summit

Meta-hosted summit in Menlo Park focused on applying AI, LLMs, and agents to software incidents and response.

Thumbnail for Alexey Grigorev — Founder of DataTalks.Club
Podcast
Humans of Reliability
Humans of Reliability

Alexey Grigorev — Founder of DataTalks.Club

How a chain of reasonable-sounding decisions led an AI coding agent to run terraform destroy against a live production database — and the guardrails that separate moving fast from losing everything.

Thumbnail for AI Made Developers 25% More Productive. It Also Tripled Our Incident Rate.
Talk
AI DevSummit
AI DevSummit New York 2026

AI Made Developers 25% More Productive. It Also Tripled Our Incident Rate.

At AI DevSummit New York: how AI-assisted coding is reshaping software delivery — the productivity gains, the surge in incident rates, and what reliability teams need to do about it.

Thumbnail for Hamed Silatani — Co-Founder and CEO at Uptime Labs
Podcast
Humans of Reliability
Humans of Reliability

Hamed Silatani — Co-Founder and CEO at Uptime Labs

Every pilot trains for engine failure, but most engineers face their first major incident unprepared — why communication, coordination, and decision-making need deliberate simulation practice.

Thumbnail for The Science of On-Call Burnout: Why "How Are You Doing?" Always Gets "Fine"
Talk
DevOpsDays Austin
DevOpsDays Austin 2026

The Science of On-Call Burnout: Why "How Are You Doing?" Always Gets "Fine"

Why traditional burnout detection fails in engineering teams. Connecting Maslach's burnout research and the Copenhagen Burnout Inventory to observable on-call metrics, and demonstrating that observed data outperforms subjective self-reporting.

Thumbnail for Gandhi Kumar — Principal Incident Commander at Twilio
Podcast
Humans of Reliability
Humans of Reliability

Gandhi Kumar — Principal Incident Commander at Twilio

The golden hour — why the first 15 minutes of an incident decide everything, from mitigation strategy to customer trust, and why human psychology matters more than tools.

Thumbnail for Cliff Snyder — Senior SRE at Multimedia LLC
Podcast
Humans of Reliability
Humans of Reliability

Cliff Snyder — Senior SRE at Multimedia LLC

From 600 to 6,000 — an 18-month transformation replacing legacy tooling with a purpose-built incident management platform, scaling operational responsibility from 600 SREs to 6,000 engineers across the org.

Thumbnail for Ganesh Datta — Co-Founder & CTO at Cortex
Podcast
Humans of Reliability
Humans of Reliability

Ganesh Datta — Co-Founder & CTO at Cortex

AI didn't change the game, it just exposed your bottlenecks — how platform engineering and SRE teams solve identical human problems through influence rather than authority, and why AI amplifies existing bottlenecks instead of transforming operations.

Thumbnail for Reliability Rebels Podcast — Guest Appearance
Podcast
Reliability Rebels
Reliability Rebels

Reliability Rebels Podcast — Guest Appearance

Reliability Rebels Guest appearance to discuss reliability engineering, AI, and modern incident response.

Thumbnail for Dana Lawson — CTO at Netlify
Podcast
Humans of Reliability
Humans of Reliability

Dana Lawson — CTO at Netlify

Fear, identity, and flaky tests — why SRE resistance to AI agents stems from identity and control concerns rather than the technology itself, and practical strategies for adopting AI-driven reliability tools starting with low-risk tasks.

Thumbnail for The Software Leaders Uncensored Podcast
Podcast
Software Leaders Uncensored
Software Leaders Uncensored

The Software Leaders Uncensored Podcast

A candid conversation on software leadership, reliability engineering, and the evolving role of AI in modern engineering teams.

Thumbnail for How AI is Changing Incident Response (SRE Deep Dive)
Podcast
The Rollback Café
The Rollback Café

How AI is Changing Incident Response (SRE Deep Dive)

From faster detection to smarter debugging — where AI delivers practical value in incident response, its limitations in production environments, and how teams can integrate AI while maintaining reliability standards.

Thumbnail for From Vibes to Outages: When AI Writes the Code You Debug
Talk
Conf42 Site Reliability Engineering 2026

From Vibes to Outages: When AI Writes the Code You Debug

LLM-assisted development introduces hard-to-trace bugs, AI-generated tests that mirror flawed logic, and hallucinated dependencies. Exploring operational consequences and actionable strategies from AI-powered incident tools to incident vibing.

Thumbnail for Will Wilson — CEO at Antithesis
Podcast
Humans of Reliability
Humans of Reliability

Will Wilson — CEO at Antithesis

The incident you never had — deterministic simulation testing, why conventional testing misses bugs that cause real outages, and how simulation-based approaches improve software reliability.

Thumbnail for Google SRE NYC Tech Talk — On-Call Health
Talk
Google SRE NYC
Google SRE NYC Tech Talks

Google SRE NYC Tech Talk — On-Call Health

The science behind responder overload and introducing On-Call Health, an open-source tool for detecting incident responder fatigue. At Google NYC Pier 57.

Thumbnail for Stephen Townshend — SRE Team Lead at isDown
Podcast
Humans of Reliability
Humans of Reliability

Stephen Townshend — SRE Team Lead at isDown

Burnout doesn't ask permission — the physiology of burnout from an incident responder's perspective, warning signs, recovery, and the systemic changes the industry must make.

Thumbnail for Swizec Teller — Bestselling Author
Podcast
Humans of Reliability
Humans of Reliability

Swizec Teller — Bestselling Author

Code is cheap, reliability isn't — owning production in the AI era, the hidden complexity of SRE work, and why human ownership remains essential.

Thumbnail for Dileshni Jayasinghe — VP of Technology at commonsku
Podcast
Humans of Reliability
Humans of Reliability

Dileshni Jayasinghe — VP of Technology at commonsku

Democratizing reliability — empowering non-engineers with operational power, incident management as a muscle, and AI-powered postmortems.

Thumbnail for Tomás Hernando Koffman — Co-founder at Not Diamond
Podcast
Humans of Reliability
Humans of Reliability

Tomás Hernando Koffman — Co-founder at Not Diamond

99%+ accuracy on a moving target — model deprecation, reliability with LLMs, and treating prompts as architectural components.

Thumbnail for Julien Simon — VP and Chief Evangelist
Podcast
Humans of Reliability
Humans of Reliability

Julien Simon — VP and Chief Evangelist

A conversation about developer advocacy, AI evangelism, and building reliable systems at scale.

Thumbnail for Conor Bronsdon — Head of Content at Galileo
Podcast
Humans of Reliability
Humans of Reliability

Conor Bronsdon — Head of Content at Galileo

Developer awareness, content strategy, and the human side of reliability engineering.

Thumbnail for Tea, Pipelines, and Retries: A Practical Guide to MLOps
Panel Moderation
SREcon EMEA 2025
SREcon EMEA 2025

Tea, Pipelines, and Retries: A Practical Guide to MLOps

Discussion at SREcon EMEA 2025 on how AI is transforming the software development lifecycle — CI/CD pipelines, deployments, scaling, monitoring, incident management, reliability tooling, and emerging disciplines like LLMOps.

Thumbnail for AI Meets Reliability
Panel Moderation
Rootly AI Labs
Rootly AI Labs

AI Meets Reliability

Panel exploring AI-driven automation and observability with leaders from NVIDIA, OpenAI, Baseten, Replit, and Weights & Biases on scaling operations and reducing MTTR.

Thumbnail for Rob Zuber — CTO at CircleCI
Podcast
Humans of Reliability
Humans of Reliability

Rob Zuber — CTO at CircleCI

The end of good code, AI throughput, and what reliability means at CI/CD scale.

Thumbnail for MCPs and the Next Wave of Reliability w/ Rootly AI, WorkOS, Block, Microsoft & Groq
Panel Moderation
Rootly AI Labs
Rootly AI Labs

MCPs and the Next Wave of Reliability w/ Rootly AI, WorkOS, Block, Microsoft & Groq

Panel at the AWS GenAI Loft in San Francisco on MCPs, incident automation, observability, and generative AI tooling for reliability.

Thumbnail for Shery Brauner — SVP of Razor Group
Podcast
Humans of Reliability
Humans of Reliability

Shery Brauner — SVP of Razor Group

Leading engineering at scale in e-commerce aggregation and the operational challenges that come with it.

Thumbnail for Frontiers of AI: Building with Rootly AI, Zscaler, CircleCI, Fireworks AI & Google DeepMind
Panel Moderation
Rootly AI Labs
Rootly AI Labs

Frontiers of AI: Building with Rootly AI, Zscaler, CircleCI, Fireworks AI & Google DeepMind

Panel at Google HQ with 300+ attendees exploring real-world Gemini models, reinforcement learning, next-gen agent systems, and AI reliability.

Thumbnail for Brian Shaw — SVP of Infrastructure and Core Banking
Podcast
Humans of Reliability
Humans of Reliability

Brian Shaw — SVP of Infrastructure and Core Banking

Building and operating core banking infrastructure with a focus on uptime and regulatory compliance.

Thumbnail for Exploring AI's Role in Incident Management
Podcast
Techstrong TV
Techstrong TV

Exploring AI's Role in Incident Management

Interview with Alan Shimel at PlatformCon NYC 2025 on how AI applies to incident management and reliability engineering — triage, root cause analysis, and why AI enhances rather than replaces engineers.

Thumbnail for Rootly Roundtable: From Weak Signals to Confident Fixes
Panel Moderation
Rootly
Rootly

Rootly Roundtable: From Weak Signals to Confident Fixes

Invite-only roundtable on filtering alert noise, enriching alerts with automated ownership context, and streamlining bug validation workflows.

Thumbnail for David Owczarek — Former Engineering Director
Podcast
Humans of Reliability
Humans of Reliability

David Owczarek — Former Engineering Director

Lessons from leading engineering teams through organizational change and system migrations.

Thumbnail for Ryan Lockard — VP of Platform Engineering at CVS Health
Podcast
Humans of Reliability
Humans of Reliability

Ryan Lockard — VP of Platform Engineering at CVS Health

Platform engineering at healthcare scale — compliance, reliability, and developer experience.

Thumbnail for Cosmo Wolfe — Head of Technology at Metronome
Podcast
Humans of Reliability
Humans of Reliability

Cosmo Wolfe — Head of Technology at Metronome

Building reliable billing infrastructure and the unique challenges of usage-based pricing systems.

Thumbnail for Chase Roberts — COO at Northflank
Podcast
Humans of Reliability
Humans of Reliability

Chase Roberts — COO at Northflank

Operating a developer platform company and the reliability challenges of running infrastructure for others.

Thumbnail for Rootly Roundtable: Centralized vs. Distributed Incident Teams
Panel Moderation
Rootly
Rootly

Rootly Roundtable: Centralized vs. Distributed Incident Teams

Invite-only roundtable comparing centralized specialist responders with distributed on-call models — which strategy works for your org.

Thumbnail for Justin Reock — Deputy CTO at DX
Podcast
Humans of Reliability
Humans of Reliability

Justin Reock — Deputy CTO at DX

Developer experience, engineering metrics, and what makes teams productive and reliable.

Thumbnail for The Future of AI-Driven Reliability
Panel Moderation
Rootly AI Labs
Rootly AI Labs

The Future of AI-Driven Reliability

Panel with leaders from a16z, Y Combinator, and Google Cloud on how MCP servers and agent-to-agent communication are revolutionizing developer tools. Demos from Anthropic, Sentry, Postman, and Browserbase.

Thumbnail for Incident Vibing: The Self-Healing System
Podcast
Adventures in DevOps
Adventures in DevOps

Incident Vibing: The Self-Healing System

Tracing the arc from ingesting logs at LinkedIn/SlideShare to LLM-driven RCA today. How fine-tuning, MCP, and incident vibing are reshaping SRE.

Thumbnail for Marino Wijay — Staff Solutions Architect at Kong
Podcast
Humans of Reliability
Humans of Reliability

Marino Wijay — Staff Solutions Architect at Kong

API gateway architecture, service mesh patterns, and reliability at the network edge.

Thumbnail for Mark Quigley — Head of Platform Engineering
Podcast
Humans of Reliability
Humans of Reliability

Mark Quigley — Head of Platform Engineering

Building internal developer platforms and the organizational dynamics of platform teams.

Thumbnail for Navigating DORA: A Guide to the EU's Digital Operational Resilience Act
Podcast
Data Defenders Forum
Data Defenders Forum

Navigating DORA: A Guide to the EU's Digital Operational Resilience Act

Breaking down the EU's Digital Operational Resilience Act — what it means for financial services and tech companies operating in Europe.

Thumbnail for Vibe Coding Is Here — But Are You Ready for Incident Vibing?
Article
The New Stack
The New Stack

Vibe Coding Is Here — But Are You Ready for Incident Vibing?

If developers are vibe coding, SREs are now incident vibing. What happens when AI-generated code meets production reality.

Thumbnail for Dan Slimmon — Incident Response Trainer
Podcast
Humans of Reliability
Humans of Reliability

Dan Slimmon — Incident Response Trainer

Training teams for incident response — chaos engineering, game days, and building organizational muscle memory.

Thumbnail for Rootly Roundtable: The State of AI in Incident Management
Panel Moderation
Rootly
Rootly

Rootly Roundtable: The State of AI in Incident Management

Invite-only roundtable examining AI's role in incident response — separating practical applications from hype with industry leaders.

Thumbnail for Mariano Cocirio — Staff Software Engineer at Vercel
Podcast
Humans of Reliability
Humans of Reliability

Mariano Cocirio — Staff Software Engineer at Vercel

Frontend infrastructure reliability, edge computing, and the operational side of Vercel's platform.

Thumbnail for Adriana Villela — Principal DevRel at Dynatrace
Podcast
Humans of Reliability
Humans of Reliability

Adriana Villela — Principal DevRel at Dynatrace

Observability culture, OpenTelemetry adoption, and the human side of monitoring and debugging.

Thumbnail for Sean Goedecke — Staff Software Engineer at GitHub
Podcast
Humans of Reliability
Humans of Reliability

Sean Goedecke — Staff Software Engineer at GitHub

Reliability at GitHub scale — incident management, deployment practices, and engineering culture.

Thumbnail for SRE-skills-bench
Project
GitHub
GitHub

SRE-skills-bench

A benchmark suite for evaluating AI agents on real-world SRE tasks — incident diagnosis, runbook execution, and infrastructure troubleshooting.

12
Thumbnail for On-Call-Health
Project
GitHub
GitHub

On-Call-Health

Detects early warning signs of on-call engineer burnout by pulling data from Rootly, PagerDuty, Linear, GitHub, and Slack.

32
Thumbnail for AI-Driven Incident Resolution — Hype or Reality
Talk
DevOpsDays Austin
DevOpsDays Austin

AI-Driven Incident Resolution — Hype or Reality

A grounded look at what LLMs can and can't do in an incident response workflow today, drawn from real production experiments at Rootly AI Labs.

Thumbnail for Automated Management of a Distributed Computing System
Project
US Patent Office

Automated Management of a Distributed Computing System

US Patent (US9674031B2) — A machine learning-powered self-healing infrastructure that monitors distributed systems, detects anomalies, and automatically applies remedies based on historical incident data. Co-designed at LinkedIn.

Thumbnail for First 5 Commands When I Connect on a Linux Server
Article
Linux.com
Linux.com

First 5 Commands When I Connect on a Linux Server

The essential first five commands every sysadmin runs when SSH-ing into a Linux server — from uptime to disk usage.

Thumbnail for How DevOps Failed 60K Users
Article
Linux.com
Linux.com

How DevOps Failed 60K Users

A postmortem on a DevOps failure that affected 60,000 users — lessons in monitoring gaps, deployment practices, and incident response.

Thumbnail for DevOps Students Learn the Value of Uptime With 3 a.m. Calls
Article
Linux.com
Linux.com

DevOps Students Learn the Value of Uptime With 3 a.m. Calls

Teaching DevOps through real-world on-call rotations — how students learn the value of monitoring, alerting, and system reliability the hard way.