Back to all work

Topic

AI Agents

My writing, talks, podcasts, and projects about AI Agents.

Selected work

Thumbnail for Eduardo Ordax — Principal GTM GenAI at AWS
Podcast
Humans of Reliability
Humans of Reliability

Eduardo Ordax — Principal GTM GenAI at AWS

The reality of GenAI in production — why organizational culture is the biggest blocker, scaling non-deterministic LLM systems, and what separates AI winners from experimenters.

Thumbnail for LLMs Broke the SRE Runbook. Now What?
Article
The New Stack
The New Stack

LLMs Broke the SRE Runbook. Now What?

AI-generated code is outpacing traditional runbooks. How SRE teams are adapting their incident response playbooks for the LLM era.

Thumbnail for From Vibes to Outages: Riding the AI Code Wave
Talk
SREcon EMEA 2025
SREcon EMEA 2025

From Vibes to Outages: Riding the AI Code Wave

AI-assisted coding is exploding — but acceleration doesn't mean reliability. Real examples of hard-to-trace LLM bugs, hallucinated dependencies, and operational fallout for lean SRE teams.

Thumbnail for Will LLMs and Vibe Coding Fuel a Developer Renaissance?
Article
The New Stack
The New Stack

Will LLMs and Vibe Coding Fuel a Developer Renaissance?

Exploring whether AI-assisted coding tools will democratize software development or create new categories of hard-to-debug production issues.

Thumbnail for Why Are Agent Protocols Like MCP and A2A Needed?
Article
The New Stack
The New Stack

Why Are Agent Protocols Like MCP and A2A Needed?

Breaking down the Model Context Protocol (MCP) and Agent-to-Agent (A2A) standards — why interoperability matters for the next wave of AI agents.

Thumbnail for Is AI-assisted coding an incident magnet?
Article
LeadDev
LeadDev

Is AI-assisted coding an incident magnet?

AI-generated code ships faster but introduces subtle bugs that are harder to trace. What engineering leaders need to know about the reliability trade-offs.

Thumbnail for Anthony Alcaraz — GTM Agentic Engineering Lead at AWS
Podcast
Humans of Reliability
Humans of Reliability

Anthony Alcaraz — GTM Agentic Engineering Lead at AWS

Your AI agents are lost — why context engineering and knowledge graphs give agents the structure they need for retrieval, memory, planning, reasoning, and learning from business outcomes.

Thumbnail for Nir Soudry — Head of R&D at 7AI
Podcast
Humans of Reliability
Humans of Reliability

Nir Soudry — Head of R&D at 7AI

AI vs. AI — how autonomous attackers are scaling phishing and malware, how defenders can use agents to cut alert noise, and where accountability belongs when AI investigates and remediates incidents.

Thumbnail for Eran Kampf — VP of Engineering at Twingate
Podcast
Humans of Reliability
Humans of Reliability

Eran Kampf — VP of Engineering at Twingate

Why Twingate stopped shipping features to rebuild reliability — from active-active multi-region architecture and smaller blast radiuses to preserving human ownership as agentic coding accelerates delivery.

Thumbnail for AIOps Summit
Talk
AIOps Summit
AIOps Summit 2026

AIOps Summit

Meta-hosted summit in Menlo Park focused on applying AI, LLMs, and agents to software incidents and response.

Thumbnail for Alexey Grigorev — Founder of DataTalks.Club
Podcast
Humans of Reliability
Humans of Reliability

Alexey Grigorev — Founder of DataTalks.Club

How a chain of reasonable-sounding decisions led an AI coding agent to run terraform destroy against a live production database — and the guardrails that separate moving fast from losing everything.

Thumbnail for AI Made Developers 25% More Productive. It Also Tripled Our Incident Rate.
Talk
AI DevSummit
AI DevSummit New York 2026

AI Made Developers 25% More Productive. It Also Tripled Our Incident Rate.

At AI DevSummit New York: how AI-assisted coding is reshaping software delivery — the productivity gains, the surge in incident rates, and what reliability teams need to do about it.

Thumbnail for Maria Vechtomova — Co-founder at Cauchy
Podcast
Humans of Reliability
Humans of Reliability

Maria Vechtomova — Co-founder at Cauchy

LLM observability — how monitoring principles from MLOps apply to large language model applications, the gaps that went overlooked for nine years, and best practices for observing AI systems.

Thumbnail for Ganesh Datta — Co-Founder & CTO at Cortex
Podcast
Humans of Reliability
Humans of Reliability

Ganesh Datta — Co-Founder & CTO at Cortex

AI didn't change the game, it just exposed your bottlenecks — how platform engineering and SRE teams solve identical human problems through influence rather than authority, and why AI amplifies existing bottlenecks instead of transforming operations.

Thumbnail for Reliability Rebels Podcast — Guest Appearance
Podcast
Reliability Rebels
Reliability Rebels

Reliability Rebels Podcast — Guest Appearance

Reliability Rebels Guest appearance to discuss reliability engineering, AI, and modern incident response.

Thumbnail for Dana Lawson — CTO at Netlify
Podcast
Humans of Reliability
Humans of Reliability

Dana Lawson — CTO at Netlify

Fear, identity, and flaky tests — why SRE resistance to AI agents stems from identity and control concerns rather than the technology itself, and practical strategies for adopting AI-driven reliability tools starting with low-risk tasks.

Thumbnail for Will Wilson — CEO at Antithesis
Podcast
Humans of Reliability
Humans of Reliability

Will Wilson — CEO at Antithesis

The incident you never had — deterministic simulation testing, why conventional testing misses bugs that cause real outages, and how simulation-based approaches improve software reliability.

Thumbnail for Swizec Teller — Bestselling Author
Podcast
Humans of Reliability
Humans of Reliability

Swizec Teller — Bestselling Author

Code is cheap, reliability isn't — owning production in the AI era, the hidden complexity of SRE work, and why human ownership remains essential.

Thumbnail for OpenClaw Demo Night w/ Rootly AI, Convex, Sentry & DigitalOcean
Talk
OpenClaw Demo Night

OpenClaw Demo Night w/ Rootly AI, Convex, Sentry & DigitalOcean

An evening of AI demos and networking in Toronto — presenting Rootly AI alongside teams from Sentry, Red Brick Labs, Convex, and DigitalOcean.

Thumbnail for Dileshni Jayasinghe — VP of Technology at commonsku
Podcast
Humans of Reliability
Humans of Reliability

Dileshni Jayasinghe — VP of Technology at commonsku

Democratizing reliability — empowering non-engineers with operational power, incident management as a muscle, and AI-powered postmortems.

Thumbnail for Tomás Hernando Koffman — Co-founder at Not Diamond
Podcast
Humans of Reliability
Humans of Reliability

Tomás Hernando Koffman — Co-founder at Not Diamond

99%+ accuracy on a moving target — model deprecation, reliability with LLMs, and treating prompts as architectural components.

Thumbnail for Developers lose focus 1,200 times a day — how MCP could change that
Article
VentureBeat
VentureBeat

Developers lose focus 1,200 times a day — how MCP could change that

Context switching kills developer productivity. How the Model Context Protocol (MCP) can reduce tool fragmentation and keep engineers in flow.

Thumbnail for Julien Simon — VP and Chief Evangelist
Podcast
Humans of Reliability
Humans of Reliability

Julien Simon — VP and Chief Evangelist

A conversation about developer advocacy, AI evangelism, and building reliable systems at scale.

Thumbnail for Tea, Pipelines, and Retries: A Practical Guide to MLOps
Panel Moderation
SREcon EMEA 2025
SREcon EMEA 2025

Tea, Pipelines, and Retries: A Practical Guide to MLOps

Discussion at SREcon EMEA 2025 on how AI is transforming the software development lifecycle — CI/CD pipelines, deployments, scaling, monitoring, incident management, reliability tooling, and emerging disciplines like LLMOps.

Thumbnail for AI Meets Reliability
Panel Moderation
Rootly AI Labs
Rootly AI Labs

AI Meets Reliability

Panel exploring AI-driven automation and observability with leaders from NVIDIA, OpenAI, Baseten, Replit, and Weights & Biases on scaling operations and reducing MTTR.

Thumbnail for Rob Zuber — CTO at CircleCI
Podcast
Humans of Reliability
Humans of Reliability

Rob Zuber — CTO at CircleCI

The end of good code, AI throughput, and what reliability means at CI/CD scale.

Thumbnail for MCPs and the Next Wave of Reliability w/ Rootly AI, WorkOS, Block, Microsoft & Groq
Panel Moderation
Rootly AI Labs
Rootly AI Labs

MCPs and the Next Wave of Reliability w/ Rootly AI, WorkOS, Block, Microsoft & Groq

Panel at the AWS GenAI Loft in San Francisco on MCPs, incident automation, observability, and generative AI tooling for reliability.

Thumbnail for AI-First Platform Engineering: 3 Signals From PlatformCon
Article
The New Stack
The New Stack

AI-First Platform Engineering: 3 Signals From PlatformCon

Three emerging patterns from PlatformCon that signal how AI is reshaping internal developer platforms and platform team workflows.

Thumbnail for Frontiers of AI: Building with Rootly AI, Zscaler, CircleCI, Fireworks AI & Google DeepMind
Panel Moderation
Rootly AI Labs
Rootly AI Labs

Frontiers of AI: Building with Rootly AI, Zscaler, CircleCI, Fireworks AI & Google DeepMind

Panel at Google HQ with 300+ attendees exploring real-world Gemini models, reinforcement learning, next-gen agent systems, and AI reliability.

Thumbnail for AI Security Demo Night w/ Rootly AI, Okta, Panther, Tailscale & More
Talk
AI Security Demo Night

AI Security Demo Night w/ Rootly AI, Okta, Panther, Tailscale & More

Live demos at Okta HQ in San Francisco — presenting Rootly AI alongside cybersecurity startups tackling identity, endpoint security, and threat response.

Thumbnail for Exploring AI's Role in Incident Management
Podcast
Techstrong TV
Techstrong TV

Exploring AI's Role in Incident Management

Interview with Alan Shimel at PlatformCon NYC 2025 on how AI applies to incident management and reliability engineering — triage, root cause analysis, and why AI enhances rather than replaces engineers.

Thumbnail for How AI is Fueling the Developer Renaissance
Panel Moderation
Rootly AI Labs
Rootly AI Labs

How AI is Fueling the Developer Renaissance

Panel at the AWS GenAI Loft with leaders from a16z, AWS, Dagger, Braintrust, Baseten, and Arize AI on how AI is transforming developer workflows.

Thumbnail for The Future of AI-Driven Reliability
Panel Moderation
Rootly AI Labs
Rootly AI Labs

The Future of AI-Driven Reliability

Panel with leaders from a16z, Y Combinator, and Google Cloud on how MCP servers and agent-to-agent communication are revolutionizing developer tools. Demos from Anthropic, Sentry, Postman, and Browserbase.

Thumbnail for Incident Vibing: The Self-Healing System
Podcast
Adventures in DevOps
Adventures in DevOps

Incident Vibing: The Self-Healing System

Tracing the arc from ingesting logs at LinkedIn/SlideShare to LLM-driven RCA today. How fine-tuning, MCP, and incident vibing are reshaping SRE.

Thumbnail for Vibe Coding Is Here — But Are You Ready for Incident Vibing?
Article
The New Stack
The New Stack

Vibe Coding Is Here — But Are You Ready for Incident Vibing?

If developers are vibe coding, SREs are now incident vibing. What happens when AI-generated code meets production reality.

Thumbnail for Rootly Roundtable: The State of AI in Incident Management
Panel Moderation
Rootly
Rootly

Rootly Roundtable: The State of AI in Incident Management

Invite-only roundtable examining AI's role in incident response — separating practical applications from hype with industry leaders.

Thumbnail for Rootly-MCP-server
Project
GitHub
GitHub

Rootly-MCP-server

An MCP server for Rootly — enabling AI agents to interact with incident management workflows via the Model Context Protocol.

39
Thumbnail for SRE-skills-bench
Project
GitHub
GitHub

SRE-skills-bench

A benchmark suite for evaluating AI agents on real-world SRE tasks — incident diagnosis, runbook execution, and infrastructure troubleshooting.

12
Thumbnail for MCP-Sylvain-Kalache
Project
GitHub
GitHub

MCP-Sylvain-Kalache

A personal MCP server that exposes my portfolio data — articles, podcasts, talks, panels, GitHub projects, news, bio, timeline, and live weather — as queryable tools for AI agents.

0
Thumbnail for Introducing Canyon — an AI for developer self-service
Talk
Platform Engineering

Introducing Canyon — an AI for developer self-service

Canyon lets developers provision resources, troubleshoot issues, and interact with internal tooling via natural-language prompts.

Thumbnail for From hype to impact: How AI is reshaping platform engineering
Talk
PlatformCon
PlatformCon

From hype to impact: How AI is reshaping platform engineering

Moving past the buzzwords — where AI is actually changing developer workflows, internal developer portals, and what it means for platform teams.

Thumbnail for EU AI Act Secrets Revealed
Podcast
Data Defenders Forum
Data Defenders Forum

EU AI Act Secrets Revealed

What the EU AI Act actually requires — risk classifications, compliance timelines, and what it means for AI-driven products.

Thumbnail for Intelligent Document Processing Compliance, from Stone Tablets to Digital Docs
Podcast
Data Defenders Forum
Data Defenders Forum

Intelligent Document Processing Compliance, from Stone Tablets to Digital Docs

The evolution of document processing compliance — from physical records to AI-powered intelligent document processing systems.

Thumbnail for AI-Driven Incident Resolution — Hype or Reality
Talk
DevOpsDays Austin
DevOpsDays Austin

AI-Driven Incident Resolution — Hype or Reality

A grounded look at what LLMs can and can't do in an incident response workflow today, drawn from real production experiments at Rootly AI Labs.

Thumbnail for Navigating LLMs challenges in data security & compliance
Podcast
Data Defenders Forum
Data Defenders Forum

Navigating LLMs challenges in data security & compliance

The unique data security and compliance challenges that large language models introduce — from training data to inference outputs.

Thumbnail for For companies that use ML, labeled data is the key differentiator
Article
TechCrunch
TechCrunch

For companies that use ML, labeled data is the key differentiator

Why Tesla leads on ADAS — and what it teaches every company about the strategic value of training data, annotation pipelines, and the $6B labeling market.