
The State of AI Observability
Q2 2026
The concept. AI systems break in production every few weeks and somebody writes it up in public. Collect those write-ups, sort them by how the thing actually failed, and put one question to each: what would have caught this earlier, and does that exist yet. Sometimes it does and nobody had it switched on. Sometimes it does not, and that is where the gap actually sits.
How it was built. Datadog’s documentation and conference talks came down, with incident retrospectives, cost analyses and the OpenTelemetry and MCP specs. Every claim traces to a source and carries a date where that matters. What holds regardless of release is kept apart from what expires with the next one, so you can see which half will rot first.
What’s inside. Why the three-pillar model cannot see AI workloads, and why that is structural rather than a setting you forgot. What tracing does when the system answers differently every time. Where operational monitoring ends and evaluation monitoring begins, which teams keep confusing. What it costs and what drives the bill. Documented failures sorted by category. And what happens when the monitoring is itself a model.
It assumes you know what a span is. It will not sell you a vendor, and it tells you which of its own claims will expire first.