Observability for AI Agents: What to Log, Trace, and Alert On
Beyond token counts and latency dashboards — a practical framework for what to log, trace, and alert on in production AI agent systems, covering run IDs, tool-call latency, hallucination detection, cost-per-run tracking, and human-review escalation.