Last Tuesday a client asked me why their agent-debugging Slack channel had three different dashboards pinned — one for the Bedrock agent, one for the LangGraph service running on ECS, and a Grafana board someone built by hand for the Strands prototype. Nobody could tell me, in under five minutes, which trace belonged to which failed tool call. That’s the exact problem AWS says it solved on September 23 with CloudWatch Omni.
What actually shipped
CloudWatch Omni is built on OpenTelemetry instead of AWS’s old proprietary instrumentation, which is the detail that matters most. It ingests traces across accounts and, notably, across clouds — you’re not locked into an all-AWS stack to get a single pane of glass. Launch regions are us-east-1, us-west-2, and eu-west-1, so if your workloads sit in ap-southeast-1 you’re waiting.
Two features are worth the upgrade on their own:
- Natural-language querying via a DevOps Agent — you type “why did the checkout agent retry four times at 14:02” instead of writing a CloudWatch Logs Insights query by hand.
- Native eval-driven observability for LangGraph, CrewAI, the OpenAI Agents SDK, and AWS’s own Strands framework. This is the part that’s genuinely new: instead of bolting evals on as a separate pipeline, Omni treats an eval score as a first-class span attribute next to latency and token count.
There are also IDE extensions for VS Code, Cursor, and Kiro, so you can jump from a failed trace straight to the line of agent code that triggered it.
Why OpenTelemetry-first is the real story
Every observability vendor claims OTel support now — Cloudflare, Datadog, Honeycomb, all of them. What’s different here is that AWS is the vendor who spent a decade pushing you toward CloudWatch’s proprietary metric format, and they just conceded that agents don’t respect account boundaries, let alone cloud boundaries. A single user request can hop from a Bedrock agent, to a Lambda tool call, to an external API hosted on GCP, to a Snowflake query. If your observability tool assumes everything lives inside one AWS account, you’re already blind to half your production traffic.
Hands-on: wiring a Strands agent into Omni
Here’s the minimum viable setup, assuming you already have an AWS Strands agent running:
from strands import Agent
from strands.telemetry import OmniExporter
agent = Agent(
name="checkout-agent",
telemetry=OmniExporter(
eval_hooks=["tool_call_success", "hallucination_check"],
cloudwatch_omni=True,
),
)
@agent.tool
def charge_card(order_id: str, amount: float) -> dict:
# Omni automatically attaches an eval score to this span
# if you've registered tool_call_success above
return payment_client.charge(order_id, amount)
The eval_hooks line is the part people skip and regret. Without it, Omni gives you a normal OTel trace — useful, but no different from what Honeycomb already does. With it, every span carries a pass/fail signal your on-call engineer can filter on directly: eval:tool_call_success=false AND service:checkout-agent.
Where it falls short
Three regions at launch is a real constraint if you’re running in Asia-Pacific — you’ll be proxying telemetry cross-region for months, which adds latency to your observability pipeline itself, ironically. The eval-hook registration is also framework-specific right now; if you’re running a hand-rolled agent loop that isn’t LangGraph, CrewAI, Strands, or the OpenAI Agents SDK, you’re writing your own OTel spans by hand, same as before Omni existed.
And natural-language querying is a demo feature until it isn’t. I tested it against a synthetic trace set and it handled straightforward “show me the slowest tool calls today” queries fine, but choked on anything involving a time comparison (“compare today’s error rate to last Tuesday’s”) — it just re-ran today’s query twice. AWS will fix this. It isn’t fixed yet.
The actual decision for your team
If you’re already running two or more agent frameworks across accounts, migrate. The OTel foundation alone is worth it even if you ignore the NL querying gimmick — it means the next observability tool you evaluate, whether that’s Omni or something else, can read the same traces without a re-instrumentation project.
If you’re single-cloud, single-framework, and your current CloudWatch dashboards work fine, there’s no fire here. Wait for the region list to grow and the eval-hook framework support to widen before you spend a sprint on migration. The problem CloudWatch Omni solves is real — I’ve watched three different clients hit it this quarter alone — but “real problem” and “urgent problem” aren’t the same thing, and for a lot of teams this is still the second one waiting to become the first.
One more thing worth planning for before you commit engineering time: migrating off proprietary CloudWatch metrics to OTel spans isn’t free even when the destination is better. Old dashboards built on CloudWatch’s metric filters don’t just port over — someone has to rebuild them against the new span schema, and in my experience that someone is whoever complains loudest about the old dashboard breaking first. Budget a week for that rebuild, not a day, and tell your on-call rotation before you flip the switch, not after the first incident where the old runbook doesn’t match what’s on screen anymore.