10 Best AI Agent Observability Tools in 2026
Short answer: Arize AX Pro at $50 a month when you need the tool path of an agent, not a prompt log. LangSmith Plus at $39 a seat when that agent is already LangGraph. Pydantic Logfire Team at $49 a month when the agent and the API beside it should land in one trace.
A looping tool call is a different failure from a bad prompt, so a prompt log will not catch it. Dupple data: Toolradar, our software directory, tracks 41 AI observability tools as of September 2026. Prompt logs and drift monitors sit in best AI observability tools and LLM observability tools. Every price on this page was checked on the vendor's own site on September 23, 2026.
Quick comparison
| Tool | Best for | Starting price | Standout |
|---|---|---|---|
| Arize AX | The path the agent took | $50/mo Pro | 50,000 spans; custom scorers on Enterprise |
| LangSmith | LangGraph agents you also deploy | $39/seat Plus | Base traces last 14 days |
| Langfuse | A cloud price and a real self-host | $29/mo Core | Spans and scores count as units |
| Pydantic Logfire | The agent and the app on one timeline | $49/mo Team | 10 million records, then per million |
| Opik | A flat cloud bill, or the full open-source build | $19/mo Pro | 100,000 spans and up to 50 members |
| AgentOps | Replay of multi-agent runs | From $40/mo Pro | Free plan stops at 5,000 events |
| Laminar | Long browser and tool agents | $30/mo Starter | Billed per GB, plus a Signals credit |
| LangWatch | Simulations before live traffic | €29/core-seat Growth | A step is an event, not a chat |
| Maxim | Simulate, then keep the live logs | $29/seat Professional | That plan keeps logs for 7 days |
| Datadog | An infra account you already pay | $160/mo annual, first 100k spans | Only LLM spans are billed |
What an AI agent observability tool does
An AI agent observability tool records a multi-step agent run: each model call, tool call, and handoff, plus the cost and the outcome. A prompt log never shows a tool the agent called twice.
Watch the agents in AI agents, agent platforms, and coding agents. A downed database or host still belongs on an observability platform, and retrieval quality is a separate buy in RAG tools, since neither shows up in an agent trace.
How we chose
The AI observability category on Toolradar lists 41 tools, updated September 2026. From that set I ranked 10 a team can buy when the thing to see is an agent run: tool spans, sessions, handoffs, and the bill. Vendor sites supplied the plan names and prices, rechecked on September 23, 2026. This ranking did not run a month of production traffic through each product, and nobody paid for a slot. See Toolradar's agent observability guide.
Arize AX, the pick when you need the path
Arize AX is the hosted product for agent trajectories, with trajectory views on every plan, and Phoenix is the open-source tool you run locally.
Who it's for: A team that wants the graph of tool calls and can stay on SaaS until custom scorer code is required.
The pricing page puts AX Pro at $50 a month: 30 days of retention, 10 GB, 25 Signal issues, and 50,000 spans, billed by usage rather than by seats. AX Free keeps 15 days, 1 GB, 10 Signal issues, and 25,000 spans, a prototype rather than a production month. Users and evals are free on both, hosting can be the US, the EU, or Canada, and your own scorer plus self-hosting wait for Enterprise.
The catch: Phoenix stays local, so the span cap and the Signal issue cap never apply to it.
The default when the question is which tool the agent called, and the wrong buy when the scorer must be your own code.
LangSmith, when the agent is LangGraph
LangSmith is the trace and deploy product from the LangChain team, so settle the framework first in AI agent frameworks.
Who it's for: Apps on LangChain or LangGraph that may also want Deployment or Engine on the same bill.
On the pricing page, Developer covers only one seat and 5,000 base traces a month with no seat fee, then pay as you go. Plus is $39 a seat a month for unlimited seats, a base allowance of 10,000 traces, and one small serverless deployment. A compute unit is $1.50, LangChain estimates an Engine run at about 5 to 30 units, and thirty units is $45 on top of the seats. A removed seat is not credited mid-month, and Enterprise is the self-host plan.
The catch: Base traces last 14 days, so last month's incident may already be gone. The pricing FAQ still lists extended traces at 400 days, while the usage and billing doc caps new SaaS extended traces at 180 days from September 14, 2026. Evaluators upgrade a base trace by default. One trace stops at 25,000 runs, and without a card Developer hard-stops at the monthly cap.
Right for a LangGraph codebase you also want to deploy, and the wrong bill off that stack.
Langfuse, when you may self-host the agent graph
Langfuse traces agents and runs LLM-as-a-judge evals, including when you self-host, with session tracking on every cloud plan.
Who it's for: A production agent where a second engineer already breaks Hobby's 2-user cap.
On the pricing page, Hobby is free with 50,000 units a month and 30 days before that history drops. Core is $29 a month for 100,000 units, 90 days, and no user cap. Further units cost $8 per 100,000 up to 1 million, then $7, then $6.50, then $6 per 100,000. That meter is structure, not chats. Pro is $199 a month for 3 years of data, SSO is a Teams add-on at $300 a month, Enterprise is $2,499 a month, and self-hosting has no license fee.
The catch: A unit is a trace, a span, an event, a generation, or a score. The published Core example for 1 million units a month is $101. Hobby ingests 1,000 requests a minute and Core raises that to 4,000, so a spike can force the upgrade on rate.
The readable cloud price, and the self-host to pick, once you budget spans and scores rather than chats.
Pydantic Logfire, when the agent is not the whole app
Pydantic Logfire traces agent runs and the services around them on one OpenTelemetry timeline, including Pydantic AI, LangChain, and the OpenAI SDK.
Who it's for: A team whose agent calls your own API and wants that API on the same trace.
The pricing page includes $20 of usage a month on every plan, equal to 10 million records. Personal is free and hard-capped there, with 1 seat and 30 days. Team is $49 a month, with 5 seats included, extra seats at $25 up to 12 seats, and 30 days. Past the allowance, records are $2 per million and scores are not metered. Growth is $249 a month with unlimited seats and up to 90 days. Built-in gateway keys carry a 5% markup on Team, and your own keys do not, up to 3 of them.
The catch: Personal stops at the included records, so a second engineer means a paid plan. The spending cap is an email to accounts@pydantic.dev. A 13th seat cannot stay on Team, and SSO on day one is an Enterprise conversation.
The right trace when the bug might be your own service.
Opik, the flat cloud price with a real self-host
Opik, from Comet, logs agent spans and ships test suites, a playground, and the same codebase self-hosted.
Who it's for: A team that wants one monthly number, or the open-source build, and can accept a short cloud archive.
The pricing page lists the open-source build at no charge with the full observability and agent-testing set. Free Cloud is 25,000 spans a month, up to 10 members, and 60 days. Pro is $19 a month, not per seat, with up to 50 members, 100,000 spans, and 60 days. Extra spans on Pro are $5 per 100,000, stretching retention from 60 days to 400 days is a separate per-span charge, and Enterprise adds SSO plus the compliance list.
The catch: Guardrails for PII, topic, and custom rules are marked for the self-hosted deployment, so the cloud plan will not block a bad call. Fifty-one members moves you to Enterprise.
The cheapest published cloud price on this list when the Pro span allowance covers the month, and the self-host when guardrails must run in your network.
AgentOps, replay first and compliance later
AgentOps traces LLM calls, tools, and multi-agent interactions, with replay of a run. The homepage names OpenAI, CrewAI, and Autogen among the integrations, enough to start a prototype.
Who it's for: Engineers debugging a multi-agent prototype who can wait on SSO and a self-host.
The homepage lists Basic at no charge up to 5,000 events, with cost tracking and replay analytics, a prototype cap rather than a production one. Pro starts at $40 a month and lists an unlimited event limit, unlimited log retention, session export, and role-based permissions. Enterprise is custom and adds the SLA, custom SSO, on-prem on AWS, GCP, or Azure, and SOC 2, HIPAA, and NIST AI RMF.
The catch: "Starts at" is not a ceiling, so a busy month can leave that Pro line. AgentOps does not define an event, so you cannot forecast the free cap from chat volume, and compliance plus self-hosting are a separate purchase.
A focused replay tool while the agent is still a prototype, and a sales process once security review starts.
Laminar, when the agent is long and the bill should be gigabytes
Laminar traces browser and tool agents, with a debugger and a plain-language Signals layer that reads traces for failures you describe.
Who it's for: Teams whose agents run for many steps, especially in a browser, and who would rather pay for stored data.
The pricing page lists Free at 1 GB, a $5 Signals credit, 7 days, 1 project, and 1 seat, with no data overage, so you cannot buy more storage for a long browser agent. Starter is $30 a month for 3 GB, then $2 per GB, a $15 Signals credit, 30 days, and unlimited projects and seats. Pro is $150 a month for 10 GB, then $1.50 per GB, a larger Signals credit, 6 months, and Slack support, and Enterprise is custom and includes on-prem. The on-page calculator is labeled an estimate, because Signals are billed by tokens the internal agent spends, not by your model's tokens.
The catch: Two meters means a quiet storage month can still spend the Signals credit if you analyze every run. Self-hosting the open-source images leaves Signals and Slack alerts out until an enterprise license key is added.
The volume buy for long agents, once you accept that Signals are not in the free self-host.
LangWatch, when you want to simulate the user first
LangWatch traces agent graphs and runs multi-turn simulations, including a judge that can pass or fail a turn, before that traffic exists in production.
Who it's for: A team that gates releases on scenarios, and that can read a price in euros.
The pricing page lists Developer at no charge: 50,000 events a month, 14 days, 2 users, and 3 simulations, a look rather than a release suite. Growth is €29 per core seat a month, with 200,000 events, then €5 per 100,000, 30 days of retention, and €3 per GB past that window. Lite users are unlimited, so a viewer does not take a core seat. Enterprise holds SSO and the self-host compliance extras, and you can self-host with Docker Compose.
The catch: An event is a step, and the pricing FAQ counts every LLM call, tool call, retrieval, evaluation, or simulation step, so one user turn is several events. Four core seats are €116 a month before that meter.
The pick when the release gate is a simulated conversation, and a poor fit if you needed a dollar price and a year of logs.
Maxim, simulate on a seat and store logs for a week
Maxim packages agent simulation, evaluation, and logs. Bifrost, the gateway, is a separate plan, so these prices are the agent platform and not the gateway.
Who it's for: A small team that will simulate agents and then watch a week of production logs.
The evals pricing page lists Developer free for up to 3 seats, 1 workspace, up to 10,000 logs a month, and 3 days of retention, with no overage allowed, so a Friday incident can vanish by Monday. Professional is $29 a seat a month, with a 14-day trial, unlimited seats, up to 3 workspaces, up to 100,000 logs, 7 days, simulation runs, online evals, and extra logs at $1 per 10,000. Business is $49 a seat a month, with up to 500,000 logs, 30 days, RBAC, and PII management, and Enterprise adds SSO and in-VPC.
The catch: Seven days will not support a monthly review, five Professional seats are $145 before overage, and a yearly audit trail is Enterprise because even Business stops at 30 days.
Useful when simulation and a short production log should be one vendor, rather than a monthly archive.
Datadog, if the on-call rotation already lives there
Datadog Agent Observability traces agents beside the services you already page on. Host monitoring stays a different product, compared in observability platforms.
Who it's for: Teams that already page from Datadog and do not want a second on-call tool.
The product page includes 40,000 LLM spans a month at no charge, with 15-day retention. The pricing list sets the first 100,000 spans at $160 a month on an annual bill, with $200 month to month and $240 on demand. Annual overage at 15-day retention is $3.50 per 10,000 spans, month to month it is $4.20 per 10,000, and on demand it is $5 per 10,000. Only provider calls are billed, so tool, workflow, agent, embedding, and retrieval spans are free, while an eval call to a model still counts.
The catch: A month-to-month account does not get the annual rate. Ninety-day trace retention is $7.50 per 10,000 spans on the annual column, on top of the span block.
The agent add-on to buy when Datadog is already on the invoice.
What you will actually pay
A five-seat LangSmith Plus team is $195 a month before any traces. Langfuse's published 1 million unit month on Core is the $101 example, and it counts spans and scores. On Datadog's annual rate, 200,000 spans past the first block, at $3.50 per 10,000, add $70, which is $230 before a longer retention add-on, and only on an annual contract.
How to choose for the agent in front of you
Use Arize for the tool path, LangSmith when the agent framework is LangGraph, Langfuse to self-host, and Logfire when your own API is on the trace. Laminar fits a long run, LangWatch or Maxim fit a simulation, Opik is the flat or open-source build, and AgentOps is replay before a security review. Use Datadog only beside a bill you already pay. A gateway is a different buy, in LLM gateways, and training pipelines belong in MLOps tools.
Common mistakes when pricing an agent trace
Treating a Langfuse unit as one chat spends the allowance on spans and scores before the chat count looks large.
Leaving the LangSmith evaluator upgrade on pushes traces out of the short base window, and a removed seat is not credited.
Using Datadog's annual column on a month-to-month account skips the higher list prices for the same first block.
Treating AgentOps Pro as a flat bill ignores the words "starts at". AgentOps does not define an event, and SOC 2 plus on-prem sit on Enterprise.
Adding a 13th person to Logfire Team breaks the seat cap and forces Growth.
Reading Laminar gigabytes as raw tokens ignores compression and the second meter, and the open-source self-host omits Signals.
Buying Maxim Professional for a monthly review still leaves a short log window, because longer retention starts on Business.
FAQ
What is the best AI agent observability tool in 2026?
Arize AX Pro, for most teams whose product is an agent: $50 a month for 50,000 spans, with trajectory views on that plan and custom code evaluators held for Enterprise. LangSmith Plus at $39 a seat fits LangGraph and is the wrong seat if the app is not on that framework. Langfuse Core fits a team that may self-host. Those 41 directory entries are not 41 versions of the same agent tracer.
How much does AI agent observability cost in 2026?
In September 2026, published entry points run from Opik Pro at $19 a month and Laminar Starter at $30 a month, through AgentOps Pro starting at $40 a month, up to Langfuse Pro at $199 a month and Logfire Growth at $249 a month. Datadog's month-to-month block is $200 if the contract is not annual. LangWatch Growth is €29 per core seat. A five-seat LangSmith Plus team is $195 a month before traces.
Is there a free AI agent observability tool in 2026?
Langfuse Hobby includes 50,000 units and 2 users, and self-hosting has no license fee. Phoenix runs on your machine, and Arize AX Free includes 25,000 spans. Opik Free Cloud includes 25,000 spans, and the open-source build includes agent testing. Logfire Personal includes 10 million records, then pauses instead of billing. AgentOps Basic includes 5,000 events, Datadog includes 40,000 LLM spans, LangWatch Developer includes 50,000 events, and Maxim Developer includes 10,000 logs for 3 days.
Should I use Arize or LangSmith for an agent?
Use Arize when the framework might change: Pro keeps 30 days, and your own scorer code waits for Enterprise. Use LangSmith when the app is LangGraph: Plus includes 10,000 base traces that last 14 days, evaluators can extend that retention, and the September 14, 2026 docs cap the longer window at 180 days on SaaS, with at most 25,000 runs in one trace.
Does our current Datadog contract include agent traces?
The agent view sits next to services you already monitor, and it is a separate line from the infrastructure invoice. The free allowance is 40,000 LLM spans, the annual price for 100,000 spans is $160 a month, and on demand that block is $240. Only provider calls are billed, including calls made inside evals, and tool spans are not the line item.
Is Langfuse or Logfire the better self-host?
Langfuse runs with no license fee, and Core is the cloud comparison at 100,000 units and 90 days. Logfire's self-host is Enterprise, while Team keeps 30 days and a 12-seat cap. Pick Langfuse for the agent graph on your own machines, and Logfire when your own services must share the timeline.
How is agent observability different from LLM observability?
Agent observability keeps the tool calls, handoffs, and session, so you can see a loop. LLM observability is the prompt, the completion, the token cost, and often a quality score. You want both when a bad answer and a bad tool choice can each ship, and the prompt-and-score comparison is LLM observability tools.
Bottom line
Start on Arize AX Pro when you need the path the agent took, and skip it when the scorer must be your own code. Pay for LangSmith only if LangGraph is already the framework, and use Langfuse when a self-host matters. Put Logfire on the trace when the bug may be your API, and add Datadog only if that account already exists.
The price moves on this list are what Dupple X covers each week, you can start the trial, and the about page names the editorial team.
Cite this: Dupple, "10 Best AI Agent Observability Tools in 2026", September 2026.
