10 Best AI Observability Tools in 2026
Short answer: Langfuse Core at $29 a month when you want a published cloud price and a free self-host. Arize AX Pro at $50 a month when the product is an agent and Phoenix can stay local. LangSmith Plus at $39 a seat when the code is already LangChain or LangGraph.
I would not buy a second prompt logger when the failure is a drifting feature store or an agent that loops on tools. Dupple data: Toolradar, our software directory, tracks 41 AI observability tools as of September 2026. Prompt-level tracing is a separate comparison in best LLM observability tools, and host monitoring stays in observability platforms. Every price below was checked on the vendor's own site on September 23, 2026.
Quick comparison
| Tool | Best for | Starting price | Standout |
|---|---|---|---|
| Langfuse | A cloud price plus a real self-host | $29/mo Core | Spans and scores count as units |
| Arize AX | Agent paths, with Phoenix local | $50/mo Pro | 50k spans, custom code evals on Enterprise |
| LangSmith | LangChain or LangGraph apps | $39/seat Plus | Evaluators can extend trace retention |
| Braintrust | Teams that score runs on purpose | $249/mo Pro | Extra scores are $1.50 or $2.50 per 1,000 |
| Fiddler | Guardrails, then a per-trace meter | $0.002/trace | SSO sits on the usage plan |
| Galileo | Evals inside a trace cap | $100/mo on yearly billing | Guardrails are an Enterprise feature |
| Helicone | A proxy in front of the provider | $79/mo Pro | Hobby keeps logs for 7 days |
| Weights & Biases | Experiments already live in W&B | $60/mo Pro | Extra Weave ingestion is $0.10/MB |
| Datadog | An account that already pays for infra | $160/mo annual for 100k spans | Only LLM spans are billed |
| Evidently | Drift and evals you will host | No license fee | Hosted cloud is no longer sold |
What an AI observability tool does
An AI observability tool records a model or agent in production: the trace, the quality score, the token cost, and whether the inputs drifted, which an uptime dashboard does not show you.
Langfuse, LangSmith, and Helicone are trace logs for the model record. Arize, Fiddler, Braintrust, and Galileo add scores and agent paths for a score, a guardrail, or a tool loop. Evidently is the library you run yourself when the question is drift. Datadog and Weights & Biases win only when the rest of the stack is already there. Training pipelines belong in MLOps tools, not on this list.
How we chose
Toolradar's AI observability category lists 41 tools, updated September 2026, and I ranked 10 a team can buy for traces, agent paths, scores, guardrails, or drift. Plan names and prices come from each vendor's site on September 23, 2026. I did not send a month of production traffic through each product, no vendor paid for a slot, and the directory's own order is a different list from this one.
I left WhyLabs off the ranking because its pricing page says the company is discontinuing operations, while whylogs and langkit stay open source. A tool that is shutting down is not a system of record for production agents.
Langfuse, the default when you may self-host
Langfuse traces agents, versions prompts, and runs LLM-as-a-judge evals, including when you may later self-host.
Who it's for: A production app past a weekend log, where a second engineer already breaks Hobby's 2-user cap.
On the pricing page, Hobby is free with 50,000 units a month and 30 days of data, so anything older than a month is gone. Core is $29 a month with 100,000 units, 90 days, and unlimited users.
Further units are $8 per 100,000 up to 1 million, then $7, then $6.50, then $6 per 100,000. Pro is $199 a month when you need 3 years of data and SOC 2. SSO is a Teams add-on at $300 a month, and Enterprise is $2,499 a month. Self-hosting has no license fee, and early-stage startups can get 50% off the first year.
The catch: A unit is a trace, a span, an event, a generation, or a score, so one request with a trace, five spans, and five scores is 11 units of structure, not chats. Langfuse's published example for 1 million units a month on Core totals $101. Hobby ingests 1,000 requests a minute and allows one annotation queue, while Core raises that to 4,000 requests a minute and 3 queues.
Default cloud buy, and the cleanest self-host, once you count spans and scores rather than user messages.
Arize AX, the pick when the product is an agent
Arize AX is the hosted product, and Phoenix is the open-source local tool. Retrieval quality is a separate buy, in RAG tools and vector databases.
Who it's for: Teams shipping agents who want trajectory views and online evals, and who can live with SaaS until Enterprise.
The pricing page lists AX Pro at $50 a month: 50,000 spans, 10 GB, 30 days of retention, and 25 Signal issues. AX Free is 25,000 spans, 1 GB, 15 days, and 10 Signal issues, with unlimited users and unlimited evals on both, so the upgrade is volume rather than seats, and the hosted regions are the US, EU, or Canada. Enterprise is custom, adds self-hosting, and is where custom code evaluators appear.
The catch: Custom code evaluators sit on Enterprise, so a team that needs its own scorer code is not on Pro. Phoenix has no hosted span cap on that page because it runs locally, and the alerting lives on AX.
The agent pick when Phoenix is the prototype and Pro's span allowance covers production.
LangSmith, only if the app is already LangChain
LangSmith is the trace and deploy product from the LangChain team, and that depth is wasted on a raw SDK, so settle the framework first in AI agent frameworks.
Who it's for: Apps on LangChain or LangGraph, including teams that want Deployment, Fleet, or Engine on the same bill.
The pricing page lists Developer at no seat fee for 1 seat and 5,000 base traces a month, then pay as you go. Plus is $39 a seat a month, with unlimited seats, 10,000 base traces, and one free small serverless deployment. A compute unit is $1.50, and LangChain estimates an Engine run at about 5 to 30 of those units. Removed seats are not credited, and Enterprise is the self-host and hybrid plan.
The catch: The usage docs say a change that started September 14, 2026 caps extended SaaS retention at 180 days, while the pricing FAQ still says 400 days, so budget the shorter window. Evaluators and automation rules upgrade a base trace to extended retention by default, and extended traces cost more, while base traces last 14 days. One trace stops at 25,000 runs, and without a card Developer hard-stops at 5,000 traces a month.
Right for a LangChain or LangGraph codebase, and the wrong default for everything else.
Braintrust, when you pay for scores
Braintrust prices a platform fee separately from processed data and from scores.
Who it's for: Teams that will score a real share of traces, and that can wait until Enterprise for SAML SSO.
The pricing page lists Starter with no platform fee, $10 in model credits, 1 GB then $4 per GB, 10,000 scores then $2.50 per 1,000, and 14 days of retention. Pro is $249 a month, with $100 in model credits, 5 GB then $3 per GB, 50,000 scores then $1.50 per 1,000, and 30 days. Storage past that window is $0.50 per GB a month, up to 180 days, and users are unlimited on Starter, so seats are not the meter.
The catch: Twenty thousand scores on Starter are the included 10,000 plus $25. One hundred thousand scores on Pro are the included 50,000 plus $75, on top of the platform fee. SAML is Enterprise-only, and the Pro card's RBAC line should be read against the order form.
Buy it when scores are the product, and skip it when a trace log is the whole job.
Fiddler, guardrails free and traces on a meter
Fiddler puts real-time guardrails on the free plan and observability for agentic and predictive systems on the usage plan. The free plan is not a free trace store.
Who it's for: Teams that want hallucination, toxicity, PII, prompt-injection, and jailbreak checks, and that will pay per trace for experiments and custom evaluators.
The pricing page states guardrail latency under 80 milliseconds. Developer is $0.002 per trace, which is $2 per 1,000 traces, so 100,000 traces are $200. That plan adds custom evaluators, a bring-your-own judge, and SSO, on SaaS. Enterprise adds VPC or on-prem and does not publish a second rate.
The catch: Developer has no monthly trace cap, so a busy month keeps the meter running. The free tier does not include the observability product.
The per-trace option when guardrails and predictive monitoring belong in one vendor.
Galileo, evals up to 50,000 traces
Galileo packages observability and evaluations with a free allotment and a Pro cap you can read before the contract.
Who it's for: A team inside 50,000 traces a month that does not need real-time guardrails yet, because unlimited custom evals are already on the free plan.
The pricing page lists Free at 5,000 traces a month and unlimited users. Pro is $100 a month on yearly billing, with a 33% yearly saving shown on the page, 50,000 traces, standard RBAC, and Slack support, though the page says the price scales with traces. Enterprise adds unlimited traces, VPC or on-prem, and real-time guardrails.
The catch: Guardrails are Enterprise, not Pro, so the published seat will not block a bad output. A yearly price with a scaling clause is not a ceiling, and the page does not print a month-to-month rate beside that yearly figure.
A clear eval seat inside that trace band, and a sales conversation once guardrails are required.
Helicone, a proxy with a short free memory
Helicone logs requests through a gateway or your own provider keys, and broader gateway choices are in LLM gateways.
Who it's for: Teams that want one proxy, with caching and fallbacks, and that can accept 7 days of memory on the free plan.
The pricing page includes 10,000 requests, 1 GB, 1 seat, 7 days, and 10 logs a minute on Hobby. Pro is $79 a month, with a 7-day trial, unlimited seats, 1 month of retention, 1,000 logs a minute, and usage-based pricing after 10,000 requests. Team is $799 a month, with SOC 2, HIPAA, 3 months of retention, and 15,000 logs a minute. Gateway credits are 0% markup plus payment processing fees. Startups under 2 years old and under $5 million raised can get 50% off the first year.
The catch: Ten logs a minute will drop a production spike, and 7 days will not support a monthly review, so Hobby only proves the wiring. Pro does not print a per-request overage, and provider token cost is a separate bill.
The proxy buy when the log should sit on the same hop as the model call.
Weights & Biases, if the experiments already live there
Weights & Biases puts Weave tracing next to experiment tracking.
Who it's for: Groups that already log training runs in W&B. A company of 50 or more employees is outside Pro and moves to Enterprise.
The pricing page lists Free with up to 5 model seats, 5 GB of storage, and 1 GB of Weave ingestion a month. Pro starts at $60 a month, with up to 10 model seats, 100 GB then $0.03 per GB, and 1.5 GB of Weave ingestion then $0.10 per MB. Pro is limited to teams under 50 employees and includes a 30-day trial. The personal self-host license is one seat and forbids corporate use.
The catch: Two hundred megabytes past the Pro Weave allowance is $20, because overage is priced per megabyte. The public grid does not print a Weave seat count, so confirm it before you add collaborators.
Sensible when W&B is already the system of record, and an expensive trace log if you arrived only for production monitoring.
Datadog, if the infra account already exists
Datadog Agent Observability traces agents beside the rest of the account. Infra tools are compared on the observability platforms page.
Who it's for: Teams whose on-call already lives in Datadog, not a startup buying its first tracer.
The product page includes 40,000 LLM spans a month free, with 15-day retention. The pricing list prices the first 100,000 spans at $160 a month on an annual bill, $200 month to month, and $240 on demand.
Extra spans at 15-day retention are $3.50 per 10,000 annually, $4.20 month to month, and $5 on demand. At 30, 60, and 90 days the annual rates per 10,000 are $5, $6.50, and $7.50. Only provider calls are billed, while tool, workflow, agent, embedding, and retrieval spans are free, and eval LLM calls still count. Infrastructure Pro on that list is $15 a host a month annually, on a different line.
The catch: Three hundred thousand spans on the annual rate are the first block plus $70, which is $230 before longer retention, and tool spans are not what the invoice counts.
The right add-on when Datadog is already paid for, and a large new bill when it is not.
Evidently, self-host only
Evidently is the open-source library for evals, tests, and drift on LLM apps, agents, and predictive models. The hosted login is not what the docs offer now.
Who it's for: Teams that will run the platform themselves. It is the wrong pick if you wanted a vendor-hosted UI next week.
The product page describes an Apache 2.0 library, listing 7,500+ GitHub stars, 40 million+ downloads, and 100+ metrics, and it prints no cloud price. The cloud setup doc says Cloud is no longer available as SaaS, and Enterprise is a conversation rather than a rate card.
The catch: A hosted-seat budget will stall, because your team owns storage, upgrades, and on-call for the monitor itself.
The self-host for drift and tests, after you accept that Cloud is not a plan you can sign up for.
How to choose for the failure in front of you
Use Langfuse when you need to see what the model did and you may self-host, Arize when the tool path matters, and LangSmith when the agent framework is LangChain or LangGraph. Use Braintrust or Galileo when the release gate is a score, and Fiddler when it is a guardrail on a meter.
Use Helicone when the integration has to be a proxy, Weights & Biases or Datadog only beside a bill you already pay, and Evidently when you will operate the service. The products being watched are in AI agents, agent platforms, and AI for coding.
Common mistakes when pricing this stack
Counting Langfuse units as chats spends the allowance on spans and scores. The published 1 million unit Core example is $101 a month, and it counts structure, not messages.
Leaving LangSmith evaluators on the default upgrade moves traces off the 14-day base window. Extended retention costs more, and the September 14, 2026 docs cap SaaS at 180 days while the pricing FAQ still says 400.
Quoting Datadog's annual rate on a month-to-month account ignores the $200 and $240 list prices for the same first 100,000 spans. A 300,000-span annual month is $230 before longer retention, and tool spans are not the billed unit.
Treating Fiddler's free plan as free traces buys guardrails, then meets a per-trace Developer rate with no written cap.
Putting 50 people on W&B Pro breaks a rule that Pro is for fewer than 50 employees. At $0.10 per MB, 200 MB over the 1.5 GB Weave allowance is $20.
Buying Galileo Pro for guardrails still leaves real-time guardrails on Enterprise. The free 5,000 traces and unlimited custom evals are the trial.
Asking for Evidently Cloud follows a closed signup, and the current doc says that SaaS is gone, so the choice is self-host or Enterprise.
FAQ
What is the best AI observability tool in 2026?
Langfuse, for most teams that want a cloud plan and a self-host: Core is $29 a month with 100,000 units, and self-hosting has no license fee. Arize AX Pro at $50 a month fits agents once Phoenix is not enough locally. LangSmith Plus at $39 a seat fits LangChain or LangGraph. The 41 tools in the directory are not 41 copies of the same tracer.
How much do AI observability tools cost in 2026?
In September 2026, published entry points run from Fiddler at $0.002 per trace and Helicone Pro at $79 a month, through Galileo Pro at $100 a month on yearly billing, up to Braintrust Pro at $249 a month. Datadog's annual block for the first 100,000 LLM spans is $160. Langfuse's 1 million unit Core example is $101, and Evidently publishes no hosted price. Five LangSmith Plus seats are $195 a month before traces.
Is there a free AI observability tool in 2026?
Langfuse Hobby includes 50,000 units, 30 days, and 2 users. Arize AX Free includes 25,000 spans, and Phoenix runs locally with no license fee. Galileo Free includes 5,000 traces, Helicone Hobby includes 10,000 requests and 7 days, and Datadog includes 40,000 LLM spans. Braintrust Starter has no platform fee and then meters data and scores, Fiddler's free plan is guardrails rather than stored traces, and Evidently's library is Apache 2.0, with its hosted cloud closed.
Should I use Langfuse or LangSmith?
Use Langfuse when the framework might change and a self-host matters. A unit is a trace, a span, or a score, and Core includes 100,000 of them. Use LangSmith when the app is on LangChain or LangGraph. Plus includes 10,000 base traces a seat, those traces last 14 days, and evaluators can extend retention, which the September 14, 2026 docs cap at 180 days on SaaS. One LangSmith trace holds many steps, up to 25,000 runs.
Is Datadog enough if we already pay for it?
It covers LLM spans next to services you already monitor, and it is not included in the infra bill. Free is 40,000 LLM spans, then 100,000 spans are $160 a month on the annual bill. Only provider calls are billed, including calls made inside evals, and host monitoring stays a separate line.
What happened to WhyLabs and Evidently Cloud?
WhyLabs says it is discontinuing operations, so it is not a system of record for production agents. whylogs and langkit are not a hosted successor with a price. Evidently's docs say Cloud is no longer sold as SaaS. The library remains Apache 2.0, with 100+ metrics on the product page, and Enterprise hosting is a quote.
Do I need an observability platform as well?
You need one when the failure might be the database, the GPU, or the path outside the model. AI observability tells you the model or agent was wrong or expensive. An observability platform tells you the service was down.
Bottom line
Start on Langfuse Core when you want a price you can read and a self-host you can run, move to Arize when the product is an agent, and pay for LangSmith only if LangChain or LangGraph is already the framework. Add Datadog or Weights & Biases when those accounts exist, and do not budget Evidently as a hosted login.
Dupple X is the weekly briefing behind these price changes, or start the trial. No vendor paid for a slot on this page, and the about page names the editorial team.
Cite this: Dupple, "10 Best AI Observability Tools in 2026", September 2026.
