LLM Observability
Tools for tracing, evaluating, and debugging LLM applications and agent runs in production.
-
LangSmithLangChain's observability and evaluation platform for LLM applications, tracing every step of a chain or agent run and scoring outputs against test datasets.
-
LangfuseOpen-source LLM observability platform for tracing, evaluating, and debugging prompts and agent runs, self-hostable or available as a managed cloud service.
-
HeliconeOpen-source LLM observability proxy that sits between your app and a model provider, logging requests, costs, and latency with a one-line integration change.
-
BraintrustEvaluation-first platform for LLM applications, treating prompt and model changes like code changes that need regression tests before shipping.
-
PortkeyAI gateway for production LLM apps, combining provider routing and failover with caching, guardrails, and observability in one managed layer.
-
Arize Phoenix
Open-source LLM observability and evaluation platform built on OpenTelemetry, with tracing, datasets, experiments and prompt tooling.
-
Opik
Comet's open-source platform for logging traces, running evaluations and monitoring LLM apps and agents in development and production.
-
W&B Weave
Weights & Biases toolkit for tracing, evaluating and monitoring LLM applications, with automatic call logging and scorers.