AI evaluation platforms
8 apps ranked by verified data. See how ranking works
Opik is an open-source LLM evaluation platform from Comet for tracing, evaluating, and monitoring LLM applications, with datasets, LLM-as-judge metrics, and a prompt playground.
Langfuse is an open-source LLM engineering platform providing tracing, prompt management, evaluations, datasets, and analytics for LLM applications; self-hostable or cloud.
Phoenix is an open-source observability and evaluation library from Arize AI for tracing, evaluating, and troubleshooting LLM applications, built on OpenTelemetry.
Braintrust is an evaluation and observability platform for AI products, with datasets, LLM-as-judge scorers, prompt playground, and production logging.
HoneyHive is an AI evaluation and observability platform for testing, tracing, and monitoring LLM applications, with datasets, LLM-as-judge evaluators, and human review.
LangSmith is a platform for debugging, testing, evaluating, and monitoring LLM applications, with tracing, datasets, and LLM-as-judge evaluations.
Weave is the Weights & Biases toolkit for tracking, evaluating, and monitoring LLM applications, with tracing, scorers, and experiment comparison.