AltCatalog

Alternatives to Arize Phoenix

Open-source LLM tracing and evaluation from Arize AI.

Alt Score

Alt Score · 96

How this alternative ranks. How it works →

Verified coverage 90%96
Visibility 10%100

Verified coverage = sourced yes and partial answers in the category checklist. Visibility = relative app-page views on AltCatalog over 30 days, neutral below 500 category views. Payment, votes, and vendor opinions are never score inputs.

#3 of 8 in AI evaluation platforms

Arize Phoenix ranks #3 of 8 in AI evaluation platforms, with an Alt Score of 96. It is licensed under Elastic License 2.0 (ELv2), free from Free and available on the web. 13 of 13 checklist rows are verified against a public source.

Phoenix is an open-source observability and evaluation library from Arize AI for tracing, evaluating, and troubleshooting LLM applications, built on OpenTelemetry.

Most compared with Comet OpikLangfuseBraintrust
Official site Suggest an edit Data history Work on Arize Phoenix? Claim this page
Who it's for

AI engineers and teams building LLM applications and agents who need to trace, evaluate, and debug their systems in development or production — from individual developers self-hosting for free to teams needing full data control or air-gapped deployments.

What you get

Tracing built on OpenTelemetry, LLM-as-judge and code-based evaluations, a prompt playground with prompt versioning, dataset management, and experiment comparison, plus token cost tracking and human annotation tools — delivered as a source-available (Elastic License 2.0) platform that is free to self-host with no feature gates, or usable via a free Phoenix Cloud instance with 10 GiB of storage.

How it works

You instrument your application with OpenTelemetry/OpenInference to send traces to a Phoenix instance (self-hosted via Docker, Kubernetes/Helm, or run in Phoenix Cloud), then use the web UI or the Python/TypeScript SDK to run evaluations on those traces, annotate them, manage and test prompt versions, and run experiments against datasets — including gating changes in CI with the pytest integration.

Why people leave Arize Phoenix

Dashed reasons are sourced facts; the rest are opinions. Vendors can dispute.

Sign in to add a reason — new reasons go through moderation before appearing.

Ranked alternatives

Ordered by Alt Score. Click any score to see the breakdown.

Sponsored Paid slot. Never affects the ranked order below. Promote here →
01

Opik is an open-source LLM evaluation platform from Comet for tracing, evaluating, and monitoring LLM applications, with datasets, LLM-as-judge metrics, and a prompt playground.

Apache-2.0 Web
Alt Score

Alt Score · 100

How this alternative ranks. How it works →

Verified coverage 90%100
Visibility 10%100

Verified coverage = sourced yes and partial answers in the category checklist. Visibility = relative app-page views on AltCatalog over 30 days, neutral below 500 category views. Payment, votes, and vendor opinions are never score inputs.

02

Langfuse is an open-source LLM engineering platform providing tracing, prompt management, evaluations, datasets, and analytics for LLM applications; self-hostable or cloud.

MIT Web
Alt Score

Alt Score · 100

How this alternative ranks. How it works →

Verified coverage 90%100
Visibility 10%100

Verified coverage = sourced yes and partial answers in the category checklist. Visibility = relative app-page views on AltCatalog over 30 days, neutral below 500 category views. Payment, votes, and vendor opinions are never score inputs.

03

Braintrust is an evaluation and observability platform for AI products, with datasets, LLM-as-judge scorers, prompt playground, and production logging.

Proprietary (platform); MIT (autoevals library) Web
Alt Score

Alt Score · 96

How this alternative ranks. How it works →

Verified coverage 90%96
Visibility 10%100

Verified coverage = sourced yes and partial answers in the category checklist. Visibility = relative app-page views on AltCatalog over 30 days, neutral below 500 category views. Payment, votes, and vendor opinions are never score inputs.

04

HoneyHive is an AI evaluation and observability platform for testing, tracing, and monitoring LLM applications, with datasets, LLM-as-judge evaluators, and human review.

Proprietary Web
Alt Score

Alt Score · 93

How this alternative ranks. How it works →

Verified coverage 90%92
Visibility 10%100

Verified coverage = sourced yes and partial answers in the category checklist. Visibility = relative app-page views on AltCatalog over 30 days, neutral below 500 category views. Payment, votes, and vendor opinions are never score inputs.

05

LangSmith is a platform for debugging, testing, evaluating, and monitoring LLM applications, with tracing, datasets, and LLM-as-judge evaluations.

Proprietary (platform); MIT (client SDK) Web
Alt Score

Alt Score · 93

How this alternative ranks. How it works →

Verified coverage 90%92
Visibility 10%100

Verified coverage = sourced yes and partial answers in the category checklist. Visibility = relative app-page views on AltCatalog over 30 days, neutral below 500 category views. Payment, votes, and vendor opinions are never score inputs.

06

Weave is the Weights & Biases toolkit for tracking, evaluating, and monitoring LLM applications, with tracing, scorers, and experiment comparison.

Apache-2.0 Web
Alt Score

Alt Score · 93

How this alternative ranks. How it works →

Verified coverage 90%92
Visibility 10%100

Verified coverage = sourced yes and partial answers in the category checklist. Visibility = relative app-page views on AltCatalog over 30 days, neutral below 500 category views. Payment, votes, and vendor opinions are never score inputs.

07

Helicone is an open-source LLM observability platform offering request logging, cost and token tracking, caching, prompt management, and evaluations via a proxy or async SDK.

Apache License 2.0 Web
Alt Score

Alt Score · 76

How this alternative ranks. How it works →

Verified coverage 90%73
Visibility 10%100

Verified coverage = sourced yes and partial answers in the category checklist. Visibility = relative app-page views on AltCatalog over 30 days, neutral below 500 category views. Payment, votes, and vendor opinions are never score inputs.

Feature comparison

Rows come from the AI evaluation platforms checklist (17 rows). Human-verified cells only. ? means the value has not been verified.

Comparing Arize Phoenix Comet Opik × Langfuse × Braintrust × HoneyHive × LangSmith ×
+ Add app
Helicone W&B Weave
AI evaluation platforms checklist Arize PhoenixComet OpikLangfuseBraintrustHoneyHiveLangSmith
Pricing model
Starts at
License
Platforms
Open source
Self-hostable
Tracing / observability
Prompt playground & versioning
Dataset management
LLM-as-judge evals
Human annotation / review
Online (production) monitoring
A/B experiments
CI / regression testing
Cost & token tracking
Framework-agnostic SDK
Free tier
verified pending unknown (?) Click any cell to view its source or propose a value
Sources & verification 18

Every fact and feature listed for Arize Phoenix is verified against its own pages. Each alternative is sourced on its own page.

FAQ

Yes. Braintrust, HoneyHive and LangSmith have a free tier or are fully free. Free-tier limits in the comparison table are verified and dated.

Comet Opik, Langfuse and Braintrust — every license claim links its source.