The best free LLM evaluation and observability tools (5)

Every option in the catalog for when you need to trace, score and debug what your AI app actually did — 5 LLM evaluation and observability tools, 5 of them free or freemium, with the pricing tier printed on every row.

5tools listed
5free or freemium
1with no paid plan at all
1sub-categories covered

Best Free LLM evaluation and observability tools — straight answers

What are the best free LLM evaluation and observability tools?

Flocci AI Tools lists 5 of them, ordered free tiers first: Arize Phoenix, Braintrust, Langfuse, Promptfoo and W&B Weave. Every one of them is free or has a free tier. Each entry states what it uniquely does rather than a score.

Which of these are completely free?

1 on this page are marked fully free with no paid plan at all: Arize Phoenix. The remaining 4 are freemium, trial or paid, and each is labelled with its exact tier so nothing surprises you at sign-up.

How is this list ordered?

Free tools first, then freemium, then trial, then paid, and alphabetically inside each tier. Nobody pays to appear higher: Flocci AI Tools carries no ads, no sponsored slots and no paid listings, so the order reflects price and nothing else.

All 5 tools, free tiers first

Browse LLM Eval & Observability →
Best Free LLM evaluation and observability tools: pricing tier, category and the capability each tool is listed for.
ToolPricingTypeListed for
Arize PhoenixFreeLLM Eval & ObservabilityFully open-source, self-hostable with zero setup via `uvx` (pip/conda also supported)
BraintrustFreemiumLLM Eval & ObservabilityFree Starter plan ships model credits and scored evals with no card required
LangfuseFreemiumLLM Eval & ObservabilityFully open-source and self-hostable for free via Docker Compose/Kubernetes, not just a hosted SaaS
PromptfooFreemiumLLM Eval & ObservabilityOpen-source CLI/library for evaluating prompts, models and RAG pipelines side by side, runs in CI
W&B WeaveFreemiumLLM Eval & ObservabilityAgent-native tracing model with sessions, steps, tools and sub-agents as first-class concepts (not generic spans)

Arize Phoenix

LLM Eval & Observability
free
  • Fully open-source, self-hostable with zero setup via `uvx` (pip/conda also supported)
  • Combines tracing, evals, datasets, experiments and prompt playground in one local-first tool
  • AI engineering agent (PXI) built in for automated troubleshooting of traces

Braintrust

LLM Eval & Observability
freemium
  • Free Starter plan ships model credits and scored evals with no card required
  • Unlimited users/projects/datasets/playgrounds even on the free tier
  • Built-in playground for side-by-side prompt/model comparison against production traces

Langfuse

LLM Eval & Observability
freemium
  • Fully open-source and self-hostable for free via Docker Compose/Kubernetes, not just a hosted SaaS
  • Hobby cloud plan is free with no credit card, 50k observability units/month
  • Combines tracing, prompt management, evals and datasets in one open platform

Promptfoo

LLM Eval & Observability
freemium
  • Open-source CLI/library for evaluating prompts, models and RAG pipelines side by side, runs in CI
  • Automated red-teaming to surface prompt injection, jailbreak and data-leak vulnerabilities
  • 300,000+ user community with an enterprise tier for teams needing hosted guardrails

W&B Weave

LLM Eval & Observability
freemium
  • Agent-native tracing model with sessions, steps, tools and sub-agents as first-class concepts (not generic spans)
  • Pre-built safety/quality scorers for toxicity, bias, PII and hallucination detection out of the box
  • Built on Weights & Biases' existing ML-experiment infrastructure, useful for teams already on W&B

Related collections

All collections →