AI Models & Local Execution (5)

All 5 LLM evaluation and observability tools in the Flocci catalog, 5 of them free or freemium. Pricing tier and real capabilities on every entry — no ads, no sponsored placement.

Langfuse

LLM Eval & Observability
freemium
  • Fully open-source and self-hostable for free via Docker Compose/Kubernetes, not just a hosted SaaS
  • Hobby cloud plan is free with no credit card, 50k observability units/month
  • Combines tracing, prompt management, evals and datasets in one open platform

Braintrust

LLM Eval & Observability
freemium
  • Free Starter plan ships model credits and scored evals with no card required
  • Unlimited users/projects/datasets/playgrounds even on the free tier
  • Built-in playground for side-by-side prompt/model comparison against production traces

Promptfoo

LLM Eval & Observability
freemium
  • Open-source CLI/library for evaluating prompts, models and RAG pipelines side by side, runs in CI
  • Automated red-teaming to surface prompt injection, jailbreak and data-leak vulnerabilities
  • 300,000+ user community with an enterprise tier for teams needing hosted guardrails

W&B Weave

LLM Eval & Observability
freemium
  • Agent-native tracing model with sessions, steps, tools and sub-agents as first-class concepts (not generic spans)
  • Pre-built safety/quality scorers for toxicity, bias, PII and hallucination detection out of the box
  • Built on Weights & Biases' existing ML-experiment infrastructure, useful for teams already on W&B

Arize Phoenix

LLM Eval & Observability
free
  • Fully open-source, self-hostable with zero setup via `uvx` (pip/conda also supported)
  • Combines tracing, evals, datasets, experiments and prompt playground in one local-first tool
  • AI engineering agent (PXI) built in for automated troubleshooting of traces

LLM Eval & Observability — straight answers

What are the best free LLM evaluation and observability tools?

Flocci AI Tools lists 5 LLM evaluation and observability tools, ordered free tiers first: Arize Phoenix, Braintrust, Langfuse, Promptfoo and W&B Weave. All of them are free or have a free tier. Every entry states what it uniquely does, not a score.

Are there completely free LLM evaluation and observability tools?

Yes — 1 of these are marked fully free with no paid plan attached: Arize Phoenix. The rest are freemium, trial or paid, and each row states which.

What should I look for in LLM evaluation and observability tools?

Start with what you need to do: trace, score and debug what your AI app actually did. Then check the pricing tier, because "free" here means no paid plan at all while "freemium" means a free tier under paid plans. Every tool on this page lists both.

See the full list →

Looking for a ranked shortlist instead? See the best free LLM evaluation and observability tools.