Promptfoo

Open-source CLI/library for evaluating prompts, models and RAG pipelines side by side, runs in CI

What Promptfoo does

  • Open-source CLI/library for evaluating prompts, models and RAG pipelines side by side, runs in CI
  • Automated red-teaming to surface prompt injection, jailbreak and data-leak vulnerabilities
  • 300,000+ user community with an enterprise tier for teams needing hosted guardrails

Promptfoo — straight answers

What is Promptfoo?

Promptfoo is listed under LLM Eval & Observability, in the AI Models & Local Execution category on Flocci AI Tools. Open-source CLI/library for evaluating prompts, models and RAG pipelines side by side, runs in CI. It is freemium — a usable free tier with paid plans above it, and it lives at promptfoo.dev.

Is Promptfoo free?

Promptfoo is freemium: there is a free tier you can work in and paid plans above it. It is one of 391 freemium tools in the catalog, all labelled the same way so you are never surprised by a paywall.

What can Promptfoo do?

Promptfoo does 3 things the catalog singles out: Open-source cli/library for evaluating prompts, models and rag pipelines side by side, runs in ci; automated red-teaming to surface prompt injection, jailbreak and data-leak vulnerabilities; 300,000+ user community with an enterprise tier for teams needing hosted guardrails.

What is the best free alternative to Promptfoo?

Arize Phoenix is the closest free alternative: it sits in the same LLM Eval & Observability sub-category and is free. Braintrust, Langfuse and W&B Weave also start free. The full list is on the alternatives page.

See the full list →

Promptfoo alternatives

Compare all alternatives →

Langfuse

LLM Eval & Observability
freemium
  • Fully open-source and self-hostable for free via Docker Compose/Kubernetes, not just a hosted SaaS
  • Hobby cloud plan is free with no credit card, 50k observability units/month
  • Combines tracing, prompt management, evals and datasets in one open platform

Braintrust

LLM Eval & Observability
freemium
  • Free Starter plan ships model credits and scored evals with no card required
  • Unlimited users/projects/datasets/playgrounds even on the free tier
  • Built-in playground for side-by-side prompt/model comparison against production traces

W&B Weave

LLM Eval & Observability
freemium
  • Agent-native tracing model with sessions, steps, tools and sub-agents as first-class concepts (not generic spans)
  • Pre-built safety/quality scorers for toxicity, bias, PII and hallucination detection out of the box
  • Built on Weights & Biases' existing ML-experiment infrastructure, useful for teams already on W&B

Arize Phoenix

LLM Eval & Observability
free
  • Fully open-source, self-hostable with zero setup via `uvx` (pip/conda also supported)
  • Combines tracing, evals, datasets, experiments and prompt playground in one local-first tool
  • AI engineering agent (PXI) built in for automated troubleshooting of traces

Ollama

Local LLM Runners
free
  • The default way to run LLMs locally
  • One command to run Llama, Qwen, DeepSeek…
  • OpenAI-compatible local API
  • Free and open-source

LM Studio

Local LLM Runners
free
  • Polished desktop GUI for local models
  • Discover, download and chat with GGUF models
  • Local server with OpenAI-compatible API
  • Free for personal use

Head-to-head comparisons