AI Models & Local Execution (12)

All 12 local LLM runners in the Flocci catalog, 12 of them free or freemium. Pricing tier and real capabilities on every entry — no ads, no sponsored placement.

Ollama

Local LLM Runners
free
  • The default way to run LLMs locally
  • One command to run Llama, Qwen, DeepSeek…
  • OpenAI-compatible local API
  • Free and open-source

LM Studio

Local LLM Runners
free
  • Polished desktop GUI for local models
  • Discover, download and chat with GGUF models
  • Local server with OpenAI-compatible API
  • Free for personal use

Jan

Local LLM Runners
free
  • Open-source, offline ChatGPT alternative
  • No-config, privacy-first GUI
  • Runs fully on your device
  • Free forever

GPT4All

Local LLM Runners
free
  • Private local chat with your documents
  • Runs on modest laptops
  • LocalDocs RAG built in
  • Free and open-source

Open WebUI

Local LLM Runners
free
  • Self-hosted ChatGPT-style interface
  • Works with Ollama and OpenAI-compatible APIs
  • RAG, tools and multi-user
  • Free and open-source

AnythingLLM

Local LLM Runners
free
  • All-in-one local RAG and agents app
  • Chat with your docs privately
  • Works with local or cloud models
  • Free and open-source

llama.cpp

Local LLM Runners
free
  • Pure C/C++ inference engine with minimal dependencies, runs on CPU-only hardware
  • Powers the GGUF format used by Ollama, LM Studio and most local runners under the hood
  • Broad quantization support (2-bit to 8-bit) for running large models on consumer hardware

vLLM

Local LLM Runners
free
  • PagedAttention memory management for high-throughput multi-request GPU serving
  • Supports 200+ model architectures with an OpenAI-compatible API server
  • Originated at UC Berkeley Sky Computing Lab, now the de facto standard for self-hosted production LLM serving

SGLang

Local LLM Runners
free
  • RadixAttention for automatic prefix-cache reuse across requests
  • Day-0 support for newly released open models (e.g. Kimi K3) among the fastest of any runner
  • Scales from single GPU to distributed multi-node clusters with speculative decoding

KoboldCpp

Local LLM Runners
free
  • Ships as a single portable executable with a bundled web UI (KoboldAI Lite)
  • Generates text, image, video and speech from one local app
  • Runs on CPU or GPU with no installation required

MLX-LM (Apple)

Local LLM Runners
free
  • Purpose-built for Apple Silicon's unified memory architecture, faster than llama.cpp for many models on Mac
  • Built-in LoRA/QLoRA fine-tuning support, not just inference
  • Direct Hugging Face Hub integration for pulling and pushing quantized models

Msty

Local LLM Runners
freemium
  • Nexus module: unified gateway for managing model connections, credentials and usage policy across providers
  • Governance features (SSO, audit logs, zero telemetry) aimed at private/enterprise use
  • Combines local model chat, agents (Go), and knowledge packages (Stack) in one app

Local LLM Runners — straight answers

What are the best free local LLM runners?

Flocci AI Tools lists 12 local LLM runners, ordered free tiers first: AnythingLLM, GPT4All, Jan, KoboldCpp and llama.cpp. All of them are free or have a free tier. Every entry states what it uniquely does, not a score.

Are there completely free local LLM runners?

Yes — 11 of these are marked fully free with no paid plan attached: AnythingLLM, GPT4All, Jan, KoboldCpp and llama.cpp. The rest are freemium, trial or paid, and each row states which.

What should I look for in local LLM runners?

Start with what you need to do: run a model on your own machine, fully offline. Then check the pricing tier, because "free" here means no paid plan at all while "freemium" means a free tier under paid plans. Every tool on this page lists both.

See the full list →

Looking for a ranked shortlist instead? See the best free local LLM runners.