AI Models & Local Execution (67)

All 67 ai models & local execution tools in the Flocci catalog, 64 of them free or freemium. Pricing tier and real capabilities on every entry — no ads, no sponsored placement.

Ollama

Local LLM Runners
free
  • The default way to run LLMs locally
  • One command to run Llama, Qwen, DeepSeek…
  • OpenAI-compatible local API
  • Free and open-source

LM Studio

Local LLM Runners
free
  • Polished desktop GUI for local models
  • Discover, download and chat with GGUF models
  • Local server with OpenAI-compatible API
  • Free for personal use

Jan

Local LLM Runners
free
  • Open-source, offline ChatGPT alternative
  • No-config, privacy-first GUI
  • Runs fully on your device
  • Free forever

GPT4All

Local LLM Runners
free
  • Private local chat with your documents
  • Runs on modest laptops
  • LocalDocs RAG built in
  • Free and open-source

Open WebUI

Local LLM Runners
free
  • Self-hosted ChatGPT-style interface
  • Works with Ollama and OpenAI-compatible APIs
  • RAG, tools and multi-user
  • Free and open-source

AnythingLLM

Local LLM Runners
free
  • All-in-one local RAG and agents app
  • Chat with your docs privately
  • Works with local or cloud models
  • Free and open-source

llama.cpp

Local LLM Runners
free
  • Pure C/C++ inference engine with minimal dependencies, runs on CPU-only hardware
  • Powers the GGUF format used by Ollama, LM Studio and most local runners under the hood
  • Broad quantization support (2-bit to 8-bit) for running large models on consumer hardware

vLLM

Local LLM Runners
free
  • PagedAttention memory management for high-throughput multi-request GPU serving
  • Supports 200+ model architectures with an OpenAI-compatible API server
  • Originated at UC Berkeley Sky Computing Lab, now the de facto standard for self-hosted production LLM serving

SGLang

Local LLM Runners
free
  • RadixAttention for automatic prefix-cache reuse across requests
  • Day-0 support for newly released open models (e.g. Kimi K3) among the fastest of any runner
  • Scales from single GPU to distributed multi-node clusters with speculative decoding

KoboldCpp

Local LLM Runners
free
  • Ships as a single portable executable with a bundled web UI (KoboldAI Lite)
  • Generates text, image, video and speech from one local app
  • Runs on CPU or GPU with no installation required

MLX-LM (Apple)

Local LLM Runners
free
  • Purpose-built for Apple Silicon's unified memory architecture, faster than llama.cpp for many models on Mac
  • Built-in LoRA/QLoRA fine-tuning support, not just inference
  • Direct Hugging Face Hub integration for pulling and pushing quantized models

Msty

Local LLM Runners
freemium
  • Nexus module: unified gateway for managing model connections, credentials and usage policy across providers
  • Governance features (SSO, audit logs, zero telemetry) aimed at private/enterprise use
  • Combines local model chat, agents (Go), and knowledge packages (Stack) in one app

Hugging Face

Model Hosting & Inference APIs
freemium
  • The hub for open models and datasets
  • Spaces to host and demo apps
  • Inference endpoints and providers
  • Free accounts and hosting

Replicate

Model Hosting & Inference APIs
freemium
  • Run open models with one API call
  • Thousands of community models
  • Deploy your own with Cog
  • Pay-per-second, free trial credits

Groq

Model Hosting & Inference APIs
freemium
  • Ultra-fast inference on custom LPUs
  • Sub-second responses for open models
  • OpenAI-compatible API
  • Generous free tier

Together AI

Model Hosting & Inference APIs
freemium
  • Fast, cheap open-model inference
  • 200+ models via one API
  • Fine-tuning and dedicated endpoints
  • Free starter credits

OpenRouter

Model Hosting & Inference APIs
freemium
  • One API for 400+ models across providers
  • Automatic fallbacks and price routing
  • Some models free to use
  • Unified billing

Fireworks AI

Model Hosting & Inference APIs
freemium
  • High-performance open-model serving
  • Fine-tuning and function calling
  • Vision and audio models
  • Free trial credits

Cerebras Inference

Model Hosting & Inference APIs
freemium
  • Runs open models on Cerebras Wafer-Scale Engine hardware at dramatically higher tokens/sec than GPU-based APIs
  • New accounts get free credits across all Cerebras-hosted models
  • Popular backend choice for latency-sensitive agentic and coding tools

SambaNova Cloud

Model Hosting & Inference APIs
  • Hosts open models (DeepSeek, MiniMax, Llama) on SambaNova's own RDU chips for high-speed inference
  • Per-token pricing with no platform subscription required
  • Competes directly with Cerebras/Groq on raw tokens/sec for open-weight models

DeepInfra

Model Hosting & Inference APIs
  • Very low per-token pricing across a large catalog of open models (text, image, audio)
  • Flex scheduling tier at 0.8x price for non-production/batch workloads
  • No subscription; usage-tier system auto-scales billing thresholds

Novita AI

Model Hosting & Inference APIs
freemium
  • Single API surface for 200+ LLM/image/video/audio models plus on-demand GPU and bare-metal instances
  • Agent Sandbox for secure agent runtimes alongside model hosting
  • 'Free to start' onboarding credits across model and GPU products

Nebius Token Factory

Model Hosting & Inference APIs
  • European GPU-cloud-backed inference API hosting major open models (Llama, Qwen, DeepSeek, etc.)
  • Rebranded in 2026 from 'Nebius AI Studio' to 'Nebius Token Factory'
  • Backed by Nebius's own large-scale GPU cloud infrastructure rather than reselling capacity

Cloudflare Workers AI

Model Hosting & Inference APIs
freemium
  • 10,000 free 'Neurons' of inference per day on every account, resetting daily
  • Runs open models (Llama, Mistral, Whisper, etc.) on Cloudflare's global edge network for low-latency serverless inference
  • Bundled with Workers/Pages for building full serverless AI apps without managing GPUs

Baseten

Model Hosting & Inference APIs
freemium
  • Basic plan is pay-as-you-go with free starter credits for new accounts
  • Deploys custom/fine-tuned models, not just a fixed model catalog, with per-second GPU billing
  • Hosts current open models (DeepSeek, GPT-OSS, Kimi K3) as ready-to-call APIs alongside custom deployment

Google AI Studio

Model Hosting & Inference APIs
freemium
  • Free-tier API access to Gemini Flash-family models with generous token allowances, no credit card required to start
  • Browser-based playground for prompt design, multimodal testing, and one-click API key generation
  • Batch API offers 50% cost reduction once on the paid tier

Llama (Meta)

Notable Open Models
free
  • Meta's flagship open-weight family (Llama 4)
  • Multimodal and long-context variants
  • Run locally or via any host
  • Free, permissive license

Qwen (Alibaba)

Notable Open Models
free
  • State-of-the-art open MoE models (Qwen3.x)
  • Top-tier coding and multilingual skills
  • Runs on a single high-RAM Mac
  • Free open weights

DeepSeek Models

Notable Open Models
free
  • Open V3-series and R1 reasoning models
  • Efficient MoE, low-cost inference
  • Strong math and code
  • Free open weights

Gemma (Google)

Notable Open Models
free
  • Google's lightweight open models (Gemma 4)
  • Runs in as little as 16GB RAM
  • On-device and multimodal variants
  • Free and open

GPT-OSS (OpenAI)

Notable Open Models
free
  • OpenAI's open-weight reasoning models
  • Run locally or self-host
  • Configurable reasoning effort
  • Free, permissive license

Pleias

Notable Open Models
free
  • Fully open-weight, open-dataset small language models (350M-3B) explicitly built for EU AI Act compliance
  • Trained exclusively on the 2-trillion-token open Common Corpus, not scraped/licensed text
  • Purpose-built RAG models (Pleias-RAG-350M/1B) with built-in citation tracking, Apache 2.0 licensed

GLM (Z.ai)

Notable Open Models
freemium
  • GLM-4.5-Flash and GLM-4.7-Flash text models are free via API
  • Open-weight GLM checkpoints on Hugging Face (MIT-style license for GLM-4.5-Air)
  • Strong agentic/coding benchmark results rivaling DeepSeek and Qwen

IBM Granite

Notable Open Models
free
  • Apache 2.0 licensed enterprise-grade model family
  • Hybrid Mamba/transformer architecture (Granite 4) for lower memory footprint
  • Built-in focus on governance, provenance and enterprise compliance

OLMo (Ai2)

Notable Open Models
free
  • Only major model family shipping full pretraining data (Dolma), code and intermediate checkpoints, not just weights
  • Olmo 3 ships Base, Think, and Instruct variants at 7B/32B
  • OlmoCore training framework and Open Instruct pipeline are also open source

Hunyuan (Tencent)

Notable Open Models
free
  • Open-weight large text MoE models (Hunyuan A13B/Hy3 series) with permissive commercial license
  • Hunyuan3D line for fast image-to-3D generation (under 1 second)
  • Hy-MT translation models covering dozens of language pairs

Falcon (TII)

Notable Open Models
free
  • Falcon H1 hybrid transformer-Mamba architecture for efficient long-context inference
  • Falcon Mamba: pure state-space model variant, no attention layers
  • Free commercial use under TII's permissive open license

EXAONE (LG AI Research)

Notable Open Models
free
  • EXAONE 4.5: LG's first open-weight vision-language model for industrial use cases
  • 33B flagship plus FP8/AWQ quantized variants for efficient local deployment
  • Backed by LG's enterprise/manufacturing domain tuning

MiniMax M2 (open weights)

Notable Open Models
free
  • Open-weight MoE text/reasoning model line distinct from MiniMax's Hailuo video product already listed
  • Actively iterated through 2026 (M2.1, M2.5, M2.7 checkpoints)
  • Competitive on agentic coding and tool-use benchmarks among open Chinese models

Phi (Microsoft)

Notable Open Models
free
  • MIT-licensed small language models optimized for on-device and edge deployment
  • Phi-4-mini and Phi-4-multimodal handle text, vision and audio in one small model
  • Distributed via Azure AI Foundry, Hugging Face, and Ollama

Nemotron (NVIDIA)

Notable Open Models
free
  • NVIDIA's open post-trained/distilled model family built on Llama and NVIDIA's own architectures
  • Includes top-ranked open embedding models (Nemotron Embed, Llama-Embed-Nemotron)
  • OpenReasoning-Nemotron: distilled reasoning models for local/edge inference

OpenAI Whisper

Speech & Transcription Engines
free
  • Open-source speech-to-text, run locally free
  • 99+ languages and translation
  • Robust to accents and noise
  • Powers countless apps

Deepgram

Speech & Transcription Engines
freemium
  • Fast, accurate speech-to-text API
  • Real-time streaming and diarization
  • Voice-agent building blocks
  • Free credits to start

AssemblyAI

Speech & Transcription Engines
freemium
  • Speech-to-text with audio intelligence
  • Summaries, topics and sentiment
  • Speaker diarization
  • Free tier for developers

NVIDIA Parakeet

Speech & Transcription Engines
free
  • Tops the Hugging Face Open ASR Leaderboard with 6.05% average WER at 0.6B params
  • CC-BY-4.0 license, free for commercial and non-commercial use
  • RTFx ~3,386 — extremely fast inference relative to accuracy

Mistral Voxtral

Speech & Transcription Engines
freemium
  • Apache 2.0 open-weight models (24B and 3B) downloadable from Hugging Face
  • Built-in Q&A/summarization over audio, not just raw transcription
  • Function-calling directly from voice input for agent workflows

Speechmatics

Speech & Transcription Engines
freemium
  • Free starting credit with no card required, 56+ language support
  • Model-training discount program cuts costs 33% for data-sharing customers
  • On-premises/VPC deployment for privacy-sensitive enterprise use

Soniox

Speech & Transcription Engines
freemium
  • Real-time speech-to-text priced well under typical competitor rates
  • Combines transcription + translation in one streaming API call
  • Diarization and speech translation bundled at no extra cost

ElevenLabs Scribe

Speech & Transcription Engines
freemium
  • Scribe v2 Realtime transcribes in under 150ms latency
  • Dynamic audio-event tagging (laughter, footsteps) beyond plain words
  • Keyterm prompting locks in up to 1,000 specified terms for accuracy

WhisperX

Speech & Transcription Engines
free
  • 70x real-time transcription speed via batched inference on top of Whisper
  • Adds wav2vec2 forced alignment for accurate word-level timestamps Whisper lacks natively
  • Built-in speaker diarization and VAD preprocessing to cut hallucinations

Unsloth

Fine-Tuning & Training Frameworks
free
  • Fine-tunes LLMs, diffusion, TTS and embedding models 2x faster with ~70% less VRAM than standard Hugging Face training
  • Free Google Colab notebooks let anyone fine-tune open models like Llama/Qwen/GLM on a free T4 GPU
  • Supports LoRA, QLoRA, full fine-tuning, GRPO and DPO reinforcement/preference tuning in one library

Axolotl

Fine-Tuning & Training Frameworks
free
  • Config-file-driven fine-tuning workflow (YAML) covering LoRA, QLoRA, full fine-tuning, preference tuning and RL
  • Broad multi-model and multimodal training support with GPU-efficiency optimizations built in
  • Popular choice for reproducible post-training recipes shared across the open-model community

LLaMA-Factory

Fine-Tuning & Training Frameworks
free
  • Zero-code Web UI (LLaMA Board) for fine-tuning 100+ open models without writing training scripts
  • Supports full-tuning, LoRA, 2/3/4/5/6/8-bit QLoRA, DPO, PPO, GaLore and PiSSA in one framework
  • Used internally by Amazon, NVIDIA and Aliyun for open-model post-training

Langfuse

LLM Eval & Observability
freemium
  • Fully open-source and self-hostable for free via Docker Compose/Kubernetes, not just a hosted SaaS
  • Hobby cloud plan is free with no credit card, 50k observability units/month
  • Combines tracing, prompt management, evals and datasets in one open platform

Braintrust

LLM Eval & Observability
freemium
  • Free Starter plan ships model credits and scored evals with no card required
  • Unlimited users/projects/datasets/playgrounds even on the free tier
  • Built-in playground for side-by-side prompt/model comparison against production traces

Promptfoo

LLM Eval & Observability
freemium
  • Open-source CLI/library for evaluating prompts, models and RAG pipelines side by side, runs in CI
  • Automated red-teaming to surface prompt injection, jailbreak and data-leak vulnerabilities
  • 300,000+ user community with an enterprise tier for teams needing hosted guardrails

W&B Weave

LLM Eval & Observability
freemium
  • Agent-native tracing model with sessions, steps, tools and sub-agents as first-class concepts (not generic spans)
  • Pre-built safety/quality scorers for toxicity, bias, PII and hallucination detection out of the box
  • Built on Weights & Biases' existing ML-experiment infrastructure, useful for teams already on W&B

Arize Phoenix

LLM Eval & Observability
free
  • Fully open-source, self-hostable with zero setup via `uvx` (pip/conda also supported)
  • Combines tracing, evals, datasets, experiments and prompt playground in one local-first tool
  • AI engineering agent (PXI) built in for automated troubleshooting of traces

LMArena (Arena.ai)

Leaderboards & Benchmarks
free
  • Crowd-voted 'Battle Mode' blind head-to-head comparisons across text, image, and coding/agent models
  • The most-cited human-preference leaderboard in the industry, referenced in nearly every model launch announcement
  • Rebranded from LMArena to Arena.ai in 2026, now also covering agent and design-to-code leaderboards

Artificial Analysis

Leaderboards & Benchmarks
freemium
  • Independent cross-provider benchmarking of intelligence, speed and price for the same model across different hosting APIs
  • Artificial Analysis Intelligence Index aggregates multiple benchmarks into one comparable score
  • Covers coding agents, image/video, speech and vertical (finance/legal/health) evaluations, not just chat

SWE-bench

Leaderboards & Benchmarks
free
  • The standard benchmark for evaluating autonomous coding agents on real GitHub issue resolution
  • Multiple official variants (Verified, Multimodal, Multilingual, Lite) for different evaluation needs
  • Companion tooling (SWE-agent, mini-SWE-agent) for running the benchmark yourself, fully open

Design Arena

Leaderboards & Benchmarks
free
  • Crowdsourced Bradley-Terry (Elo-style) ranking of AI models specifically on design/aesthetic quality, not raw capability
  • Separate leaderboards for Website, UI Component, Game Dev, Data Viz, 3D, Image, Video and Logo generation
  • Real-time rankings from user votes across 140+ countries rather than a fixed static benchmark

Epoch AI

Leaderboards & Benchmarks
free
  • Independent, non-vendor research institute tracking AI compute, training cost and capability trends over time
  • Maintains FrontierMath and the Epoch Capabilities Index, benchmarks designed to resist saturation/gaming
  • Interactive Data Explorers covering 3,200+ tracked models, data centers, chips and companies, free to browse

LeRobot (Hugging Face)

Robotics & Embodied AI
free
  • Open PyTorch library for training real-world robot control policies (ACT, Diffusion Policy, VLA models)
  • Direct integration with Hugging Face Hub for sharing datasets and pretrained robot policies
  • Works with low-cost hobbyist robot arms (SO-100/101), not just simulation

NVIDIA Isaac GR00T

Robotics & Embodied AI
free
  • Open, Apache-2.0-licensed vision-language-action foundation model for generalist humanoid robots (GR00T N1.7)
  • Opened to all developers worldwide Apr 2026 (previously partner-only), running on DGX Cloud
  • Full pipeline: pretrained model + simulation fine-tuning + deployment tools, not just a research checkpoint

Genesis

Robotics & Embodied AI
free
  • Unified multi-physics simulation engine plus photorealistic renderer for robotics/embodied-AI training, fully open source
  • 30,000+ GitHub stars, actively updated as recently as Aug 2026
  • Pythonic API designed to be dramatically faster than prior simulators for RL/robot-policy training

AI Models & Local Execution — straight answers

What ai models & local execution tools does Flocci list?

Flocci AI Tools lists 67 ai models & local execution tools, ordered free tiers first: AnythingLLM, Arize Phoenix, Axolotl, DeepSeek Models and Design Arena. 64 of the 67 are free or have a free tier. Every entry states what it uniquely does, not a score.

Which ai models & local execution tools are completely free?

Yes — 40 of these are marked fully free with no paid plan attached: AnythingLLM, Arize Phoenix, Axolotl, DeepSeek Models and Design Arena. The rest are freemium, trial or paid, and each row states which.

How is AI Models & Local Execution organised?

Into 8 sub-categories — Local LLM Runners, Model Hosting & Inference APIs, Notable Open Models, Speech & Transcription Engines, Fine-Tuning & Training Frameworks, LLM Eval & Observability, Leaderboards & Benchmarks, Robotics & Embodied AI — holding 67 tools between them. Each sub-category has its own page with the same pricing labels and capability lists.