64 of these 67 tools are free or have a free tier. Sub-categories: Local LLM Runners, Model Hosting & Inference APIs, Notable Open Models, Speech & Transcription Engines, Fine-Tuning & Training Frameworks, LLM Eval & Observability, Leaderboards & Benchmarks, Robotics & Embodied AI. Open AI Models & Local Execution in the app.
OllamaFree The default way to run LLMs locally · One command to run Llama, Qwen, DeepSeek… · OpenAI-compatible local API · Free and open-source https://ollama.com/ · alternatives
LM StudioFree Polished desktop GUI for local models · Discover, download and chat with GGUF models · Local server with OpenAI-compatible API · Free for personal use https://lmstudio.ai/ · alternatives
JanFree Open-source, offline ChatGPT alternative · No-config, privacy-first GUI · Runs fully on your device · Free forever https://jan.ai/ · alternatives
GPT4AllFree Private local chat with your documents · Runs on modest laptops · LocalDocs RAG built in · Free and open-source https://www.nomic.ai/gpt4all · alternatives
Open WebUIFree Self-hosted ChatGPT-style interface · Works with Ollama and OpenAI-compatible APIs · RAG, tools and multi-user · Free and open-source https://openwebui.com/ · alternatives
AnythingLLMFree All-in-one local RAG and agents app · Chat with your docs privately · Works with local or cloud models · Free and open-source https://anythingllm.com/ · alternatives
llama.cppFree Pure C/C++ inference engine with minimal dependencies, runs on CPU-only hardware · Powers the GGUF format used by Ollama, LM Studio and most local runners under the hood · Broad quantization support (2-bit to 8-bit) for running large models on consumer hardware https://github.com/ggml-org/llama.cpp · alternatives
vLLMFree PagedAttention memory management for high-throughput multi-request GPU serving · Supports 200+ model architectures with an OpenAI-compatible API server · Originated at UC Berkeley Sky Computing Lab, now the de facto standard for self-hosted production LLM serving https://github.com/vllm-project/vllm · alternatives
SGLangFree RadixAttention for automatic prefix-cache reuse across requests · Day-0 support for newly released open models (e.g. Kimi K3) among the fastest of any runner · Scales from single GPU to distributed multi-node clusters with speculative decoding https://github.com/sgl-project/sglang · alternatives
KoboldCppFree Ships as a single portable executable with a bundled web UI (KoboldAI Lite) · Generates text, image, video and speech from one local app · Runs on CPU or GPU with no installation required https://github.com/LostRuins/koboldcpp · alternatives
MLX-LM (Apple)Free Purpose-built for Apple Silicon's unified memory architecture, faster than llama.cpp for many models on Mac · Built-in LoRA/QLoRA fine-tuning support, not just inference · Direct Hugging Face Hub integration for pulling and pushing quantized models https://github.com/ml-explore/mlx-lm · alternatives
MstyFreemium Nexus module: unified gateway for managing model connections, credentials and usage policy across providers · Governance features (SSO, audit logs, zero telemetry) aimed at private/enterprise use · Combines local model chat, agents (Go), and knowledge packages (Stack) in one app https://msty.ai/ · alternatives
Hugging FaceFreemium The hub for open models and datasets · Spaces to host and demo apps · Inference endpoints and providers · Free accounts and hosting https://huggingface.co/ · alternatives
ReplicateFreemium Run open models with one API call · Thousands of community models · Deploy your own with Cog · Pay-per-second, free trial credits https://replicate.com/ · alternatives
GroqFreemium Ultra-fast inference on custom LPUs · Sub-second responses for open models · OpenAI-compatible API · Generous free tier https://groq.com/ · alternatives
Together AIFreemium Fast, cheap open-model inference · 200+ models via one API · Fine-tuning and dedicated endpoints · Free starter credits https://www.together.ai/ · alternatives
OpenRouterFreemium One API for 400+ models across providers · Automatic fallbacks and price routing · Some models free to use · Unified billing https://openrouter.ai/ · alternatives
Fireworks AIFreemium High-performance open-model serving · Fine-tuning and function calling · Vision and audio models · Free trial credits https://fireworks.ai/ · alternatives
Cerebras InferenceFreemium Runs open models on Cerebras Wafer-Scale Engine hardware at dramatically higher tokens/sec than GPU-based APIs · New accounts get free credits across all Cerebras-hosted models · Popular backend choice for latency-sensitive agentic and coding tools https://www.cerebras.ai/inference · alternatives
SambaNova CloudPaid Hosts open models (DeepSeek, MiniMax, Llama) on SambaNova's own RDU chips for high-speed inference · Per-token pricing with no platform subscription required · Competes directly with Cerebras/Groq on raw tokens/sec for open-weight models https://cloud.sambanova.ai/ · alternatives
DeepInfraPaid Very low per-token pricing across a large catalog of open models (text, image, audio) · Flex scheduling tier at 0.8x price for non-production/batch workloads · No subscription; usage-tier system auto-scales billing thresholds https://deepinfra.com/ · alternatives
Novita AIFreemium Single API surface for 200+ LLM/image/video/audio models plus on-demand GPU and bare-metal instances · Agent Sandbox for secure agent runtimes alongside model hosting · 'Free to start' onboarding credits across model and GPU products https://novita.ai/ · alternatives
Nebius Token FactoryPaid European GPU-cloud-backed inference API hosting major open models (Llama, Qwen, DeepSeek, etc.) · Rebranded in 2026 from 'Nebius AI Studio' to 'Nebius Token Factory' · Backed by Nebius's own large-scale GPU cloud infrastructure rather than reselling capacity https://tokenfactory.nebius.com/ · alternatives
Cloudflare Workers AIFreemium 10,000 free 'Neurons' of inference per day on every account, resetting daily · Runs open models (Llama, Mistral, Whisper, etc.) on Cloudflare's global edge network for low-latency serverless inference · Bundled with Workers/Pages for building full serverless AI apps without managing GPUs https://developers.cloudflare.com/workers-ai/ · alternatives
BasetenFreemium Basic plan is pay-as-you-go with free starter credits for new accounts · Deploys custom/fine-tuned models, not just a fixed model catalog, with per-second GPU billing · Hosts current open models (DeepSeek, GPT-OSS, Kimi K3) as ready-to-call APIs alongside custom deployment https://www.baseten.co/ · alternatives
Google AI StudioFreemium Free-tier API access to Gemini Flash-family models with generous token allowances, no credit card required to start · Browser-based playground for prompt design, multimodal testing, and one-click API key generation · Batch API offers 50% cost reduction once on the paid tier https://aistudio.google.com/ · alternatives
Llama (Meta)Free Meta's flagship open-weight family (Llama 4) · Multimodal and long-context variants · Run locally or via any host · Free, permissive license https://www.llama.com/ · alternatives
Qwen (Alibaba)Free State-of-the-art open MoE models (Qwen3.x) · Top-tier coding and multilingual skills · Runs on a single high-RAM Mac · Free open weights https://qwen.ai/ · alternatives
DeepSeek ModelsFree Open V3-series and R1 reasoning models · Efficient MoE, low-cost inference · Strong math and code · Free open weights https://www.deepseek.com/ · alternatives
Gemma (Google)Free Google's lightweight open models (Gemma 4) · Runs in as little as 16GB RAM · On-device and multimodal variants · Free and open https://ai.google.dev/gemma · alternatives
PleiasFree Fully open-weight, open-dataset small language models (350M-3B) explicitly built for EU AI Act compliance · Trained exclusively on the 2-trillion-token open Common Corpus, not scraped/licensed text · Purpose-built RAG models (Pleias-RAG-350M/1B) with built-in citation tracking, Apache 2.0 licensed https://pleias.fr/ · alternatives
Mistral Open Models (Magistral, Devstral, Voxtral)Free Devstral: open coding-agent model tuned for SWE-bench-style tasks · Magistral: open reasoning model with visible chain-of-thought · Voxtral: open-weights speech model for transcription and voice agents https://mistral.ai/news · alternatives
GLM (Z.ai)Freemium GLM-4.5-Flash and GLM-4.7-Flash text models are free via API · Open-weight GLM checkpoints on Hugging Face (MIT-style license for GLM-4.5-Air) · Strong agentic/coding benchmark results rivaling DeepSeek and Qwen https://z.ai/ · alternatives
IBM GraniteFree Apache 2.0 licensed enterprise-grade model family · Hybrid Mamba/transformer architecture (Granite 4) for lower memory footprint · Built-in focus on governance, provenance and enterprise compliance https://www.ibm.com/granite · alternatives
OLMo (Ai2)Free Only major model family shipping full pretraining data (Dolma), code and intermediate checkpoints, not just weights · Olmo 3 ships Base, Think, and Instruct variants at 7B/32B · OlmoCore training framework and Open Instruct pipeline are also open source https://allenai.org/olmo · alternatives
Hunyuan (Tencent)Free Open-weight large text MoE models (Hunyuan A13B/Hy3 series) with permissive commercial license · Hunyuan3D line for fast image-to-3D generation (under 1 second) · Hy-MT translation models covering dozens of language pairs https://huggingface.co/tencent · alternatives
Falcon (TII)Free Falcon H1 hybrid transformer-Mamba architecture for efficient long-context inference · Falcon Mamba: pure state-space model variant, no attention layers · Free commercial use under TII's permissive open license https://falconllm.tii.ae/ · alternatives
EXAONE (LG AI Research)Free EXAONE 4.5: LG's first open-weight vision-language model for industrial use cases · 33B flagship plus FP8/AWQ quantized variants for efficient local deployment · Backed by LG's enterprise/manufacturing domain tuning https://www.lgresearch.ai/exaone · alternatives
MiniMax M2 (open weights)Free Open-weight MoE text/reasoning model line distinct from MiniMax's Hailuo video product already listed · Actively iterated through 2026 (M2.1, M2.5, M2.7 checkpoints) · Competitive on agentic coding and tool-use benchmarks among open Chinese models https://huggingface.co/MiniMaxAI · alternatives
Phi (Microsoft)Free MIT-licensed small language models optimized for on-device and edge deployment · Phi-4-mini and Phi-4-multimodal handle text, vision and audio in one small model · Distributed via Azure AI Foundry, Hugging Face, and Ollama https://azure.microsoft.com/en-us/products/phi · alternatives
Nemotron (NVIDIA)Free NVIDIA's open post-trained/distilled model family built on Llama and NVIDIA's own architectures · Includes top-ranked open embedding models (Nemotron Embed, Llama-Embed-Nemotron) · OpenReasoning-Nemotron: distilled reasoning models for local/edge inference https://huggingface.co/nvidia · alternatives
OpenAI WhisperFree Open-source speech-to-text, run locally free · 99+ languages and translation · Robust to accents and noise · Powers countless apps https://github.com/openai/whisper · alternatives
DeepgramFreemium Fast, accurate speech-to-text API · Real-time streaming and diarization · Voice-agent building blocks · Free credits to start https://deepgram.com/ · alternatives
AssemblyAIFreemium Speech-to-text with audio intelligence · Summaries, topics and sentiment · Speaker diarization · Free tier for developers https://www.assemblyai.com/ · alternatives
NVIDIA ParakeetFree Tops the Hugging Face Open ASR Leaderboard with 6.05% average WER at 0.6B params · CC-BY-4.0 license, free for commercial and non-commercial use · RTFx ~3,386 — extremely fast inference relative to accuracy https://huggingface.co/nvidia/parakeet-tdt-0.6b-v2 · alternatives
Mistral VoxtralFreemium Apache 2.0 open-weight models (24B and 3B) downloadable from Hugging Face · Built-in Q&A/summarization over audio, not just raw transcription · Function-calling directly from voice input for agent workflows https://mistral.ai/news/voxtral · alternatives
SpeechmaticsFreemium Free starting credit with no card required, 56+ language support · Model-training discount program cuts costs 33% for data-sharing customers · On-premises/VPC deployment for privacy-sensitive enterprise use https://www.speechmatics.com/ · alternatives
SonioxFreemium Real-time speech-to-text priced well under typical competitor rates · Combines transcription + translation in one streaming API call · Diarization and speech translation bundled at no extra cost https://soniox.com/ · alternatives
ElevenLabs ScribeFreemium Scribe v2 Realtime transcribes in under 150ms latency · Dynamic audio-event tagging (laughter, footsteps) beyond plain words · Keyterm prompting locks in up to 1,000 specified terms for accuracy https://elevenlabs.io/speech-to-text · alternatives
WhisperXFree 70x real-time transcription speed via batched inference on top of Whisper · Adds wav2vec2 forced alignment for accurate word-level timestamps Whisper lacks natively · Built-in speaker diarization and VAD preprocessing to cut hallucinations https://github.com/m-bain/whisperX · alternatives
UnslothFree Fine-tunes LLMs, diffusion, TTS and embedding models 2x faster with ~70% less VRAM than standard Hugging Face training · Free Google Colab notebooks let anyone fine-tune open models like Llama/Qwen/GLM on a free T4 GPU · Supports LoRA, QLoRA, full fine-tuning, GRPO and DPO reinforcement/preference tuning in one library https://unsloth.ai/ · alternatives
AxolotlFree Config-file-driven fine-tuning workflow (YAML) covering LoRA, QLoRA, full fine-tuning, preference tuning and RL · Broad multi-model and multimodal training support with GPU-efficiency optimizations built in · Popular choice for reproducible post-training recipes shared across the open-model community https://axolotl.ai/ · alternatives
LLaMA-FactoryFree Zero-code Web UI (LLaMA Board) for fine-tuning 100+ open models without writing training scripts · Supports full-tuning, LoRA, 2/3/4/5/6/8-bit QLoRA, DPO, PPO, GaLore and PiSSA in one framework · Used internally by Amazon, NVIDIA and Aliyun for open-model post-training https://github.com/hiyouga/LLaMA-Factory · alternatives
LangfuseFreemium Fully open-source and self-hostable for free via Docker Compose/Kubernetes, not just a hosted SaaS · Hobby cloud plan is free with no credit card, 50k observability units/month · Combines tracing, prompt management, evals and datasets in one open platform https://langfuse.com/ · alternatives
BraintrustFreemium Free Starter plan ships model credits and scored evals with no card required · Unlimited users/projects/datasets/playgrounds even on the free tier · Built-in playground for side-by-side prompt/model comparison against production traces https://www.braintrust.dev/ · alternatives
PromptfooFreemium Open-source CLI/library for evaluating prompts, models and RAG pipelines side by side, runs in CI · Automated red-teaming to surface prompt injection, jailbreak and data-leak vulnerabilities · 300,000+ user community with an enterprise tier for teams needing hosted guardrails https://www.promptfoo.dev/ · alternatives
W&B WeaveFreemium Agent-native tracing model with sessions, steps, tools and sub-agents as first-class concepts (not generic spans) · Pre-built safety/quality scorers for toxicity, bias, PII and hallucination detection out of the box · Built on Weights & Biases' existing ML-experiment infrastructure, useful for teams already on W&B https://wandb.ai/site/weave/ · alternatives
Arize PhoenixFree Fully open-source, self-hostable with zero setup via `uvx` (pip/conda also supported) · Combines tracing, evals, datasets, experiments and prompt playground in one local-first tool · AI engineering agent (PXI) built in for automated troubleshooting of traces https://phoenix.arize.com/ · alternatives
LMArena (Arena.ai)Free Crowd-voted 'Battle Mode' blind head-to-head comparisons across text, image, and coding/agent models · The most-cited human-preference leaderboard in the industry, referenced in nearly every model launch announcement · Rebranded from LMArena to Arena.ai in 2026, now also covering agent and design-to-code leaderboards https://arena.ai/ · alternatives
Artificial AnalysisFreemium Independent cross-provider benchmarking of intelligence, speed and price for the same model across different hosting APIs · Artificial Analysis Intelligence Index aggregates multiple benchmarks into one comparable score · Covers coding agents, image/video, speech and vertical (finance/legal/health) evaluations, not just chat https://artificialanalysis.ai/ · alternatives
SWE-benchFree The standard benchmark for evaluating autonomous coding agents on real GitHub issue resolution · Multiple official variants (Verified, Multimodal, Multilingual, Lite) for different evaluation needs · Companion tooling (SWE-agent, mini-SWE-agent) for running the benchmark yourself, fully open https://www.swebench.com/ · alternatives
Design ArenaFree Crowdsourced Bradley-Terry (Elo-style) ranking of AI models specifically on design/aesthetic quality, not raw capability · Separate leaderboards for Website, UI Component, Game Dev, Data Viz, 3D, Image, Video and Logo generation · Real-time rankings from user votes across 140+ countries rather than a fixed static benchmark https://www.designarena.ai/ · alternatives
Epoch AIFree Independent, non-vendor research institute tracking AI compute, training cost and capability trends over time · Maintains FrontierMath and the Epoch Capabilities Index, benchmarks designed to resist saturation/gaming · Interactive Data Explorers covering 3,200+ tracked models, data centers, chips and companies, free to browse https://epoch.ai/ · alternatives
LeRobot (Hugging Face)Free Open PyTorch library for training real-world robot control policies (ACT, Diffusion Policy, VLA models) · Direct integration with Hugging Face Hub for sharing datasets and pretrained robot policies · Works with low-cost hobbyist robot arms (SO-100/101), not just simulation https://github.com/huggingface/lerobot · alternatives
NVIDIA Isaac GR00TFree Open, Apache-2.0-licensed vision-language-action foundation model for generalist humanoid robots (GR00T N1.7) · Opened to all developers worldwide Apr 2026 (previously partner-only), running on DGX Cloud · Full pipeline: pretrained model + simulation fine-tuning + deployment tools, not just a research checkpoint https://developer.nvidia.com/isaac/gr00t · alternatives
GenesisFree Unified multi-physics simulation engine plus photorealistic renderer for robotics/embodied-AI training, fully open source · 30,000+ GitHub stars, actively updated as recently as Aug 2026 · Pythonic API designed to be dramatically faster than prior simulators for RL/robot-policy training https://github.com/Genesis-Embodied-AI/genesis-world · alternatives
This page is a plain-HTML mirror of the same catalog data rendered by the Flocci AI Tools app at aitools.flocci.in. It exists because the app is a single-page application; the content is identical. Generated 2026-08-20.