The best free ai models & local execution tools (67)

All 67 ai models & local execution tools in the catalog across 8 sub-categories, ordered free tiers first.

67tools listed
64free or freemium
40with no paid plan at all
8sub-categories covered

Best Free AI Models & Local Execution Tools — straight answers

What are the best free ai models & local execution tools?

Flocci AI Tools lists 67 of them, ordered free tiers first: AnythingLLM, Arize Phoenix, Axolotl, DeepSeek Models and Design Arena. 64 of the 67 are free or have a free tier. Each entry states what it uniquely does rather than a score.

Which of these are completely free?

40 on this page are marked fully free with no paid plan at all: AnythingLLM, Arize Phoenix, Axolotl, DeepSeek Models and Design Arena. The remaining 27 are freemium, trial or paid, and each is labelled with its exact tier so nothing surprises you at sign-up.

How is this list ordered?

Free tools first, then freemium, then trial, then paid, and alphabetically inside each tier. Nobody pays to appear higher: Flocci AI Tools carries no ads, no sponsored slots and no paid listings, so the order reflects price and nothing else.

All 67 tools, free tiers first

Best Free AI Models & Local Execution Tools: pricing tier, category and the capability each tool is listed for.
ToolPricingTypeListed for
AnythingLLMFreeLocal LLM RunnersAll-in-one local RAG and agents app
Arize PhoenixFreeLLM Eval & ObservabilityFully open-source, self-hostable with zero setup via `uvx` (pip/conda also supported)
AxolotlFreeFine-Tuning & Training FrameworksConfig-file-driven fine-tuning workflow (YAML) covering LoRA, QLoRA, full fine-tuning, preference tuning and RL
DeepSeek ModelsFreeNotable Open ModelsOpen V3-series and R1 reasoning models
Design ArenaFreeLeaderboards & BenchmarksCrowdsourced Bradley-Terry (Elo-style) ranking of AI models specifically on design/aesthetic quality, not raw capability
Epoch AIFreeLeaderboards & BenchmarksIndependent, non-vendor research institute tracking AI compute, training cost and capability trends over time
EXAONE (LG AI Research)FreeNotable Open ModelsEXAONE 4.5: LG's first open-weight vision-language model for industrial use cases
Falcon (TII)FreeNotable Open ModelsFalcon H1 hybrid transformer-Mamba architecture for efficient long-context inference
Gemma (Google)FreeNotable Open ModelsGoogle's lightweight open models (Gemma 4)
GenesisFreeRobotics & Embodied AIUnified multi-physics simulation engine plus photorealistic renderer for robotics/embodied-AI training, fully open source
GPT-OSS (OpenAI)FreeNotable Open ModelsOpenAI's open-weight reasoning models
GPT4AllFreeLocal LLM RunnersPrivate local chat with your documents
Hunyuan (Tencent)FreeNotable Open ModelsOpen-weight large text MoE models (Hunyuan A13B/Hy3 series) with permissive commercial license
IBM GraniteFreeNotable Open ModelsApache 2.0 licensed enterprise-grade model family
JanFreeLocal LLM RunnersOpen-source, offline ChatGPT alternative
KoboldCppFreeLocal LLM RunnersShips as a single portable executable with a bundled web UI (KoboldAI Lite)
LeRobot (Hugging Face)FreeRobotics & Embodied AIOpen PyTorch library for training real-world robot control policies (ACT, Diffusion Policy, VLA models)
Llama (Meta)FreeNotable Open ModelsMeta's flagship open-weight family (Llama 4)
LLaMA-FactoryFreeFine-Tuning & Training FrameworksZero-code Web UI (LLaMA Board) for fine-tuning 100+ open models without writing training scripts
llama.cppFreeLocal LLM RunnersPure C/C++ inference engine with minimal dependencies, runs on CPU-only hardware
LM StudioFreeLocal LLM RunnersPolished desktop GUI for local models
LMArena (Arena.ai)FreeLeaderboards & BenchmarksCrowd-voted 'Battle Mode' blind head-to-head comparisons across text, image, and coding/agent models
MiniMax M2 (open weights)FreeNotable Open ModelsOpen-weight MoE text/reasoning model line distinct from MiniMax's Hailuo video product already listed
Mistral Open Models (Magistral, Devstral, Voxtral)FreeNotable Open ModelsDevstral: open coding-agent model tuned for SWE-bench-style tasks
MLX-LM (Apple)FreeLocal LLM RunnersPurpose-built for Apple Silicon's unified memory architecture, faster than llama.cpp for many models on Mac
Nemotron (NVIDIA)FreeNotable Open ModelsNVIDIA's open post-trained/distilled model family built on Llama and NVIDIA's own architectures
NVIDIA Isaac GR00TFreeRobotics & Embodied AIOpen, Apache-2.0-licensed vision-language-action foundation model for generalist humanoid robots (GR00T N1.7)
NVIDIA ParakeetFreeSpeech & Transcription EnginesTops the Hugging Face Open ASR Leaderboard with 6.05% average WER at 0.6B params
OllamaFreeLocal LLM RunnersThe default way to run LLMs locally
OLMo (Ai2)FreeNotable Open ModelsOnly major model family shipping full pretraining data (Dolma), code and intermediate checkpoints, not just weights
Open WebUIFreeLocal LLM RunnersSelf-hosted ChatGPT-style interface
OpenAI WhisperFreeSpeech & Transcription EnginesOpen-source speech-to-text, run locally free
Phi (Microsoft)FreeNotable Open ModelsMIT-licensed small language models optimized for on-device and edge deployment
PleiasFreeNotable Open ModelsFully open-weight, open-dataset small language models (350M-3B) explicitly built for EU AI Act compliance
Qwen (Alibaba)FreeNotable Open ModelsState-of-the-art open MoE models (Qwen3.x)
SGLangFreeLocal LLM RunnersRadixAttention for automatic prefix-cache reuse across requests
SWE-benchFreeLeaderboards & BenchmarksThe standard benchmark for evaluating autonomous coding agents on real GitHub issue resolution
UnslothFreeFine-Tuning & Training FrameworksFine-tunes LLMs, diffusion, TTS and embedding models 2x faster with ~70% less VRAM than standard Hugging Face training
vLLMFreeLocal LLM RunnersPagedAttention memory management for high-throughput multi-request GPU serving
WhisperXFreeSpeech & Transcription Engines70x real-time transcription speed via batched inference on top of Whisper
Artificial AnalysisFreemiumLeaderboards & BenchmarksIndependent cross-provider benchmarking of intelligence, speed and price for the same model across different hosting APIs
AssemblyAIFreemiumSpeech & Transcription EnginesSpeech-to-text with audio intelligence
BasetenFreemiumModel Hosting & Inference APIsBasic plan is pay-as-you-go with free starter credits for new accounts
BraintrustFreemiumLLM Eval & ObservabilityFree Starter plan ships model credits and scored evals with no card required
Cerebras InferenceFreemiumModel Hosting & Inference APIsRuns open models on Cerebras Wafer-Scale Engine hardware at dramatically higher tokens/sec than GPU-based APIs
Cloudflare Workers AIFreemiumModel Hosting & Inference APIs10,000 free 'Neurons' of inference per day on every account, resetting daily
DeepgramFreemiumSpeech & Transcription EnginesFast, accurate speech-to-text API
ElevenLabs ScribeFreemiumSpeech & Transcription EnginesScribe v2 Realtime transcribes in under 150ms latency
Fireworks AIFreemiumModel Hosting & Inference APIsHigh-performance open-model serving
GLM (Z.ai)FreemiumNotable Open ModelsGLM-4.5-Flash and GLM-4.7-Flash text models are free via API
Google AI StudioFreemiumModel Hosting & Inference APIsFree-tier API access to Gemini Flash-family models with generous token allowances, no credit card required to start
GroqFreemiumModel Hosting & Inference APIsUltra-fast inference on custom LPUs
Hugging FaceFreemiumModel Hosting & Inference APIsThe hub for open models and datasets
LangfuseFreemiumLLM Eval & ObservabilityFully open-source and self-hostable for free via Docker Compose/Kubernetes, not just a hosted SaaS
Mistral VoxtralFreemiumSpeech & Transcription EnginesApache 2.0 open-weight models (24B and 3B) downloadable from Hugging Face
MstyFreemiumLocal LLM RunnersNexus module: unified gateway for managing model connections, credentials and usage policy across providers
Novita AIFreemiumModel Hosting & Inference APIsSingle API surface for 200+ LLM/image/video/audio models plus on-demand GPU and bare-metal instances
OpenRouterFreemiumModel Hosting & Inference APIsOne API for 400+ models across providers
PromptfooFreemiumLLM Eval & ObservabilityOpen-source CLI/library for evaluating prompts, models and RAG pipelines side by side, runs in CI
ReplicateFreemiumModel Hosting & Inference APIsRun open models with one API call
SonioxFreemiumSpeech & Transcription EnginesReal-time speech-to-text priced well under typical competitor rates
SpeechmaticsFreemiumSpeech & Transcription EnginesFree starting credit with no card required, 56+ language support
Together AIFreemiumModel Hosting & Inference APIsFast, cheap open-model inference
W&B WeaveFreemiumLLM Eval & ObservabilityAgent-native tracing model with sessions, steps, tools and sub-agents as first-class concepts (not generic spans)
DeepInfraModel Hosting & Inference APIsVery low per-token pricing across a large catalog of open models (text, image, audio)
Nebius Token FactoryModel Hosting & Inference APIsEuropean GPU-cloud-backed inference API hosting major open models (Llama, Qwen, DeepSeek, etc.)
SambaNova CloudModel Hosting & Inference APIsHosts open models (DeepSeek, MiniMax, Llama) on SambaNova's own RDU chips for high-speed inference

AnythingLLM

Local LLM Runners
free
  • All-in-one local RAG and agents app
  • Chat with your docs privately
  • Works with local or cloud models
  • Free and open-source

Arize Phoenix

LLM Eval & Observability
free
  • Fully open-source, self-hostable with zero setup via `uvx` (pip/conda also supported)
  • Combines tracing, evals, datasets, experiments and prompt playground in one local-first tool
  • AI engineering agent (PXI) built in for automated troubleshooting of traces

Axolotl

Fine-Tuning & Training Frameworks
free
  • Config-file-driven fine-tuning workflow (YAML) covering LoRA, QLoRA, full fine-tuning, preference tuning and RL
  • Broad multi-model and multimodal training support with GPU-efficiency optimizations built in
  • Popular choice for reproducible post-training recipes shared across the open-model community

DeepSeek Models

Notable Open Models
free
  • Open V3-series and R1 reasoning models
  • Efficient MoE, low-cost inference
  • Strong math and code
  • Free open weights

Design Arena

Leaderboards & Benchmarks
free
  • Crowdsourced Bradley-Terry (Elo-style) ranking of AI models specifically on design/aesthetic quality, not raw capability
  • Separate leaderboards for Website, UI Component, Game Dev, Data Viz, 3D, Image, Video and Logo generation
  • Real-time rankings from user votes across 140+ countries rather than a fixed static benchmark

Epoch AI

Leaderboards & Benchmarks
free
  • Independent, non-vendor research institute tracking AI compute, training cost and capability trends over time
  • Maintains FrontierMath and the Epoch Capabilities Index, benchmarks designed to resist saturation/gaming
  • Interactive Data Explorers covering 3,200+ tracked models, data centers, chips and companies, free to browse

EXAONE (LG AI Research)

Notable Open Models
free
  • EXAONE 4.5: LG's first open-weight vision-language model for industrial use cases
  • 33B flagship plus FP8/AWQ quantized variants for efficient local deployment
  • Backed by LG's enterprise/manufacturing domain tuning

Falcon (TII)

Notable Open Models
free
  • Falcon H1 hybrid transformer-Mamba architecture for efficient long-context inference
  • Falcon Mamba: pure state-space model variant, no attention layers
  • Free commercial use under TII's permissive open license

Gemma (Google)

Notable Open Models
free
  • Google's lightweight open models (Gemma 4)
  • Runs in as little as 16GB RAM
  • On-device and multimodal variants
  • Free and open

Genesis

Robotics & Embodied AI
free
  • Unified multi-physics simulation engine plus photorealistic renderer for robotics/embodied-AI training, fully open source
  • 30,000+ GitHub stars, actively updated as recently as Aug 2026
  • Pythonic API designed to be dramatically faster than prior simulators for RL/robot-policy training

GPT-OSS (OpenAI)

Notable Open Models
free
  • OpenAI's open-weight reasoning models
  • Run locally or self-host
  • Configurable reasoning effort
  • Free, permissive license

GPT4All

Local LLM Runners
free
  • Private local chat with your documents
  • Runs on modest laptops
  • LocalDocs RAG built in
  • Free and open-source

Hunyuan (Tencent)

Notable Open Models
free
  • Open-weight large text MoE models (Hunyuan A13B/Hy3 series) with permissive commercial license
  • Hunyuan3D line for fast image-to-3D generation (under 1 second)
  • Hy-MT translation models covering dozens of language pairs

IBM Granite

Notable Open Models
free
  • Apache 2.0 licensed enterprise-grade model family
  • Hybrid Mamba/transformer architecture (Granite 4) for lower memory footprint
  • Built-in focus on governance, provenance and enterprise compliance

Jan

Local LLM Runners
free
  • Open-source, offline ChatGPT alternative
  • No-config, privacy-first GUI
  • Runs fully on your device
  • Free forever

KoboldCpp

Local LLM Runners
free
  • Ships as a single portable executable with a bundled web UI (KoboldAI Lite)
  • Generates text, image, video and speech from one local app
  • Runs on CPU or GPU with no installation required

LeRobot (Hugging Face)

Robotics & Embodied AI
free
  • Open PyTorch library for training real-world robot control policies (ACT, Diffusion Policy, VLA models)
  • Direct integration with Hugging Face Hub for sharing datasets and pretrained robot policies
  • Works with low-cost hobbyist robot arms (SO-100/101), not just simulation

Llama (Meta)

Notable Open Models
free
  • Meta's flagship open-weight family (Llama 4)
  • Multimodal and long-context variants
  • Run locally or via any host
  • Free, permissive license

LLaMA-Factory

Fine-Tuning & Training Frameworks
free
  • Zero-code Web UI (LLaMA Board) for fine-tuning 100+ open models without writing training scripts
  • Supports full-tuning, LoRA, 2/3/4/5/6/8-bit QLoRA, DPO, PPO, GaLore and PiSSA in one framework
  • Used internally by Amazon, NVIDIA and Aliyun for open-model post-training

llama.cpp

Local LLM Runners
free
  • Pure C/C++ inference engine with minimal dependencies, runs on CPU-only hardware
  • Powers the GGUF format used by Ollama, LM Studio and most local runners under the hood
  • Broad quantization support (2-bit to 8-bit) for running large models on consumer hardware

LM Studio

Local LLM Runners
free
  • Polished desktop GUI for local models
  • Discover, download and chat with GGUF models
  • Local server with OpenAI-compatible API
  • Free for personal use

LMArena (Arena.ai)

Leaderboards & Benchmarks
free
  • Crowd-voted 'Battle Mode' blind head-to-head comparisons across text, image, and coding/agent models
  • The most-cited human-preference leaderboard in the industry, referenced in nearly every model launch announcement
  • Rebranded from LMArena to Arena.ai in 2026, now also covering agent and design-to-code leaderboards

MiniMax M2 (open weights)

Notable Open Models
free
  • Open-weight MoE text/reasoning model line distinct from MiniMax's Hailuo video product already listed
  • Actively iterated through 2026 (M2.1, M2.5, M2.7 checkpoints)
  • Competitive on agentic coding and tool-use benchmarks among open Chinese models

MLX-LM (Apple)

Local LLM Runners
free
  • Purpose-built for Apple Silicon's unified memory architecture, faster than llama.cpp for many models on Mac
  • Built-in LoRA/QLoRA fine-tuning support, not just inference
  • Direct Hugging Face Hub integration for pulling and pushing quantized models

Nemotron (NVIDIA)

Notable Open Models
free
  • NVIDIA's open post-trained/distilled model family built on Llama and NVIDIA's own architectures
  • Includes top-ranked open embedding models (Nemotron Embed, Llama-Embed-Nemotron)
  • OpenReasoning-Nemotron: distilled reasoning models for local/edge inference

NVIDIA Isaac GR00T

Robotics & Embodied AI
free
  • Open, Apache-2.0-licensed vision-language-action foundation model for generalist humanoid robots (GR00T N1.7)
  • Opened to all developers worldwide Apr 2026 (previously partner-only), running on DGX Cloud
  • Full pipeline: pretrained model + simulation fine-tuning + deployment tools, not just a research checkpoint

NVIDIA Parakeet

Speech & Transcription Engines
free
  • Tops the Hugging Face Open ASR Leaderboard with 6.05% average WER at 0.6B params
  • CC-BY-4.0 license, free for commercial and non-commercial use
  • RTFx ~3,386 — extremely fast inference relative to accuracy

Ollama

Local LLM Runners
free
  • The default way to run LLMs locally
  • One command to run Llama, Qwen, DeepSeek…
  • OpenAI-compatible local API
  • Free and open-source

OLMo (Ai2)

Notable Open Models
free
  • Only major model family shipping full pretraining data (Dolma), code and intermediate checkpoints, not just weights
  • Olmo 3 ships Base, Think, and Instruct variants at 7B/32B
  • OlmoCore training framework and Open Instruct pipeline are also open source

Open WebUI

Local LLM Runners
free
  • Self-hosted ChatGPT-style interface
  • Works with Ollama and OpenAI-compatible APIs
  • RAG, tools and multi-user
  • Free and open-source

OpenAI Whisper

Speech & Transcription Engines
free
  • Open-source speech-to-text, run locally free
  • 99+ languages and translation
  • Robust to accents and noise
  • Powers countless apps

Phi (Microsoft)

Notable Open Models
free
  • MIT-licensed small language models optimized for on-device and edge deployment
  • Phi-4-mini and Phi-4-multimodal handle text, vision and audio in one small model
  • Distributed via Azure AI Foundry, Hugging Face, and Ollama

Pleias

Notable Open Models
free
  • Fully open-weight, open-dataset small language models (350M-3B) explicitly built for EU AI Act compliance
  • Trained exclusively on the 2-trillion-token open Common Corpus, not scraped/licensed text
  • Purpose-built RAG models (Pleias-RAG-350M/1B) with built-in citation tracking, Apache 2.0 licensed

Qwen (Alibaba)

Notable Open Models
free
  • State-of-the-art open MoE models (Qwen3.x)
  • Top-tier coding and multilingual skills
  • Runs on a single high-RAM Mac
  • Free open weights

SGLang

Local LLM Runners
free
  • RadixAttention for automatic prefix-cache reuse across requests
  • Day-0 support for newly released open models (e.g. Kimi K3) among the fastest of any runner
  • Scales from single GPU to distributed multi-node clusters with speculative decoding

SWE-bench

Leaderboards & Benchmarks
free
  • The standard benchmark for evaluating autonomous coding agents on real GitHub issue resolution
  • Multiple official variants (Verified, Multimodal, Multilingual, Lite) for different evaluation needs
  • Companion tooling (SWE-agent, mini-SWE-agent) for running the benchmark yourself, fully open

Unsloth

Fine-Tuning & Training Frameworks
free
  • Fine-tunes LLMs, diffusion, TTS and embedding models 2x faster with ~70% less VRAM than standard Hugging Face training
  • Free Google Colab notebooks let anyone fine-tune open models like Llama/Qwen/GLM on a free T4 GPU
  • Supports LoRA, QLoRA, full fine-tuning, GRPO and DPO reinforcement/preference tuning in one library

vLLM

Local LLM Runners
free
  • PagedAttention memory management for high-throughput multi-request GPU serving
  • Supports 200+ model architectures with an OpenAI-compatible API server
  • Originated at UC Berkeley Sky Computing Lab, now the de facto standard for self-hosted production LLM serving

WhisperX

Speech & Transcription Engines
free
  • 70x real-time transcription speed via batched inference on top of Whisper
  • Adds wav2vec2 forced alignment for accurate word-level timestamps Whisper lacks natively
  • Built-in speaker diarization and VAD preprocessing to cut hallucinations

Artificial Analysis

Leaderboards & Benchmarks
freemium
  • Independent cross-provider benchmarking of intelligence, speed and price for the same model across different hosting APIs
  • Artificial Analysis Intelligence Index aggregates multiple benchmarks into one comparable score
  • Covers coding agents, image/video, speech and vertical (finance/legal/health) evaluations, not just chat

AssemblyAI

Speech & Transcription Engines
freemium
  • Speech-to-text with audio intelligence
  • Summaries, topics and sentiment
  • Speaker diarization
  • Free tier for developers

Baseten

Model Hosting & Inference APIs
freemium
  • Basic plan is pay-as-you-go with free starter credits for new accounts
  • Deploys custom/fine-tuned models, not just a fixed model catalog, with per-second GPU billing
  • Hosts current open models (DeepSeek, GPT-OSS, Kimi K3) as ready-to-call APIs alongside custom deployment

Braintrust

LLM Eval & Observability
freemium
  • Free Starter plan ships model credits and scored evals with no card required
  • Unlimited users/projects/datasets/playgrounds even on the free tier
  • Built-in playground for side-by-side prompt/model comparison against production traces

Cerebras Inference

Model Hosting & Inference APIs
freemium
  • Runs open models on Cerebras Wafer-Scale Engine hardware at dramatically higher tokens/sec than GPU-based APIs
  • New accounts get free credits across all Cerebras-hosted models
  • Popular backend choice for latency-sensitive agentic and coding tools

Cloudflare Workers AI

Model Hosting & Inference APIs
freemium
  • 10,000 free 'Neurons' of inference per day on every account, resetting daily
  • Runs open models (Llama, Mistral, Whisper, etc.) on Cloudflare's global edge network for low-latency serverless inference
  • Bundled with Workers/Pages for building full serverless AI apps without managing GPUs

Deepgram

Speech & Transcription Engines
freemium
  • Fast, accurate speech-to-text API
  • Real-time streaming and diarization
  • Voice-agent building blocks
  • Free credits to start

ElevenLabs Scribe

Speech & Transcription Engines
freemium
  • Scribe v2 Realtime transcribes in under 150ms latency
  • Dynamic audio-event tagging (laughter, footsteps) beyond plain words
  • Keyterm prompting locks in up to 1,000 specified terms for accuracy

Fireworks AI

Model Hosting & Inference APIs
freemium
  • High-performance open-model serving
  • Fine-tuning and function calling
  • Vision and audio models
  • Free trial credits

GLM (Z.ai)

Notable Open Models
freemium
  • GLM-4.5-Flash and GLM-4.7-Flash text models are free via API
  • Open-weight GLM checkpoints on Hugging Face (MIT-style license for GLM-4.5-Air)
  • Strong agentic/coding benchmark results rivaling DeepSeek and Qwen

Google AI Studio

Model Hosting & Inference APIs
freemium
  • Free-tier API access to Gemini Flash-family models with generous token allowances, no credit card required to start
  • Browser-based playground for prompt design, multimodal testing, and one-click API key generation
  • Batch API offers 50% cost reduction once on the paid tier

Groq

Model Hosting & Inference APIs
freemium
  • Ultra-fast inference on custom LPUs
  • Sub-second responses for open models
  • OpenAI-compatible API
  • Generous free tier

Hugging Face

Model Hosting & Inference APIs
freemium
  • The hub for open models and datasets
  • Spaces to host and demo apps
  • Inference endpoints and providers
  • Free accounts and hosting

Langfuse

LLM Eval & Observability
freemium
  • Fully open-source and self-hostable for free via Docker Compose/Kubernetes, not just a hosted SaaS
  • Hobby cloud plan is free with no credit card, 50k observability units/month
  • Combines tracing, prompt management, evals and datasets in one open platform

Mistral Voxtral

Speech & Transcription Engines
freemium
  • Apache 2.0 open-weight models (24B and 3B) downloadable from Hugging Face
  • Built-in Q&A/summarization over audio, not just raw transcription
  • Function-calling directly from voice input for agent workflows

Msty

Local LLM Runners
freemium
  • Nexus module: unified gateway for managing model connections, credentials and usage policy across providers
  • Governance features (SSO, audit logs, zero telemetry) aimed at private/enterprise use
  • Combines local model chat, agents (Go), and knowledge packages (Stack) in one app

Novita AI

Model Hosting & Inference APIs
freemium
  • Single API surface for 200+ LLM/image/video/audio models plus on-demand GPU and bare-metal instances
  • Agent Sandbox for secure agent runtimes alongside model hosting
  • 'Free to start' onboarding credits across model and GPU products

OpenRouter

Model Hosting & Inference APIs
freemium
  • One API for 400+ models across providers
  • Automatic fallbacks and price routing
  • Some models free to use
  • Unified billing

Promptfoo

LLM Eval & Observability
freemium
  • Open-source CLI/library for evaluating prompts, models and RAG pipelines side by side, runs in CI
  • Automated red-teaming to surface prompt injection, jailbreak and data-leak vulnerabilities
  • 300,000+ user community with an enterprise tier for teams needing hosted guardrails

Replicate

Model Hosting & Inference APIs
freemium
  • Run open models with one API call
  • Thousands of community models
  • Deploy your own with Cog
  • Pay-per-second, free trial credits

Soniox

Speech & Transcription Engines
freemium
  • Real-time speech-to-text priced well under typical competitor rates
  • Combines transcription + translation in one streaming API call
  • Diarization and speech translation bundled at no extra cost

Speechmatics

Speech & Transcription Engines
freemium
  • Free starting credit with no card required, 56+ language support
  • Model-training discount program cuts costs 33% for data-sharing customers
  • On-premises/VPC deployment for privacy-sensitive enterprise use

Together AI

Model Hosting & Inference APIs
freemium
  • Fast, cheap open-model inference
  • 200+ models via one API
  • Fine-tuning and dedicated endpoints
  • Free starter credits

W&B Weave

LLM Eval & Observability
freemium
  • Agent-native tracing model with sessions, steps, tools and sub-agents as first-class concepts (not generic spans)
  • Pre-built safety/quality scorers for toxicity, bias, PII and hallucination detection out of the box
  • Built on Weights & Biases' existing ML-experiment infrastructure, useful for teams already on W&B

DeepInfra

Model Hosting & Inference APIs
  • Very low per-token pricing across a large catalog of open models (text, image, audio)
  • Flex scheduling tier at 0.8x price for non-production/batch workloads
  • No subscription; usage-tier system auto-scales billing thresholds

Nebius Token Factory

Model Hosting & Inference APIs
  • European GPU-cloud-backed inference API hosting major open models (Llama, Qwen, DeepSeek, etc.)
  • Rebranded in 2026 from 'Nebius AI Studio' to 'Nebius Token Factory'
  • Backed by Nebius's own large-scale GPU cloud infrastructure rather than reselling capacity

SambaNova Cloud

Model Hosting & Inference APIs
  • Hosts open models (DeepSeek, MiniMax, Llama) on SambaNova's own RDU chips for high-speed inference
  • Per-token pricing with no platform subscription required
  • Competes directly with Cerebras/Groq on raw tokens/sec for open-weight models

Related collections

All collections →