Which LLM runs on
your machine?
Tell us what's under the hood. We'll tell you what runs, how fast, and how to install it — step by step, in plain English.
The market in brief
All briefs →The BestLLMfor kits — the reference guide by use case.
One kit per use case: the full guide, ready-to-paste configs, lifetime online access.
239 models, every angle.
The catalog's most-tracked families — one flagship model per author. Filter and jump straight into the full catalog.
Find your LLM by your need.
Don’t want to use the configurator? Our themed rankings pick the best self-hostable models by use case and hardware.
The BestLLMfor documentation.
92+ hands-on guides, tested on Windows, macOS and Linux. From your first install to advanced RAG and fine-tuning techniques.
Featured
— our essential guidesIndependent 2026 ranking of the best LLM for coding. SWE-bench scores, VRAM needs, cost per million tokens, and a verdict for cloud and local setups.
Best LLM for RAG in 2026: Claude 4.5 Sonnet for accuracy, Qwen3 32B for local, Voyage-3 for embeddings. Tested verdicts, costs, and stacks.
How to run a local LLM in 2026: pick the right hardware, install Ollama, choose your first model, and understand quantization. A no-fluff beginner guide.
Best LLMs to Run on the RTX 4060 Ti (8GB vs 16GB)
RTX 4060 Ti 8GB vs 16GB for local LLMs: real bandwidth numbers, which models fit each config, Ollama setup steps, and our 2026 buying verdict.
Which LLM Runs Best on the MacBook Air M4 (16GB, 24GB, or 32GB)?
MacBook Air M4 for local LLMs: which RAM tier (16GB, 24GB, 32GB) fits 8B-27B models, real specs, setup steps, and how it stacks up to the MacBook Pro M4 and Mac mini.
Best LLM for MacBook Air M3: 8GB, 16GB, or 24GB
Which local LLMs actually run well on a MacBook Air M3 with 8GB, 16GB, or 24GB of RAM? Real RAM budgets, Ollama/MLX setup, Metal tuning, and the M4 upgrade math.
Best LLM for RTX 5080 (16GB): What It Can Actually Run in 2026
RTX 5080 for local LLMs in 2026: full specs, Ollama setup, Flash Attention 3 gains, and how it stacks up against the RTX 5070 Ti and RTX 4080 Super.
Which LLM Runs Best on the Radeon RX 7900 XTX (24GB)?
Radeon RX 7900 XTX for local LLMs: real specs, ROCm 6.x setup on Ubuntu, which model sizes fit in 24GB, and how it stacks up against the RTX 4090 and 3090.
Which LLM Should You Run on a Mac Studio (M2/M3/M4 Ultra)?
See which LLMs run on Mac Studio M2, M3, and M4 Ultra configs from 96GB to 512GB, with real speeds, 2026 pricing, and setup steps for local inference.
Which LLM Should You Run on an RTX 4080 or 4080 Super?
RTX 4080 and 4080 Super buyer's guide for local LLMs in 2026: real VRAM limits, which model sizes fit, used pricing, and whether to upgrade to a 5070 Ti instead.
Best LLM for MacBook Pro M3 Pro and M3 Max (18GB-128GB)
M3 Pro or M3 Max for local LLMs? We compare memory bandwidth, RAM tiers, and real tok/s across every MacBook Pro M3 configuration, from 18GB to 128GB.
Best LLM for MacBook Pro M4 Pro and M4 Max (24GB–128GB)
Which local LLM runs best on a MacBook Pro M4 Pro or M4 Max? We break down bandwidth, RAM ceilings, MLX vs llama.cpp, and Mac Studio pricing for 2026.
Best Local LLM for RTX 4070 / 4070 Super / 4070 Ti (12GB)
RTX 4070, Super, and Ti all ship with 12GB of VRAM. Here's which local LLMs actually fit, which quant levels to use, and which of the three is worth buying in 2026.
Best Local LLMs for the RTX 5070 Ti (16GB) in 2026
RTX 5070 Ti (16GB) local LLM guide: specs, what models fit at Q4/Q8, Ollama setup, Blackwell optimizations, and how it stacks up against the 5080 and a used 4090.
What LLM Can You Run on 8GB VRAM in 2026?
See exactly which LLMs fit in 8GB of VRAM in 2026, from Qwen 3.5 9B to Granite 4.2 8B, plus the quantization settings that make local LLMs run on budget GPUs.
Best LLMs for Mac Mini M4 and M4 Pro (16–64GB) in 2026
Mac mini M4 and M4 Pro for local AI in 2026: which LLMs fit 16GB–64GB of unified memory, real tok/s numbers, Ollama setup, and the best value config.
Best Local LLM for RTX 3090 (24GB) in 2026
RTX 3090 24GB in 2026: the cheapest path to 27B-35B local LLMs. What tokens/sec to expect, how loud it runs, and when to buy a 4090 or 5070 Ti instead.
What LLMs Can You Run on the RTX 3060 12GB in 2026?
RTX 3060 12GB local LLM guide for 2026: which models fit in 12GB, real tokens/sec, Ollama setup, KV cache tuning, and how it stacks up against the RTX 4060 and 5060.
How to Install text-generation-webui with CUDA on Windows
Step-by-step guide to install text-generation-webui with CUDA on Windows, verify GPU acceleration, pick the right loader, and fix common CUDA errors.
How to Self-Host TabbyAPI as an OpenAI-Compatible Endpoint
Self-host TabbyAPI as an OpenAI-compatible endpoint. Docker and manual install, config.yml, EXL2/EXL3 models, GPU requirements, and tuning for RTX 3090/4090/5090.
How to Use LlamaIndex with a Local Ollama LLM
Step-by-step guide to run LlamaIndex with a local Ollama LLM. Install commands, a working RAG script, hardware specs, model benchmarks, and a clear verdict.
How to Wire LangChain to a Local LLM — Production-Ready
Wire LangChain to a local LLM the production way: Ollama vs vLLM, timeouts, streaming, structured output, RAG, plus 2026 hardware benchmarks and cost math.
How to Use LiteLLM Router with Ollama — OpenAI-Drop-In
Configure LiteLLM Router as an OpenAI-compatible proxy for Ollama in minutes. Load balancing, fallbacks, config.yaml, and 2026 latency benchmarks.
How to Add a Local LLM to Claude Desktop via MCP
Step-by-step guide to bridging Ollama or LM Studio into Claude Desktop with an MCP server, so a local LLM handles private tasks alongside Claude in 2026.
How to Use a Local Ollama LLM as Cursor's Backend
Use a local Ollama LLM as Cursor's backend the right way: expose an OpenAI-compatible endpoint, fix the context window, and pick the best local coding model.
How to Wire Continue.dev to a Local Ollama LLM in VS Code
Wire Continue.dev to a local Ollama LLM in VS Code: exact models, config.yaml, hardware specs, and benchmarks for private, $0/month AI coding.
How to Use Aider with Ollama as a Local Copilot
Set up Aider with Ollama for a private, zero-cost local copilot. Install steps, the ollama_chat/ prefix, OLLAMA_API_BASE, context tuning, and the best coder models.
How to Self-Host OpenWebUI in Docker — 5-Minute Setup
Self-host Open WebUI in Docker in under 5 minutes. Copy-paste run command, Compose file, persistence fix, Ollama hookup, and production hardening.
How to Build llama.cpp with Metal on Mac M-Series
Build llama.cpp with Metal on Mac M-series in under 5 minutes. CMake steps, verification, M1–M4 benchmarks, tuning flags, and fixes for common errors.
How to Run Llama on Apple Silicon with MLX — Native Performance
Run Llama on Apple Silicon with MLX for native performance. Install steps, memory requirements, MLX vs llama.cpp benchmarks, quantization picks, and an OpenAI API.
How to Install Ollama with ROCm on AMD GPU
Step-by-step guide to installing Ollama with ROCm on AMD GPUs. Covers RX 6000/7000/9000 support, HSA overrides, benchmarks, and troubleshooting for 2026.
How to Run vLLM in Docker with NVIDIA CUDA — 10-Minute Setup
Run vLLM in Docker with NVIDIA CUDA in under 10 minutes. Step-by-step setup, GPU passthrough, model VRAM tables, benchmarks, and troubleshooting fixes.
How to Build llama.cpp from Source with CUDA — 2026 Guide
Step-by-step 2026 guide to build llama.cpp from source with CUDA: exact CMake flags, GPU arch codes, toolkit versions, benchmarks, and fixes for common errors.
Reviews & tests of AI tools.
When local isn’t enough, which paid services are worth your money? We reviewed 6 brands. Every review cites the free local alternative first.
Pick your path.
Three common starting points — jump straight to the one that matches you.
New to local LLMs?
Start here: what "local" means, what hardware you need, and your first model in 10 minutes.
Coding & dev work
Ranked local models for autocomplete, agents, and full coding sessions — matched to your GPU.
Replace ChatGPT at work
Keep sensitive data in-house. What actually works as a private, self-hosted swap-in.
Built around your decision, not vendor benchmarks.
Four practical tools that answer the questions you actually have when picking an LLM.
Hardware-matched rankings
Best local LLM for RTX 4090, RTX 5090, Mac M4 Max, Snapdragon X — cut through the noise with rankings that respect your VRAM, memory, and target speed.
Cost ROI: self-hosted vs API
Sliders for your monthly token volume, electricity cost, GPU amortization. Real break-even point against GPT-5, Claude, Gemini, DeepSeek — updated pricing.
Public API & MCP server
178 JSON endpoints under CC BY 4.0, free to use in your own tools. Official MCP server on GitHub for ChatGPT, Claude Desktop, and Cursor.
Independent benchmark pipeline
Continuous benchmarking against published model versions and quantizations. No press-kit numbers, no marketing decks — just tokens/sec backed by our open data API.
Independent. Skin in the game.
BestLLMfor is built and operated by Mohamed Meguedmi — one engineer, a continuous benchmark pipeline, a public data API and an open-source MCP server.
No VC, no SEO farm. One engineer obsessed with tracking every model worth running, and publishing what the numbers say — transparently.
| Models tracked | 239+ (daily) |
| Quants tested | Q4 · Q5 · Q8 · FP16 |
| Data API | 178 JSON · CC BY 4.0 |
| MCP server | Public · open source |
| Methodology | see how → |
Your GPU cheat sheet,
then a hands-on series.
Your VRAM cheat sheet now, then a short starter series (about nine emails over ten days), then only occasional updates. One-click unsubscribe.
Want the deep dive instead? See the Local Copilot Kit →