BestLLMfor Your hardware. Your LLM. Your call.
◆ The kits◆ Kits APIOpen data Find my LLM
All models · 100% local, 100% private

Which LLM runs on
your machine?

Tell us what's under the hood. We'll tell you what runs, how fast, and how to install it — step by step, in plain English.

~/configurator —— loading…
Your data stays in your browser. Nothing is sent.
Free
No signup
Open data
CC BY 4.0
Independent
No tracker
Real benchmarks
Daily updates
The BestLLMfor kits — the reference guide by use case. $24 each, or $49 for all, lifetime. See the kits →
The catalog

239 models, every angle.

The catalog's most-tracked families — one flagship model per author. Filter and jump straight into the full catalog.

397B · Apache 2.0
Qwen 3.5 397B-A17B
240 GB Q4 · 255k ctx
31B · Gemma
Gemma 4 31B
18 GB Q4 · 250k ctx
561B · NVIDIA Open Model License
Nemotron 3 Ultra Base (BF16)
325 GB Q4 · 125k ctx
1700B · MIT
DeepSeek V4 Pro 0813 1.7T
986 GB Q4 · 1024k ctx
32B · Apache 2.0
Granite 4.0 H-Small 32B-A9B
19 GB Q4 · 125k ctx
675B · Apache 2.0
Mistral Large 3 675B
405 GB Q4 · 250k ctx
405B · Llama 3.1 Community
Llama 3.1 405B Instruct
240 GB Q4 · 125k ctx
72B · Apache 2.0
Molmo 72B
42 GB Q4 · 4k ctx
24B · LFM Open License v1.0
LFM2 24B
14 GB Q4 · 32k ctx
14B · MIT
Phi-4 Reasoning 14B
9 GB Q4 · 32k ctx
35B · CC-BY-NC 4.0
Aya 23 35B
20 GB Q4 · 8k ctx
753B · MIT
GLM 5.2 753B-A40B
437 GB Q4 · 976k ctx
2800B · Kimi License
Kimi K3
1624 GB Q4 · 976k ctx
8B · MiniCPM Model License
MiniCPM-o 2.6 8B
5.5 GB Q4 · 31k ctx
10B · TII Falcon-LLM License 2.0
Falcon 3 10B Instruct
6 GB Q4 · 31k ctx
406B · Tencent Hunyuan License
Hunyuan Large 2.0
245 GB Q4 · 256k ctx
104B · CC-BY-NC 4.0
Command R+ 104B (08-2024)
60 GB Q4 · 125k ctx
3B · Apache 2.0
SmolLM3 3B
2 GB Q4 · 125k ctx
118B · OpenMDW 1.1
Laguna S 2.1
68 GB Q4 · 256k ctx
1020B · MIT
MiMo V2.5 Pro
595 GB Q4 · 976k ctx
34B · Apache 2.0
Yi 1.5 34B Chat
20 GB Q4 · 4k ctx
1000B · MIT
Ling 2.6 1T
580 GB Q4 · 256k ctx
40B · Apache 2.0
Salamandra 40B Instruct
24 GB Q4 · 8k ctx
300B · Apache 2.0
ERNIE 4.5 300B-A47B
180 GB Q4 · 128k ctx
Browse all 239 models →
Learning center

The BestLLMfor documentation.

92+ hands-on guides, tested on Windows, macOS and Linux. From your first install to advanced RAG and fine-tuning techniques.

◆ The BestLLMfor kits — the reference guide by use case · $24 the kit · $49 all, lifetime →
⌘ K
92 results
Intermediate 6 min

Best LLMs to Run on the RTX 4060 Ti (8GB vs 16GB)

RTX 4060 Ti 8GB vs 16GB for local LLMs: real bandwidth numbers, which models fit each config, Ollama setup steps, and our 2026 buying verdict.

HardwareRead →
Intermediate 10 min

Which LLM Runs Best on the MacBook Air M4 (16GB, 24GB, or 32GB)?

MacBook Air M4 for local LLMs: which RAM tier (16GB, 24GB, 32GB) fits 8B-27B models, real specs, setup steps, and how it stacks up to the MacBook Pro M4 and Mac mini.

HardwareRead →
Intermediate 8 min

Best LLM for MacBook Air M3: 8GB, 16GB, or 24GB

Which local LLMs actually run well on a MacBook Air M3 with 8GB, 16GB, or 24GB of RAM? Real RAM budgets, Ollama/MLX setup, Metal tuning, and the M4 upgrade math.

HardwareRead →
Intermediate 6 min

Best LLM for RTX 5080 (16GB): What It Can Actually Run in 2026

RTX 5080 for local LLMs in 2026: full specs, Ollama setup, Flash Attention 3 gains, and how it stacks up against the RTX 5070 Ti and RTX 4080 Super.

HardwareRead →
Intermediate 7 min

Which LLM Runs Best on the Radeon RX 7900 XTX (24GB)?

Radeon RX 7900 XTX for local LLMs: real specs, ROCm 6.x setup on Ubuntu, which model sizes fit in 24GB, and how it stacks up against the RTX 4090 and 3090.

GuidesRead →
Intermediate 9 min

Which LLM Should You Run on a Mac Studio (M2/M3/M4 Ultra)?

See which LLMs run on Mac Studio M2, M3, and M4 Ultra configs from 96GB to 512GB, with real speeds, 2026 pricing, and setup steps for local inference.

HardwareRead →
Intermediate 8 min

Which LLM Should You Run on an RTX 4080 or 4080 Super?

RTX 4080 and 4080 Super buyer's guide for local LLMs in 2026: real VRAM limits, which model sizes fit, used pricing, and whether to upgrade to a 5070 Ti instead.

HardwareRead →
Intermediate 8 min

Best LLM for MacBook Pro M3 Pro and M3 Max (18GB-128GB)

M3 Pro or M3 Max for local LLMs? We compare memory bandwidth, RAM tiers, and real tok/s across every MacBook Pro M3 configuration, from 18GB to 128GB.

HardwareRead →
Intermediate 8 min

Best LLM for MacBook Pro M4 Pro and M4 Max (24GB–128GB)

Which local LLM runs best on a MacBook Pro M4 Pro or M4 Max? We break down bandwidth, RAM ceilings, MLX vs llama.cpp, and Mac Studio pricing for 2026.

HardwareRead →
Intermediate 8 min

Best Local LLM for RTX 4070 / 4070 Super / 4070 Ti (12GB)

RTX 4070, Super, and Ti all ship with 12GB of VRAM. Here's which local LLMs actually fit, which quant levels to use, and which of the three is worth buying in 2026.

RankingsRead →
Intermediate 6 min

Best Local LLMs for the RTX 5070 Ti (16GB) in 2026

RTX 5070 Ti (16GB) local LLM guide: specs, what models fit at Q4/Q8, Ollama setup, Blackwell optimizations, and how it stacks up against the 5080 and a used 4090.

HardwareRead →
Intermediate 7 min

What LLM Can You Run on 8GB VRAM in 2026?

See exactly which LLMs fit in 8GB of VRAM in 2026, from Qwen 3.5 9B to Granite 4.2 8B, plus the quantization settings that make local LLMs run on budget GPUs.

HardwareRead →
Intermediate 7 min

Best LLMs for Mac Mini M4 and M4 Pro (16–64GB) in 2026

Mac mini M4 and M4 Pro for local AI in 2026: which LLMs fit 16GB–64GB of unified memory, real tok/s numbers, Ollama setup, and the best value config.

HardwareRead →
Intermediate 6 min

Best Local LLM for RTX 3090 (24GB) in 2026

RTX 3090 24GB in 2026: the cheapest path to 27B-35B local LLMs. What tokens/sec to expect, how loud it runs, and when to buy a 4090 or 5070 Ti instead.

RankingsRead →
Intermediate 7 min

What LLMs Can You Run on the RTX 3060 12GB in 2026?

RTX 3060 12GB local LLM guide for 2026: which models fit in 12GB, real tokens/sec, Ollama setup, KV cache tuning, and how it stacks up against the RTX 4060 and 5060.

HardwareRead →
Beginner 8 min

How to Install text-generation-webui with CUDA on Windows

Step-by-step guide to install text-generation-webui with CUDA on Windows, verify GPU acceleration, pick the right loader, and fix common CUDA errors.

GuidesRead →
Beginner 9 min

How to Self-Host TabbyAPI as an OpenAI-Compatible Endpoint

Self-host TabbyAPI as an OpenAI-compatible endpoint. Docker and manual install, config.yml, EXL2/EXL3 models, GPU requirements, and tuning for RTX 3090/4090/5090.

ToolsRead →
Beginner 9 min

How to Use LlamaIndex with a Local Ollama LLM

Step-by-step guide to run LlamaIndex with a local Ollama LLM. Install commands, a working RAG script, hardware specs, model benchmarks, and a clear verdict.

ToolsRead →
Beginner 9 min

How to Wire LangChain to a Local LLM — Production-Ready

Wire LangChain to a local LLM the production way: Ollama vs vLLM, timeouts, streaming, structured output, RAG, plus 2026 hardware benchmarks and cost math.

ToolsRead →
Beginner 9 min

How to Use LiteLLM Router with Ollama — OpenAI-Drop-In

Configure LiteLLM Router as an OpenAI-compatible proxy for Ollama in minutes. Load balancing, fallbacks, config.yaml, and 2026 latency benchmarks.

ToolsRead →
Beginner 8 min

How to Add a Local LLM to Claude Desktop via MCP

Step-by-step guide to bridging Ollama or LM Studio into Claude Desktop with an MCP server, so a local LLM handles private tasks alongside Claude in 2026.

ToolsRead →
Beginner 9 min

How to Use a Local Ollama LLM as Cursor's Backend

Use a local Ollama LLM as Cursor's backend the right way: expose an OpenAI-compatible endpoint, fix the context window, and pick the best local coding model.

ToolsRead →
Beginner 9 min

How to Wire Continue.dev to a Local Ollama LLM in VS Code

Wire Continue.dev to a local Ollama LLM in VS Code: exact models, config.yaml, hardware specs, and benchmarks for private, $0/month AI coding.

ToolsRead →
Beginner 9 min

How to Use Aider with Ollama as a Local Copilot

Set up Aider with Ollama for a private, zero-cost local copilot. Install steps, the ollama_chat/ prefix, OLLAMA_API_BASE, context tuning, and the best coder models.

ToolsRead →
Beginner 8 min

How to Self-Host OpenWebUI in Docker — 5-Minute Setup

Self-host Open WebUI in Docker in under 5 minutes. Copy-paste run command, Compose file, persistence fix, Ollama hookup, and production hardening.

ToolsRead →
Beginner 9 min

How to Build llama.cpp with Metal on Mac M-Series

Build llama.cpp with Metal on Mac M-series in under 5 minutes. CMake steps, verification, M1–M4 benchmarks, tuning flags, and fixes for common errors.

HardwareRead →
Beginner 9 min

How to Run Llama on Apple Silicon with MLX — Native Performance

Run Llama on Apple Silicon with MLX for native performance. Install steps, memory requirements, MLX vs llama.cpp benchmarks, quantization picks, and an OpenAI API.

HardwareRead →
Beginner 9 min

How to Install Ollama with ROCm on AMD GPU

Step-by-step guide to installing Ollama with ROCm on AMD GPUs. Covers RX 6000/7000/9000 support, HSA overrides, benchmarks, and troubleshooting for 2026.

ToolsRead →
Beginner 8 min

How to Run vLLM in Docker with NVIDIA CUDA — 10-Minute Setup

Run vLLM in Docker with NVIDIA CUDA in under 10 minutes. Step-by-step setup, GPU passthrough, model VRAM tables, benchmarks, and troubleshooting fixes.

ToolsRead →
Beginner 9 min

How to Build llama.cpp from Source with CUDA — 2026 Guide

Step-by-step 2026 guide to build llama.cpp from source with CUDA: exact CMake flags, GPU arch codes, toolkit versions, benchmarks, and fixes for common errors.

ToolsRead →
1–30 of 92 guides
What you get

Built around your decision, not vendor benchmarks.

Four practical tools that answer the questions you actually have when picking an LLM.

01 RANKINGS

Hardware-matched rankings

Best local LLM for RTX 4090, RTX 5090, Mac M4 Max, Snapdragon X — cut through the noise with rankings that respect your VRAM, memory, and target speed.

02 CALCULATOR

Cost ROI: self-hosted vs API

Sliders for your monthly token volume, electricity cost, GPU amortization. Real break-even point against GPT-5, Claude, Gemini, DeepSeek — updated pricing.

03 OPEN DATA

Public API & MCP server

178 JSON endpoints under CC BY 4.0, free to use in your own tools. Official MCP server on GitHub for ChatGPT, Claude Desktop, and Cursor.

04 METHOD

Independent benchmark pipeline

Continuous benchmarking against published model versions and quantizations. No press-kit numbers, no marketing decks — just tokens/sec backed by our open data API.

Who's behind this

Independent. Skin in the game.

BestLLMfor is built and operated by Mohamed Meguedmi — one engineer, a continuous benchmark pipeline, a public data API and an open-source MCP server.

No VC, no SEO farm. One engineer obsessed with tracking every model worth running, and publishing what the numbers say — transparently.

BENCHMARK PIPELINE
Models tracked239+ (daily)
Quants testedQ4 · Q5 · Q8 · FP16
Data API178 JSON · CC BY 4.0
MCP serverPublic · open source
Methodologysee how →
Newsletter

Your GPU cheat sheet,
then a hands-on series.

Your VRAM cheat sheet now, then a short starter series (about nine emails over ten days), then only occasional updates. One-click unsubscribe.

Want the deep dive instead? See the Local Copilot Kit →