https://www.canirun.ai/ CanIRun.ai [tier list] [docs] [why] Can I Run AI locally? Find out which AI models your machine can actually run. [detecting...] | [--] | [] | WebGPU Estimates based on browser APIs. Actual specs may vary. S 0 A 0 B 0 C 0 D 0 F 0 [ ] / [All grades ] [All tasks] [All providers] [Sort: Score ] List Llama 3.1 8B 1 year ago Meta * 8B Meta's versatile 8B -- great quality/speed ratio 4.1 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-07 Architecture Dense Memory -- chatcodereasoning Qwen 3.5 9B 1 month ago Alibaba * 9B Multimodal Qwen 3.5 mid-size 4.6 GB * 32K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2026-02 Architecture Dense Memory -- chatvision Phi-4 14B 1 year ago Microsoft * 14B Microsoft's reasoning-focused model 7.2 GB * 16K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-12 Architecture Dense Memory -- reasoningcode GPT-OSS 20B 7mo ago OpenAI * 21B OpenAI's open-weight MoE with configurable reasoning 10.8 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-08 Architecture MoE Active 3.6B active Memory -- chatreasoningcode Mistral Small 3.1 24B 1 year ago Mistral AI * 24B Multimodal Mistral with vision support 12.3 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-03 Architecture Dense Memory -- chatvisioncode Gemma 3 27B 1 year ago Google * 27B Google's flagship Gemma 3 model 13.8 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-03 Architecture Dense Memory -- chatvisionreasoning Qwen 2.5 Coder 32B 1 year ago Alibaba * 32B Best open-source coding model at release 16.4 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-11 Architecture Dense Memory -- code Qwen 3 32B 11mo ago Alibaba * 32B Qwen 3 flagship dense model 16.4 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-04 Architecture Dense Memory -- chatcodereasoning DeepSeek R1 Distill 32B 1 year ago DeepSeek * 32B R1 reasoning distilled into Qwen 32B -- sweet spot 16.4 GB * 64K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-01 Architecture Dense Memory -- reasoning Llama 3.3 70B 1 year ago Meta * 70B Best open model at 70B class 35.9 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-12 Architecture Dense Memory -- chatreasoningcode Llama 4 Scout 17B 11mo ago Meta * 109B MoE with 16 experts, 17B active params 55.8 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-04 Architecture MoE Active 17B active Memory -- chatvisionreasoning GPT-OSS 120B 7mo ago OpenAI * 117B OpenAI's flagship open-weight MoE -- 52.6% SWE-bench 59.9 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-08 Architecture MoE Active 5.1B active Memory -- chatreasoningcode Devstral 2 123B 3mo ago Mistral AI * 123B Dense 123B coding model -- 72.2% SWE-bench Verified 63 GB * 256K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-12 Architecture Dense Memory -- code DeepSeek R1 1 year ago DeepSeek * 671B Massive MoE reasoning model -- 37B active 343.7 GB * 64K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-01 Architecture MoE Active 37B active Memory -- reasoning DeepSeek V3.2 3mo ago DeepSeek * 685B State-of-the-art MoE -- 37B active params 350.9 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-12 Architecture MoE Active 37B active Memory -- chatcodereasoning Kimi K2 8mo ago Moonshot AI * 1T 1T-param MoE with 384 experts -- 32B active, strong agentic coding 512.2 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-07 Architecture MoE Active 32B active Memory -- chatreasoningcode All models Qwen 3.5 0.8B 1 month ago Alibaba * 0.8B Ultra-tiny model for embedded and edge 0.5 GB * 32K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2026-02 Architecture Dense Memory -- chatedge Llama 3.2 1B 1 year ago Meta * 1B Meta's smallest Llama for edge devices 0.5 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-09 Architecture Dense Memory -- chatedge Gemma 3 1B 1 year ago Google * 1B Google's tiny Gemma for on-device 0.5 GB * 32K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-03 Architecture Dense Memory -- chatedge TinyLlama 1.1B 2y ago Community * 1.1B Ultralight model for constrained devices 0.6 GB * 2K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-01 Architecture Dense Memory -- chatedge Qwen 2.5 Coder 1.5B 1 year ago Alibaba * 1.5B Ultra-lightweight coding model 0.8 GB * 32K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-11 Architecture Dense Memory -- code DeepSeek R1 1.5B 1 year ago DeepSeek * 1.5B Tiny reasoning model distilled from R1 0.8 GB * 64K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-01 Architecture Dense Memory -- reasoning Qwen 3 1.7B 11mo ago Alibaba * 1.7B Compact multilingual Qwen 3 0.9 GB * 32K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-04 Architecture Dense Memory -- chatmultilingual Qwen 3.5 2B 1 month ago Alibaba * 2B Small multimodal Qwen 3.5 1 GB * 32K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2026-02 Architecture Dense Memory -- chatmultilingual Gemma 2 2B 1 year ago Google * 2B Google's compact open model 1 GB * 8K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-06 Architecture Dense Memory -- chatedge Llama 3.2 3B 1 year ago Meta * 3B Lightweight Llama for mobile and edge 1.5 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-09 Architecture Dense Memory -- chatcode SmolLM3 3B 8mo ago HuggingFace * 3B Lightweight multilingual reasoning 1.5 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-07 Architecture Dense Memory -- chatreasoning Phi-3.5 Mini 1 year ago Microsoft * 3.8B Microsoft's efficient small model with long context 1.9 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-08 Architecture Dense Memory -- reasoningcodechat Phi-4 Mini Reasoning 11mo ago Microsoft * 3.8B Lightweight reasoning model 1.9 GB * 16K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-04 Architecture Dense Memory -- reasoning Qwen 3 4B 11mo ago Alibaba * 4B Compact Qwen 3 for general tasks 2 GB * 32K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-04 Architecture Dense Memory -- chatcode Gemma 3 4B 1 year ago Google * 4B Multimodal Gemma with 128K context 2 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-03 Architecture Dense Memory -- chatvision Qwen 3.5 4B 1 month ago Alibaba * 4B Small multimodal Qwen 3.5 2 GB * 32K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2026-02 Architecture Dense Memory -- chatmultilingual Mistral 7B v0.3 1 year ago Mistral AI * 7B High-quality 7B with sliding window attention 3.6 GB * 32K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-05 Architecture Dense Memory -- chatreasoning Qwen 2.5 7B 1 year ago Alibaba * 7B Strong multilingual and coding capabilities 3.6 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-09 Architecture Dense Memory -- chatmultilingualcode Qwen 2.5 Coder 7B 1 year ago Alibaba * 7B Dedicated coding model 3.6 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-11 Architecture Dense Memory -- code DeepSeek R1 Distill 7B 1 year ago DeepSeek * 7B R1 reasoning distilled into Qwen 7B 3.6 GB * 64K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-01 Architecture Dense Memory -- reasoning Qwen 3 8B 11mo ago Alibaba * 8B Qwen 3 with thinking mode support 4.1 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-04 Architecture Dense Memory -- chatcodereasoning Ministral 8B 1 year ago Mistral AI * 8B Mistral's efficient 8B model 4.1 GB * 32K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-10 Architecture Dense Memory -- chat Gemma 2 9B 1 year ago Google * 9B Google's best mid-size open model 4.6 GB * 8K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-06 Architecture Dense Memory -- chatreasoning GLM-4 9B 1 year ago Zhipu AI * 9B Multilingual model supporting 26 languages with 128K context 4.6 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-06 Architecture Dense Memory -- chatmultilingualcode Nemotron Nano 9B v2 9mo ago NVIDIA * 9B Hybrid Mamba2 architecture for reasoning 4.6 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-06 Architecture Dense Memory -- reasoning Llama 3.2 11B Vision 1 year ago Meta * 11B Multimodal vision and text model 5.6 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-09 Architecture Dense Memory -- chatvision Gemma 3 12B 1 year ago Google * 12B Multimodal Gemma with 128K context 6.1 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-03 Architecture Dense Memory -- chatvisionreasoning Mistral Nemo 12B 1 year ago Mistral AI * 12B Multilingual 12B with 128K context 6.1 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-07 Architecture Dense Memory -- chatmultilingual Qwen 2.5 14B 1 year ago Alibaba * 14B Excellent quality for its size class 7.2 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-09 Architecture Dense Memory -- chatmultilingualreasoning Qwen 3 14B 11mo ago Alibaba * 14B Strong all-rounder with thinking mode 7.2 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-04 Architecture Dense Memory -- chatcodereasoning DeepSeek R1 Distill 14B 1 year ago DeepSeek * 14B R1 reasoning distilled into Qwen 14B 7.2 GB * 64K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-01 Architecture Dense Memory -- reasoning LFM2 24B 4mo ago Liquid AI * 24B Hybrid MoE with convolution+attention layers -- 2.3B active 12.3 GB * 32K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-11 Architecture MoE Active 2.3B active Memory -- chatedgerag Devstral Small 2 24B 3mo ago Mistral AI * 24B Coding-focused model with 256K context -- 68% SWE-bench 12.3 GB * 256K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-12 Architecture Dense Memory -- code Gemma 2 27B 1 year ago Google * 27B Google's largest Gemma 2 model 13.8 GB * 8K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-06 Architecture Dense Memory -- chatreasoning Qwen 3.5 27B 1 month ago Alibaba * 27.8B Flagship native multimodal Qwen 3.5 14.2 GB * 256K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2026-02 Architecture Dense Memory -- chatvisionreasoning Qwen 3 30B-A3B 11mo ago Alibaba * 30B MoE with only 3.3B active -- extremely efficient 15.4 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-04 Architecture MoE Active 3.3B active Memory -- chatreasoning Nemotron 3 Nano 30B 9mo ago NVIDIA * 30B MoE with 1M context and 3B active 15.4 GB * 1024K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-06 Architecture MoE Active 3B active Memory -- chatreasoning Qwen 2.5 32B 1 year ago Alibaba * 32B High-quality reasoning and multilingual 16.4 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-09 Architecture Dense Memory -- chatmultilingualreasoning EXAONE 4.0 32B 8mo ago LG AI * 32B Hybrid reasoning, multilingual 16.4 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-07 Architecture Dense Memory -- chatreasoning OLMo 2 32B 1 year ago Allen AI * 32B Fully open research model by Allen AI 16.4 GB * 4K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-03 Architecture Dense Memory -- chatreasoning Command R 35B 2y ago Cohere * 35B Optimized for retrieval-augmented generation 17.9 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-03 Architecture Dense Memory -- chatrag Qwen 3.5 35B-A3B 1 month ago Alibaba * 35B Efficient multimodal MoE with 3B active 17.9 GB * 256K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2026-02 Architecture MoE Active 3B active Memory -- chatvision Mixtral 8x7B 2y ago Mistral AI * 47B MoE with 12.9B active params 24.1 GB * 32K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2023-12 Architecture MoE Active 12.9B active Memory -- chatcode Qwen 2.5 72B 1 year ago Alibaba * 72B Alibaba's flagship open model 36.9 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-09 Architecture Dense Memory -- chatmultilingualreasoningcode Qwen 3.5 122B-A10B 1 month ago Alibaba * 122B Large multimodal MoE with 10B active 62.5 GB * 256K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2026-02 Architecture MoE Active 10B active Memory -- chatvisionreasoning Mixtral 8x22B 1 year ago Mistral AI * 141B Large MoE with 39B active params 72.2 GB * 64K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-04 Architecture MoE Active 39B active Memory -- chatcodereasoning Qwen 3 235B-A22B 11mo ago Alibaba * 235B Massive MoE with 22B active -- frontier quality 120.4 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-04 Architecture MoE Active 22B active Memory -- chatcodereasoning Qwen 3.5 397B-A17B 1 month ago Alibaba * 397B Largest multimodal Qwen 3.5 MoE 203.4 GB * 256K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2026-02 Architecture MoE Active 17B active Memory -- chatvisionreasoningcode Llama 4 Maverick 17B-128E 11mo ago Meta * 400B Multimodal MoE with 128 experts -- 17B active, 1M context 204.9 GB * 1024K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-04 Architecture MoE Active 17B active Memory -- chatvisionreasoningcode Llama 3.1 405B 1 year ago Meta * 405B Largest open-weight dense model by Meta 207.5 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2024-07 Architecture Dense Memory -- chatreasoningcode Qwen 3 Coder 480B 8mo ago Alibaba * 480B Largest open coding MoE -- 35B active 245.9 GB * 256K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-07 Architecture MoE Active 35B active Memory -- code DeepSeek V3.1 7mo ago DeepSeek * 671B Improved V3 with hybrid thinking and tool use 343.7 GB * 128K ctx * Q2_K Q3_K_M Q4_K_M Q5_K_M Q6_K Q8_0 F16 Released 2025-08 Architecture MoE Active 37B active Memory -- chatcodereasoning No models found Try adjusting your search or filters Reset filters Data sourced from llama.cpp, Ollama and LM Studio. Built by midudev