Model compatibility
Which model can your Mac run?
Every Apple Silicon Mac runs Osaurus. Pick your chip and memory below to see which curated models fit and how fast they'll feel.
Your Mac
Gemma 4 E2B
The smallest high-precision multimodal model — the best quality that still runs on any Mac. 128K context.
Bonsai 27B 1-bit
The smallest of the Bonsai family at ~4.7 GB — a 27B-class vision model that fits almost anywhere.
Gemma 4 E4B
Smallest-download E4B build — a lower-precision fallback for RAM-constrained Macs.
Gemma 4 E4B QAT
Quantization-aware 4-bit multimodal edge model. 128K context.
ZAYA1 8B
Compact reasoning and tool-use model with CCA hybrid attention.
ZAYA1 8B
Reasoning and tool-use model with 4-bit TurboQuant experts and CCA hybrid attention.
Gemma 4 E2B
Smallest-download Gemma 4 build — runs on any Mac.
Gemma 4 E2B QAT
Quantization-aware 4-bit build — the smallest multimodal floor. 128K context.
Gemma 4 E4B
Recommended multimodal edge model — the best first-run quality in the E4B family. Images, video, audio. 128K context.
Bonsai 27B Ternary
27B-class vision model compressed to ~8 GB with ternary weights. Big-model quality on mainstream Macs.
LFM2.5 8B
Liquid AI hybrid MoE (~1B active) — high-precision, fast chat on Apple Silicon. 128K context.
Gemma 4 12B QAT
Quantization-aware 4-bit build of the 12B multimodal model. 128K context.
Ornith 1.0 35B
State-of-the-art open agentic coding for its size. Vision-language MoE at near-lossless precision. 256K context.
Gemma 4 12B
Google's multimodal workhorse — images, video, and audio at high precision. 128K context.
Ornith 1.0 9B
Vision-language model tuned for agentic coding, at near-lossless precision. 256K context.
Kimi K2.6
Frontier-class vision model, K-quant TurboQuant. A very large specialist.
MiniMax M2.7
228B agentic MoE with 2-bit TurboQuant experts — the smallest footprint of the family. 192K context.
MiniMax M2.7
228B agentic MoE with 4-bit TurboQuant experts — near-bf16 quality at a quarter of the disk. 192K context.
Hunyuan 3 Preview
295B MoE preview, K-quant TurboQuant. A very large specialist.
Nemotron-3 Ultra 550B
NVIDIA's 550B (~55B active) reasoning MoE — a showcase for the highest-memory Macs.
DeepSeek V4 Flash
Reasoning model with CSA/HSA/SWA hybrid attention, K-quant TurboQuant. A large specialist.
Mistral Medium 3.5 128B
128B with Pixtral vision — fastest decode. 256K context, 24-language coverage.
Step 3.7 Flash
Vision-language specialist, JANG K-quant.
DeepSeek V4 Flash
Reasoning model with CSA/HSA/SWA hybrid attention, 2-bit TurboQuant. A large specialist.
Ling 2.6 Flash
BailingHybrid MoE — the highest-quality Ling local path.
Mistral Medium 3.5 128B
128B with Pixtral vision, 2-bit TurboQuant text decoder — ~41 GB footprint. 256K context.
MiniMax M2.7 Small
The most-liked model on the OsaurusAI hub — an agentic MoE with TurboQuant experts. 192K context.
Qwen 3.6 35B MoE
35B MoE (~3B active) vision model with speculative decode — the precision-first sibling of the MXFP4 build. 256K context.
Ling 2.6 Flash
BailingHybrid MoE with TurboQuant experts — a smaller local footprint.
Qwen 3.6 27B
High-precision build with multi-token-prediction speculative decode — precise and fast. 256K context.
Nemotron-3 Nano 30B
NVIDIA reasoning hybrid (Mamba-2 + MoE) — the fastest decode path of the family. 262K context.
Laguna XS.2
Poolside's 33B/3B-active agentic-coding MoE — fastest decode. 131K context.
Nemotron-3 Nano 30B
Reasoning hybrid with 4-bit TurboQuant experts — near-bf16 quality. 262K context.
Holo3 35B MoE
Computer-use GUI agent with 4-bit TurboQuant experts.
Gemma 4 31B QAT
The largest dense Gemma 4, quantization-aware 4-bit. 128K context.
Holo3 35B MoE
Computer-use GUI agent — a vision MoE specialist.
Qwen 3.6 35B MoE
35B MoE vision model — best quality per byte. 256K context.
Gemma 4 12B
Smaller, lower-precision companion to the 12B MXFP8 Top Pick. 128K context.
Gemma 4 26B-A4B QAT
Quantization-aware 4-bit MoE (~4B active) vision model. 128K context.
Qwen 3.6 27B
Dense vision model with the best quality per byte — the most-downloaded model in the catalog. 256K context.
Nemotron-3 Nano 30B
Reasoning hybrid with 2-bit TurboQuant experts — the smallest footprint of the family. 262K context.
Laguna XS.2
Agentic-coding MoE with 2-bit TurboQuant experts — the smallest footprint at ~10 GB. 131K context.