Osaurus

Model compatibility

Which model can your Mac run?

Every Apple Silicon Mac runs Osaurus. Pick your chip and memory below to see which curated models fit and how fast they'll feel.


Your Mac

8 run well4 tight fit30 too largeof 42 curated models · ~11 GB usable for models
Picked for your Mac

Gemma 4 E2B

Runs Well

8-bit · 5.5 GB · Top Pick

The smallest high-precision multimodal model — the best quality that still runs on any Mac. 128K context.

~6.9 GB~39 tok/s

Bonsai 27B 1-bit

Runs Well

1-bit JANG · 4.4 GB · Top Pick

The smallest of the Bonsai family at ~4.7 GB — a 27B-class vision model that fits almost anywhere.

~5.4 GB~23 tok/sVision

Gemma 4 E4B

Runs Well

4-bit · 6.4 GB

Smallest-download E4B build — a lower-precision fallback for RAM-constrained Macs.

~8.0 GB~39 tok/sVision

Gemma 4 E4B QAT

Runs Well

QAT MXFP4 · 5.5 GB

Quantization-aware 4-bit multimodal edge model. 128K context.

~6.9 GB~39 tok/sVision

ZAYA1 8B

Runs Well

MXFP4 · 5.5 GB

Compact reasoning and tool-use model with CCA hybrid attention.

~6.8 GB~20 tok/s

ZAYA1 8B

Runs Well

JANGTQ4 · 4.6 GB

Reasoning and tool-use model with 4-bit TurboQuant experts and CCA hybrid attention.

~5.8 GB~20 tok/s

Gemma 4 E2B

Runs Well

4-bit · 4.1 GB

Smallest-download Gemma 4 build — runs on any Mac.

~5.1 GB~78 tok/s

Gemma 4 E2B QAT

Runs Well

QAT MXFP4 · 3.8 GB

Quantization-aware 4-bit build — the smallest multimodal floor. 128K context.

~4.7 GB~78 tok/s

Gemma 4 E4B

Tight Fit

8-bit · 8.4 GB · Top Pick

Recommended multimodal edge model — the best first-run quality in the E4B family. Images, video, audio. 128K context.

~10.5 GB~20 tok/sVision

Bonsai 27B Ternary

Tight Fit

Ternary JANG · 7.5 GB · Top Pick

27B-class vision model compressed to ~8 GB with ternary weights. Big-model quality on mainstream Macs.

~9.4 GB~11 tok/sVision

LFM2.5 8B

Tight Fit

MXFP8 · 8.1 GB

Liquid AI hybrid MoE (~1B active) — high-precision, fast chat on Apple Silicon. 128K context.

~10.2 GB~78 tok/s

Gemma 4 12B QAT

Tight Fit

QAT MXFP4 · 7.4 GB

Quantization-aware 4-bit build of the 12B multimodal model. 128K context.

~9.2 GB~13 tok/sVision

Ornith 1.0 35B

Too Large

MXFP8 · 34.2 GB · Top Pick

State-of-the-art open agentic coding for its size. Vision-language MoE at near-lossless precision. 256K context.

~42.7 GBVision

Gemma 4 12B

Too Large

MXFP8 · 24.9 GB · Top Pick

Google's multimodal workhorse — images, video, and audio at high precision. 128K context.

~31.2 GBVision

Ornith 1.0 9B

Too Large

MXFP8 · 9.5 GB · Top Pick

Vision-language model tuned for agentic coding, at near-lossless precision. 256K context.

~11.8 GBVision

Kimi K2.6

Too Large

JANGTQ-K · 328 GB

Frontier-class vision model, K-quant TurboQuant. A very large specialist.

~410.2 GBVision

MiniMax M2.7

Too Large

JANGTQ · 113 GB

228B agentic MoE with 2-bit TurboQuant experts — the smallest footprint of the family. 192K context.

~141.3 GB

MiniMax M2.7

Too Large

JANGTQ4 · 109 GB

228B agentic MoE with 4-bit TurboQuant experts — near-bf16 quality at a quarter of the disk. 192K context.

~136.1 GB

Hunyuan 3 Preview

Too Large

JANGTQ-K · 102 GB

295B MoE preview, K-quant TurboQuant. A very large specialist.

~126.9 GB

Nemotron-3 Ultra 550B

Too Large

JANGTQ-1L · 98.4 GB

NVIDIA's 550B (~55B active) reasoning MoE — a showcase for the highest-memory Macs.

~123.0 GB

DeepSeek V4 Flash

Too Large

JANGTQ-K · 80.0 GB

Reasoning model with CSA/HSA/SWA hybrid attention, K-quant TurboQuant. A large specialist.

~100.0 GB

Mistral Medium 3.5 128B

Too Large

MXFP4 · 79.9 GB

128B with Pixtral vision — fastest decode. 256K context, 24-language coverage.

~99.8 GBVision

Step 3.7 Flash

Too Large

JANG-K · 74.2 GB

Vision-language specialist, JANG K-quant.

~92.8 GBVision

DeepSeek V4 Flash

Too Large

JANGTQ2 · 74.2 GB

Reasoning model with CSA/HSA/SWA hybrid attention, 2-bit TurboQuant. A large specialist.

~92.7 GB

Ling 2.6 Flash

Too Large

MXFP4 · 62.6 GB

BailingHybrid MoE — the highest-quality Ling local path.

~78.3 GB

Mistral Medium 3.5 128B

Too Large

JANGTQ · 38.0 GB

128B with Pixtral vision, 2-bit TurboQuant text decoder — ~41 GB footprint. 256K context.

~47.5 GBVision

MiniMax M2.7 Small

Too Large

JANGTQ · 35.8 GB

The most-liked model on the OsaurusAI hub — an agentic MoE with TurboQuant experts. 192K context.

~44.7 GB

Qwen 3.6 35B MoE

Too Large

MXFP8 MTP · 35.0 GB

35B MoE (~3B active) vision model with speculative decode — the precision-first sibling of the MXFP4 build. 256K context.

~43.7 GBVision

Ling 2.6 Flash

Too Large

JANGTQ · 28.5 GB

BailingHybrid MoE with TurboQuant experts — a smaller local footprint.

~35.6 GB

Qwen 3.6 27B

Too Large

MXFP8 MTP · 27.1 GB

High-precision build with multi-token-prediction speculative decode — precise and fast. 256K context.

~33.9 GBVision

Nemotron-3 Nano 30B

Too Large

MXFP4 · 21.1 GB

NVIDIA reasoning hybrid (Mamba-2 + MoE) — the fastest decode path of the family. 262K context.

~26.4 GB

Laguna XS.2

Too Large

MXFP4 · 19.5 GB

Poolside's 33B/3B-active agentic-coding MoE — fastest decode. 131K context.

~24.4 GB

Nemotron-3 Nano 30B

Too Large

JANGTQ4 · 18.6 GB

Reasoning hybrid with 4-bit TurboQuant experts — near-bf16 quality. 262K context.

~23.2 GB

Holo3 35B MoE

Too Large

JANGTQ4 · 18.3 GB

Computer-use GUI agent with 4-bit TurboQuant experts.

~22.9 GB

Gemma 4 31B QAT

Too Large

QAT MXFP4 · 18.2 GB

The largest dense Gemma 4, quantization-aware 4-bit. 128K context.

~22.8 GBVision

Holo3 35B MoE

Too Large

MXFP4 · 18.0 GB

Computer-use GUI agent — a vision MoE specialist.

~22.5 GB

Qwen 3.6 35B MoE

Too Large

MXFP4 · 18.0 GB

35B MoE vision model — best quality per byte. 256K context.

~22.5 GBVision

Gemma 4 12B

Too Large

MXFP4 · 14.8 GB

Smaller, lower-precision companion to the 12B MXFP8 Top Pick. 128K context.

~18.5 GBVision

Gemma 4 26B-A4B QAT

Too Large

QAT MXFP4 · 14.6 GB

Quantization-aware 4-bit MoE (~4B active) vision model. 128K context.

~18.2 GBVision

Qwen 3.6 27B

Too Large

MXFP4 · 14.2 GB

Dense vision model with the best quality per byte — the most-downloaded model in the catalog. 256K context.

~17.7 GBVision

Nemotron-3 Nano 30B

Too Large

JANGTQ2 · 11.7 GB

Reasoning hybrid with 2-bit TurboQuant experts — the smallest footprint of the family. 262K context.

~14.7 GB

Laguna XS.2

Too Large

JANGTQ · 9.4 GB

Agentic-coding MoE with 2-bit TurboQuant experts — the smallest footprint at ~10 GB. 131K context.

~11.8 GB

Speed is a rough estimate from the chip's memory bandwidth — real numbers vary with context length, quantization, and what else your Mac is doing.

The newsletter

Release notes, new skills, and product news. No spam.

No spam. Unsubscribe anytime.