Evidence-backed models that fit common laptops and consumer GPUs.
Best low-memory chat option in the demo catalog.
A fast, memory-efficient chat model for laptops and edge devices.
Confirmed coding Runs on 12 GB VRAM.
A compact code-generation model tuned for local development workflows.
Strong reasoning with a 10 GB minimum recommendation.
An open-weight reasoning model designed for careful, inspectable local evaluations.
Independent reproduction on an 8 GB Mac.