Open source · MIT · v4.8.0
Make AI run onevery machine.
Text, voice, vision, images, video, 3D and agents, from one open runtime. It works on the laptop you already own, and scales to a room full of GPUs when you have one. When the engine we need is too heavy, too closed, or does not exist, we write it.
APEX quantization
The model you could not fit, on the card you already own.
A 35B mixture-of-experts model is 64.6 GB at full precision, which puts it out of reach of every consumer GPU. APEX gets it to 12.2 GB, and it runs at 74 tokens a second, more than twice the speed of the original. Quality barely moves. The file is an ordinary GGUF, so stock llama.cpp opens it with no patches, and 201 of them are already sitting in the LocalAI gallery.
| Build | Size | Perplexity | HellaSwag | MMLU | tg128 t/s |
|---|---|---|---|---|---|
| F16 | 64.6 GB | 6.537 | 82.5% | 41.5% | 30.4 |
| Q8_0 | 34.4 GB | 6.533 | 83.0% | 41.2% | 52.5 |
| Unsloth UD-Q8_K_XL | 45.3 GB | 6.536 | 82.5% | 41.3% | 36.4 |
| APEX Quality | 21.3 GB | 6.527 | 83.0% | 41.2% | 62.3 |
| APEX I-Quality | 21.3 GB | 6.552 | 83.5% | 41.4% | 63.1 |
| APEX Compact | 16.1 GB | 6.783 | 82.5% | 40.9% | 69.8 |
| APEX Mini | 12.2 GB | 7.088 | 81.0% | 41.3% | 74.4 |
Qwen3.5-35B-A3B on an NVIDIA DGX Spark (GB10). Perplexity on wikitext-2-raw at context 2048. Full methodology in the technical report.
nib
An agent you can drop on any box you SSH into.
One Go binary, about 20 MB, no runtime and no daemon. Press Ctrl+Space anywhere and it opens. Point it at any OpenAI-compatible endpoint, including a model running on your own laptop, and it works. Tool calls go through an approval gate you control, and Claude Code plugins load as they are.
From the team
We publish the numbers, including the ones that cost us.
APEX takes a 35B model from 64.6 GB to 12.2 GB, and from 30.4 to 74.4 tokens a second. Perplexity goes from 6.537 to 7.088. That trade is in the post too, with the command that produced it.
29 July 2026
LocalAI, from March 2023 to now
Three years, 133 releases and 224 contributors later. The four changes that mattered most were making the core small, adding agents, making it a cluster, …
Read the post →28 July 2026
What landed in LocalAI 4.8
The web interface got 3.48x lighter, gallery entries now install the build your hardware can actually run, and there is a new inference engine in the box. …
Read the post →24 July 2026
Why we write our own C and C++ engines
A 66 MiB binary instead of a 9.1 GiB virtualenv, depth estimation that beats PyTorch on CPU in half the memory, and biometrics that match insightface bit …
Read the post →
