Open source · MIT · v4.8.0
Make AI run onevery machine.
Text, voice, vision, images, video, 3D and agents, from one open runtime. It works on the laptop you already own, and scales to a room full of GPUs when you have one. When the engine we need is too heavy, too closed, or does not exist, we write it.
APEX quantization
The model you could not fit, on the card you already own.
The engine decides how fast a model runs. The weights decide whether it runs at all, so we build those too. APEX assigns a different precision to every tensor and every layer: a 35B mixture-of-experts model goes from 64.6 GB, out of reach of any consumer GPU, to 12.2 GB at 74 tokens a second. That is more than twice the speed of the original, and quality barely moves. The file is an ordinary GGUF, so stock llama.cpp opens it with no patches, and 201 of them are already sitting in the LocalAI gallery.
| Build | Size | Perplexity | HellaSwag | MMLU | tg128 t/s |
|---|---|---|---|---|---|
| F16 | 64.6 GB | 6.537 | 82.5% | 41.5% | 30.4 |
| Q8_0 | 34.4 GB | 6.533 | 83.0% | 41.2% | 52.5 |
| Unsloth UD-Q8_K_XL | 45.3 GB | 6.536 | 82.5% | 41.3% | 36.4 |
| APEX Quality | 21.3 GB | 6.527 | 83.0% | 41.2% | 62.3 |
| APEX I-Quality | 21.3 GB | 6.552 | 83.5% | 41.4% | 63.1 |
| APEX Compact | 16.1 GB | 6.783 | 82.5% | 40.9% | 69.8 |
| APEX Mini | 12.2 GB | 7.088 | 81.0% | 41.3% | 74.4 |
Qwen3.5-35B-A3B on an NVIDIA DGX Spark (GB10). Perplexity on wikitext-2-raw at context 2048. Full methodology in the technical report.
the agent
The agent is already in the binary.
Run local-ai chat and you are talking to an agent that already knows where your models are. It runs shell commands behind an approval gate you control, delegates to sub-agents, and loads MCP servers, plugins and skills. Claude Code plugins load as they are.
It is also nib, a single ~20 MB Go binary with no runtime and no daemon, so you can drop the same agent on any box you SSH into and press Ctrl+Space.
From the team
We publish the numbers, including the ones that cost us.
APEX takes a 35B model from 64.6 GB to 12.2 GB, and from 30.4 to 74.4 tokens a second. Perplexity goes from 6.537 to 7.088. That trade is in the post too, with the command that produced it.
1 August 2026
What landed in LocalAI 4.8
A new inference engine, 3D generation, one backend that serves six audio endpoints, and a web interface 3.48x lighter. 321 pull requests in eighteen days.
Read the post →29 July 2026
LocalAI, from March 2023 to now
Three years, 133 releases and 224 contributors later. The four changes that mattered most were making the core small, adding agents, making it a cluster, …
Read the post →24 July 2026
Why we write our own C and C++ engines
A 66 MiB binary instead of a 9.1 GiB virtualenv, depth estimation that beats PyTorch on CPU in half the memory, and biometrics that match insightface bit …
Read the post →
