<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Blog on LocalAI</title><link>https://localai.io/blog/</link><description>Recent content in Blog on LocalAI</description><generator>Hugo</generator><language>en-GB</language><lastBuildDate>Fri, 02 Oct 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://localai.io/blog/index.xml" rel="self" type="application/rss+xml"/><item><title>Run decision models in LocalAI</title><link>https://localai.io/blog/decision-models/</link><pubDate>Fri, 02 Oct 2026 00:00:00 +0000</pubDate><guid>https://localai.io/blog/decision-models/</guid><description>Ask named questions about text and get structured answers for routing, moderation, and model selection.</description></item><item><title>What landed in LocalAI 4.11</title><link>https://localai.io/blog/what-landed-in-localai-4-11/</link><pubDate>Fri, 02 Oct 2026 00:00:00 +0000</pubDate><guid>https://localai.io/blog/what-landed-in-localai-4-11/</guid><description>Remember speakers, configure model failover, and inspect the models running on one LocalAI host.</description></item><item><title>Remember speakers from your recordings in LocalAI</title><link>https://localai.io/blog/diarization-speaker-profiles/</link><pubDate>Thu, 01 Oct 2026 00:00:00 +0000</pubDate><guid>https://localai.io/blog/diarization-speaker-profiles/</guid><description>Name speakers from an existing conversation and recognize them in later recordings.</description></item><item><title>What landed in LocalAI 4.9</title><link>https://localai.io/blog/what-landed-in-localai-4-9/</link><pubDate>Thu, 20 Aug 2026 00:00:00 +0000</pubDate><guid>https://localai.io/blog/what-landed-in-localai-4-9/</guid><description>Authentication now denies by default, chat can compress its own history, models and backends each got one page, and vllm-cpp builds for cards other than Blackwell. 146 pull requests in thirteen days.</description></item><item><title>What landed in LocalAI 4.8</title><link>https://localai.io/blog/what-landed-in-localai-4-8/</link><pubDate>Tue, 04 Aug 2026 00:00:00 +0000</pubDate><guid>https://localai.io/blog/what-landed-in-localai-4-8/</guid><description>A new inference engine, a terminal agent in the CLI, 3D generation, and a web interface 3.48x lighter. 386 pull requests in twenty-two days.</description></item><item><title>LocalAI, from March 2023 to now</title><link>https://localai.io/blog/localai-since-march-2023/</link><pubDate>Wed, 29 Jul 2026 00:00:00 +0000</pubDate><guid>https://localai.io/blog/localai-since-march-2023/</guid><description>Three years, 133 releases and 224 contributors later. Here are the four decisions that shaped it: making the core small, adding agents, making it a cluster, and giving it eyes and ears.</description></item><item><title>Why we write our own C and C++ engines</title><link>https://localai.io/blog/why-we-write-our-own-engines/</link><pubDate>Fri, 24 Jul 2026 00:00:00 +0000</pubDate><guid>https://localai.io/blog/why-we-write-our-own-engines/</guid><description>Eighteen of our backends are C or C++ ports we wrote from scratch instead of wrapping an upstream engine. Here is why we did it and what we measured.</description></item><item><title>parakeet.cpp: the same NeMo transcript, without the Python</title><link>https://localai.io/blog/parakeet-cpp-asr-on-cpu/</link><pubDate>Fri, 05 Jun 2026 00:00:00 +0000</pubDate><guid>https://localai.io/blog/parakeet-cpp-asr-on-cpu/</guid><description>The same transcript as NVIDIA NeMo at a median 1.40x on CPU, and about 27x the speed of whisper.cpp, from one binary and one GGUF file.</description></item><item><title>LocalAI 4.3: signed backends, and the prompt cache that was off</title><link>https://localai.io/blog/what-landed-in-localai-4-3/</link><pubDate>Sun, 24 May 2026 00:00:00 +0000</pubDate><guid>https://localai.io/blog/what-landed-in-localai-4-3/</guid><description>Keyless cosign verification for backend OCI images, the llama.cpp prompt cache enabled by default, per-API-key usage attribution, and the replica-pinning bug that kept a second node idle.</description></item><item><title>LocalAI 4.2: who spoke when, and whose face is that</title><link>https://localai.io/blog/what-landed-in-localai-4-2/</link><pubDate>Mon, 11 May 2026 00:00:00 +0000</pubDate><guid>https://localai.io/blog/what-landed-in-localai-4-2/</guid><description>A /v1/audio/diarization endpoint, voice and face recognition with liveness, a drop-in Ollama API, and eleven new backends.</description></item><item><title>APEX: a 35B MoE model at 12.2 GB, and faster than F16</title><link>https://localai.io/blog/apex-moe-quantization/</link><pubDate>Fri, 10 Apr 2026 00:00:00 +0000</pubDate><guid>https://localai.io/blog/apex-moe-quantization/</guid><description>Qwen3.5-35B-A3B goes from 64.6 GB to 12.2 GB and speeds up from 30.4 to 74.4 tokens per second. Perplexity moves from 6.537 to 7.088. Here is the precision assignment that does it, and where the quality drops.</description></item><item><title>LocalAI 4.1: more than one box, and more than one user</title><link>https://localai.io/blog/what-landed-in-localai-4-1/</link><pubDate>Thu, 02 Apr 2026 00:00:00 +0000</pubDate><guid>https://localai.io/blog/what-landed-in-localai-4-1/</guid><description>Distributed cluster mode that places requests by real free VRAM, OIDC with per-user API keys and quotas, and LoRA fine-tuning that exports straight to GGUF.</description></item><item><title>LocalAI 4.0: agents in the core, and a React interface</title><link>https://localai.io/blog/what-landed-in-localai-4-0/</link><pubDate>Sat, 14 Mar 2026 00:00:00 +0000</pubDate><guid>https://localai.io/blog/what-landed-in-localai-4-0/</guid><description>Native agent orchestration with the Agenthub, a rewritten interface with Canvas mode, MCP Apps with tool streaming, and two things removed.</description></item><item><title>LocalAI 3.10: the Anthropic and Responses APIs, and one image for every GPU</title><link>https://localai.io/blog/what-landed-in-localai-3-10/</link><pubDate>Sun, 18 Jan 2026 00:00:00 +0000</pubDate><guid>https://localai.io/blog/what-landed-in-localai-3-10/</guid><description>A /v1/messages endpoint that Claude clients can talk to unchanged, Open Responses compatibility that passes the official acceptance tests, and GPU libraries moved inside the backend containers so one image works on any hardware.</description></item></channel></rss>