* [The LeaderWorkerSet API (LWS) v0.10.0](https://github.com/kubernetes-sigs/lws) – API for deploying and managing groups of pods as a single unit with leader and worker roles for multi-host inference workloads. * [QVAC sdk-v0.17.0](https://github.com/tetherto/qvac) – Local-first, cross-platform SDK for building peer-to-peer AI apps with local model inference, speech, translation, and RAG. * [OME (Open Model Engine) v1.2.2](https://github.com/ome-projects/ome) – Kubernetes operator for enterprise-grade management and serving of large language models with model lifecycle automation, runtime selection, and GPU scheduling. * [Olla v0.0.29](https://github.com/thushan/olla) – High-performance LLM proxy and load balancer providing intelligent routing, automatic failover, and unified model discovery across local and remote inference backends. * [Atlas Inference Engine b236](https://github.com/Avarok-Cybersecurity/atlas) – Pure Rust LLM inference engine focused on fast, hardware-optimized local model execution across GPUs and CPUs.