* [Ollama v0.40.0](https://github.com/ollama/ollama) – Tool for running and managing large language models. * [yzma v1.29.0](https://github.com/hybridgroup/yzma) – Go-based library for hardware-accelerated local inference with llama.cpp integration. * [HelixML 2.12.33](https://github.com/helixml/helix) – Private GenAI stack for deploying AI agents with support for RAG, API calls, vision, and efficient GPU scheduling. * [llama-swap v263](https://github.com/mostlygeek/llama-swap) – Reliable on-demand model switching between local OpenAI-compatible inference servers without restarting applications. * [llama.rn v0.13.0-rc.7](https://github.com/mybigday/llama.rn) – React Native binding for running LLaMA model inference with multimodal support including vision and audio.