* [Ollama v0.33.0](https://github.com/ollama/ollama) – Tool for running and managing large language models. * [Gerbil v1.28.0](https://github.com/lone-cloud/gerbil) – Desktop app for running Large Language Models locally with cross-platform support and integrated image generation. * [yzma v1.25.0](https://github.com/hybridgroup/yzma) – Go-based library for hardware-accelerated local inference with llama.cpp integration. * [llama\_cpp.rb v0.28.0](https://github.com/yoshoku/llama_cpp.rb) – Ruby bindings for llama.cpp, enabling easy integration of the library in Ruby applications. * [HelixML 2.12.7](https://github.com/helixml/helix) – Private GenAI stack for deploying AI agents with support for RAG, API calls, vision, and efficient GPU scheduling. * [go-llama v0.2.3](https://github.com/goccy/go-llama) – Pure Go inference engine for GGUF models, built from llama.cpp compiled to WebAssembly and translated to Go without wasm runtime. * [QVAC sdk-v0.18.2](https://github.com/tetherto/qvac) – Local-first, cross-platform SDK for building peer-to-peer AI apps with local model inference, speech, translation, and RAG. * [llama.rn v0.13.0-rc.1](https://github.com/mybigday/llama.rn) – React Native binding for running LLaMA model inference with multimodal support including vision and audio. * [llama-swap v251](https://github.com/mostlygeek/llama-swap) – Reliable on-demand model switching between local OpenAI-compatible inference servers without restarting applications.