* [node-llama-cpp v3.20.0](https://github.com/withcatai/node-llama-cpp) – Run AI models locally with Node.js bindings for llama.cpp. * [Off Grid AI v0.0.107](https://github.com/off-grid-ai/OGAM) – Offline on-device AI suite for chat, image generation, vision, speech-to-text, and tool calling across mobile and macOS, using local GGUF and hardware acceleration. * [ggrun v3.2.8](https://github.com/raketenkater/ggrun) – Auto-tuned launcher that measures multi-GPU hardware for GGUF models, picks an optimal llama.cpp/ik\_llama.cpp backend, and serves an OpenAI-compatible API. * [llama.rn v0.12.9](https://github.com/mybigday/llama.rn) – React Native binding for running LLaMA model inference with multimodal support including vision and audio.